Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Maxyelow /kenyan-code-switch-instruct-50k 🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs) A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules. Dataset Summary Total Samples: 50,000 instruction-response pairs train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.texttext-generation10K<n<100K0 likes78 downloads12d agoHugging Face02years0 /multilingual-code-switching-bench Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application) Question 1: Critical Blind Spot & Capability Gap Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/years0/multilingual-code-switching-bench.texttext-generationn<1K0 likes43 downloads3d agoHugging Face03tdnathmlenthusiast /qwen2.5-3b-codeswitch-blindspot Dataset Card: Code-Switched Agentic Reasoning Eval (Bengali / Hindi / Arabic) 16 hand-designed, adversarially-constructed reasoning tasks, each provided in 7 parallel language conditions (same semantic content, same gold answer, only the surface language changes): English, Bengali, Banglish, Hindi, Hinglish, Arabic, Arabilish. Built for evaluating whether a tool-using agent's correctness, calibration under ambiguity, and (separately, via the accompanying notebook)… See the full description on the dataset page: https://huggingface.co/datasets/tdnathmlenthusiast/qwen2.5-3b-codeswitch-blindspot.textquestion-answeringn<1K0 likes36 downloads6d agoHugging Face04MouradGad /multilingual-code-switching-bench Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application) Question 1: Critical Blind Spot & Capability Gap Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/MouradGad/multilingual-code-switching-bench.texttext-generationn<1K0 likes36 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.