datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kenyan-code-switch-instruct-50k
🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs)
A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules.
Dataset Summary
Total Samples: 50,000 instruction-response pairs
train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.multilingual-code-switching-bench
Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application)
Question 1: Critical Blind Spot & Capability Gap
Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/years0/multilingual-code-switching-bench.qwen2.5-3b-codeswitch-blindspot
Dataset Card: Code-Switched Agentic Reasoning Eval (Bengali / Hindi / Arabic)
16 hand-designed, adversarially-constructed reasoning tasks, each provided in 7 parallel
language conditions (same semantic content, same gold answer, only the surface language
changes): English, Bengali, Banglish, Hindi, Hinglish, Arabic, Arabilish.
Built for evaluating whether a tool-using agent's correctness, calibration under
ambiguity, and (separately, via the accompanying notebook)… See the full description on the dataset page: https://huggingface.co/datasets/tdnathmlenthusiast/qwen2.5-3b-codeswitch-blindspot.multilingual-code-switching-bench
Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application)
Question 1: Critical Blind Spot & Capability Gap
Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/MouradGad/multilingual-code-switching-bench.
