datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pragmatic-code-switch-blindspot
Pragmatic Blind Spots Under Roman Urdu Framing
Question 1: the blind spot
The central blind spot in multilingual evaluation is that models frequently lose pragmatic and social meaning under Roman Urdu framing, even when they understand the individual words. I write in Roman Urdu myself, and I regularly see language models misunderstand the intended meaning in everyday exchanges. While standard multilingual benchmarks evaluate formal Perso-Arabic Urdu script or… See the full description on the dataset page: https://huggingface.co/datasets/rshahbaz/pragmatic-code-switch-blindspot.codeswitch-pairs-lase
Codeswitch Pairs LASE — training corpus
1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"lang": "en | hi | te | ta",
"text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.kenyan-code-switch-instruct-50k
🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs)
A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules.
Dataset Summary
Total Samples: 50,000 instruction-response pairs
train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented)
This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.kenyan-code-switch-1m
🇰🇪 Kenyan Code-Switching Pretraining Corpus (1,000,000 Sentences / 16.99M Words)
The largest standardized, rule-audited monolingual pretraining corpus for Kenyan Code-Switching (Sheng / Technical Swahili-English Blend).
Corpus Statistics
Total Sentences: 1,000,000 sentences
Total Word Tokens: 16,990,653 words
Average Sentence Length: 16.99 words (multi-clause explanatory syntax)
Linguistic Audit Score: 100.0000% compliance across all 20 Master Blueprint rules… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-1m.qwen2.5-3b-codeswitch-blindspot
Dataset Card: Code-Switched Agentic Reasoning Eval (Bengali / Hindi / Arabic)
16 hand-designed, adversarially-constructed reasoning tasks, each provided in 7 parallel
language conditions (same semantic content, same gold answer, only the surface language
changes): English, Bengali, Banglish, Hindi, Hinglish, Arabic, Arabilish.
Built for evaluating whether a tool-using agent's correctness, calibration under
ambiguity, and (separately, via the accompanying notebook)… See the full description on the dataset page: https://huggingface.co/datasets/tdnathmlenthusiast/qwen2.5-3b-codeswitch-blindspot.arabic-english-code-switching-review-annotations
Review Annotations for Arabic-English Code-Switching Speech
This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.
The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.
Coverage and outcomes
The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.hindi-english-codeswitch-dataset
Hindi-English Code-Switch ASR Transcripts
Text transcripts and metadata for a large bilingual Hindi-English code-switch ASR training corpus, used to train Abhisingh-18/hindi-english-codeswitch-asr.
This release contains transcripts and metadata only — no audio files. Audio was sourced from multiple corpora and institutions and is not redistributed here.
Credits
Speech data collection and curation credit: SPRING Lab, IIT Madras.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/hindi-english-codeswitch-dataset.multilingual-code-switching-bench
Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application)
Question 1: Critical Blind Spot & Capability Gap
Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/years0/multilingual-code-switching-bench.multilingual-code-switching-bench
Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application)
Question 1: Critical Blind Spot & Capability Gap
Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/MouradGad/multilingual-code-switching-bench.east-african-code-switching
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
east_african_code_switching
This dataset contains utterances demonstrating code-switching between English and various East African languages, including Swahili, Luganda, and Acholi. Each sample is annotated with metadata such as region, intent, domain, sentiment, and specific switch types to support linguistic analysis. The collection covers diverse conversational registers… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/east-african-code-switching.dravidian-codeswitch
Dravidian CodeSwitch Word-level Language Dataset
AutoTrain-ready dataset for token classification.
mal-en-code-switch
