Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rshahbaz /pragmatic-code-switch-blindspot Pragmatic Blind Spots Under Roman Urdu Framing Question 1: the blind spot The central blind spot in multilingual evaluation is that models frequently lose pragmatic and social meaning under Roman Urdu framing, even when they understand the individual words. I write in Roman Urdu myself, and I regularly see language models misunderstand the intended meaning in everyday exchanges. While standard multilingual benchmarks evaluate formal Perso-Arabic Urdu script or… See the full description on the dataset page: https://huggingface.co/datasets/rshahbaz/pragmatic-code-switch-blindspot.textn<1K0 likes137 downloads11d agoHugging Face02Praxel /codeswitch-pairs-lase Codeswitch Pairs LASE — training corpus 1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair). Schema (manifest.jsonl) { "voice_id": "21m00Tcm4TlvDq8ikWAM", "lang": "en | hi | te | ta", "text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.audioaudio-classificationn<1K0 likes70 downloads5mo agoHugging Face03Maxyelow /kenyan-code-switch-instruct-50k 🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs) A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules. Dataset Summary Total Samples: 50,000 instruction-response pairs train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.texttext-generation10K<n<100K0 likes70 downloads9d agoHugging Face04oumayma03 /adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented) This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.text1K<n<10K0 likes60 downloads20d agoHugging Face05Maxyelow /kenyan-code-switch-1m 🇰🇪 Kenyan Code-Switching Pretraining Corpus (1,000,000 Sentences / 16.99M Words) The largest standardized, rule-audited monolingual pretraining corpus for Kenyan Code-Switching (Sheng / Technical Swahili-English Blend). Corpus Statistics Total Sentences: 1,000,000 sentences Total Word Tokens: 16,990,653 words Average Sentence Length: 16.99 words (multi-clause explanatory syntax) Linguistic Audit Score: 100.0000% compliance across all 20 Master Blueprint rules… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-1m.text1M<n<10M0 likes40 downloads9d agoHugging Face06tdnathmlenthusiast /qwen2.5-3b-codeswitch-blindspot Dataset Card: Code-Switched Agentic Reasoning Eval (Bengali / Hindi / Arabic) 16 hand-designed, adversarially-constructed reasoning tasks, each provided in 7 parallel language conditions (same semantic content, same gold answer, only the surface language changes): English, Bengali, Banglish, Hindi, Hinglish, Arabic, Arabilish. Built for evaluating whether a tool-using agent's correctness, calibration under ambiguity, and (separately, via the accompanying notebook)… See the full description on the dataset page: https://huggingface.co/datasets/tdnathmlenthusiast/qwen2.5-3b-codeswitch-blindspot.textquestion-answeringn<1K0 likes33 downloads4d agoHugging Face07abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes26 downloads2mo agoHugging Face08Abhisingh-18 /hindi-english-codeswitch-dataset Hindi-English Code-Switch ASR Transcripts Text transcripts and metadata for a large bilingual Hindi-English code-switch ASR training corpus, used to train Abhisingh-18/hindi-english-codeswitch-asr. This release contains transcripts and metadata only — no audio files. Audio was sourced from multiple corpora and institutions and is not redistributed here. Credits Speech data collection and curation credit: SPRING Lab, IIT Madras. Contents File… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/hindi-english-codeswitch-dataset.textautomatic-speech-recognition1M<n<10M0 likes24 downloads2mo agoHugging Face09years0 /multilingual-code-switching-bench Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application) Question 1: Critical Blind Spot & Capability Gap Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/years0/multilingual-code-switching-bench.texttext-generationn<1K0 likes23 downloads15h agoHugging Face10MouradGad /multilingual-code-switching-bench Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application) Question 1: Critical Blind Spot & Capability Gap Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/MouradGad/multilingual-code-switching-bench.texttext-generationn<1K0 likes16 downloads15h agoHugging Face11gimmy256 /east-african-code-switching This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. east_african_code_switching This dataset contains utterances demonstrating code-switching between English and various East African languages, including Swahili, Luganda, and Acholi. Each sample is annotated with metadata such as region, intent, domain, sentiment, and specific switch types to support linguistic analysis. The collection covers diverse conversational registers… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/east-african-code-switching.textn<1K0 likes14 downloads6mo agoHugging Face12aafrinaaysha /dravidian-codeswitch Dravidian CodeSwitch Word-level Language Dataset AutoTrain-ready dataset for token classification. textn<1K0 likes6 downloads1y agoHugging Face13kaumudi-ai /mal-en-code-switchtextn<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.