datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
neuromoyo-sahara-codeswitch-benchmark
NEUROMOYO — Sahara CodeSwitch Africa Benchmark
🔗 Live Benchmark Results
Interactive benchmark:
https://www.neuromoyo.app/benchmark
This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech.
🚀 Live NEUROMOYO Demo
Live application:
https://www.neuromoyo.app
The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.codeswitch-pairs-lase
Codeswitch Pairs LASE — training corpus
1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"lang": "en | hi | te | ta",
"text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.topic-classificationSaudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.CEB_code_switchednaturalnessarabic-english-code-switching-review-annotations
Review Annotations for Arabic-English Code-Switching Speech
This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.
The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.
Coverage and outcomes
The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.code-switched-medical-blindspot-eval
Code-Switched Medical Query Evaluation: Qwen2.5-1.5B-Instruct
Blind Spot
Standard AI benchmarks mostly test clean English or native-script text in a single language. They rarely test how many people in Pakistan actually type when describing a health problem: Roman-script Urdu mixed mid-sentence with English ("chest mein tightness hai"), or Sindhi. This register is informal, has no standard spelling, and is under-represented in training data and evaluation sets.… See the full description on the dataset page: https://huggingface.co/datasets/AsadChanna/code-switched-medical-blindspot-eval.progressive-code-switch
MehdiAstaraki/progressive-code-switch
Progressive code-switching retrieval-decay benchmark. Each base patent yields a cumulative ladder of documents (base__r0 clean → base__rN), where each step swaps one more chemistry term into another language / spelling / ChEBI form. One fixed question per base (about the step-1 term) is reused for every depth; the qrels carry the depth so you can measure how retrieval decays as more terms are code-switched.
Configs: corpus (ladder variant… See the full description on the dataset page: https://huggingface.co/datasets/MehdiAstaraki/progressive-code-switch.
