datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
neuromoyo-sahara-codeswitch-benchmark
NEUROMOYO — Sahara CodeSwitch Africa Benchmark
🔗 Live Benchmark Results
Interactive benchmark:
https://www.neuromoyo.app/benchmark
This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech.
🚀 Live NEUROMOYO Demo
Live application:
https://www.neuromoyo.app
The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.naturalnesscode-switched-medical-blindspot-eval
Code-Switched Medical Query Evaluation: Qwen2.5-1.5B-Instruct
Blind Spot
Standard AI benchmarks mostly test clean English or native-script text in a single language. They rarely test how many people in Pakistan actually type when describing a health problem: Roman-script Urdu mixed mid-sentence with English ("chest mein tightness hai"), or Sindhi. This register is informal, has no standard spelling, and is under-represented in training data and evaluation sets.… See the full description on the dataset page: https://huggingface.co/datasets/AsadChanna/code-switched-medical-blindspot-eval.Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset
Sinhala-English Hospitality Review Corpus
This dataset contains 4,133 user reviews collected from Facebook hotel pages across Sri Lanka, annotated at the sentence level for sentiment analysis. The reviews come from 400 hotels across 144 tourist destinations, including Colombo, Kandy, Anuradhapura, Negombo, and other major regions.
Annotation Details
All samples were manually annotated by the author to ensure consistency in sentiment labeling. We welcome volunteers… See the full description on the dataset page: https://huggingface.co/datasets/nlp-dataset/Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset.fon-code-switching-evaluation
French-Fon Code-Switching Evaluation Benchmark
Overview
This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios.
The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon.
The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.
