Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Perle-ai /ASR_Code_Switch ASR Code-Switching Benchmark A curated benchmark of 1,200 code-switching utterances (300 per language pair) for evaluating commercial ASR systems on multilingual speech with intra-sentential language switching. Paper Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German arXiv link Language pairs Split Language pair Samples Scripts egyptian_arabic_english Egyptian Arabic–English 300 Arabic + Latin… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.audioautomatic-speech-recognition1K<n<10K12 likes1k downloads5mo agoHugging Face02TigreGotico /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes631 downloads3mo agoHugging Face03ghanaopenai /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K5 likes430 downloads24d agoHugging Face04Kennethdot /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K4 likes307 downloads2mo agoHugging Face05Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_new10K<n<100K0 likes277 downloads1y agoHugging Face06FatimahEmadEldin /cafe-algerian-codeswitch-speech CAFE Algerian Codeswitch Speech This dataset contains Algerian Arabic and French code-switched speech. Repository Path: FatimahEmadEldin/cafe-algerian-codeswitch-speech audio1K<n<10K0 likes265 downloads4mo agoHugging Face07MohamedRashad /arabic-english-code-switching Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨ The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning. Citation If you use this dataset, please cite it as follows: @misc{rashad2024arabic, author = {Mohamed Rashad}, title = {arabic-english-code-switching}, year = {2024}, publisher = {Hugging Face}, url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.audioautomatic-speech-recognition10K<n<100K35 likes263 downloads2y agoHugging Face08nlpai-lab /ko_commongen_v2_code_switching 🇰🇷🇺🇸🇯🇵🇨🇳🇪🇸 KoCommonGEN v2 Code-switching This KoCommonGEN v2 Code-switching dataset consists of 99 samples for numerical commonsense reasoning, which were created relying on machine translation. The dataset can be found on Hugging Face at: nlpai-lab/ko_commongen_v2_code_switching This dataset contains code-switching data for the following languages: Korean (korean) English (english) Japanese (japan) Chinese (china) Spanish (espanol) (The code-switching data relies on… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/ko_commongen_v2_code_switching.textn<1K1 likes216 downloads2y agoHugging Face09liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes214 downloads3mo agoHugging Face10NLPC-UOM /Sinhala-English-Code-Mixed-Code-Switched-Dataset Sinhala-English-Code-Mixed-Code-Switched-Dataset This dataset contains 10,000 comments that have been annotated at the sentence level for sentiment analysis, humor detection, hate speech detection, aspect identification, and language identification. The following is the tag scheme. Sentiment - Positive, Negative, Neutral, Conflict Humor - Humorous, Non humorous Hate Speech - Hate-Inducing, Abusive, Not offensive Aspect - Network, Billing or Price, Package, Customer Service, Data… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-English-Code-Mixed-Code-Switched-Dataset.text-classification6 likes205 downloads2y agoHugging Face11Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes197 downloads1y agoHugging Face12ServiceNow-AI /asr_codeswitchedaudio1K<n<10K6 likes169 downloads3mo agoHugging Face13Kimyayd /vocal-money-codeswitch-asr-benchmark Vocal Money — Yoruba–English Code-Switched ASR Benchmark A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally code-switched Yoruba–English speech, together with the reference transcriptions and the output of every system on every clip, so that the published results can be recomputed or contradicted. Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026. Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes156 downloads2mo agoHugging Face14Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo10K<n<100K0 likes153 downloads1y agoHugging Face15fiewor /gradrai-viva-codeswitch-benchmark GradrAI Viva Code-Switched Oral Benchmark Consented, de-identified classroom-style oral answer clips used to benchmark GradrAI Viva for the Sahara CodeSwitch Africa challenge. Contents metadata.csv / metadata.jsonl: one row per clip. audio/: 16 kHz mono WAV files for Hugging Face dataset preview and ASR reuse. audio_original/: original submitted browser/Opus/WebM audio files. benchmark/: benchmark outputs (results.md, results.json) and manifest used by GradrAI… See the full description on the dataset page: https://huggingface.co/datasets/fiewor/gradrai-viva-codeswitch-benchmark.automatic-speech-recognitionn<1K1 likes148 downloads21d agoHugging Face16Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes147 downloads4mo agoHugging Face17abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes142 downloads2mo agoHugging Face18shulhaaja /id-en-codeswitch-dataset-alternative Indonesian–English Code-Switching Synthetic Speech Dataset Synthetic speech generated for the undergraduate final project "Handling Code-Switching in Automatic Speech Recognition for Low-Resource Language Pairs: An Indonesian–English Case Study", School of Electrical Engineering and Informatics, Institut Teknologi Bandung. This dataset contains synthetic audio produced from the Indonesian–English code-switching text corpora released in the companion repository below. It was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.audioautomatic-speech-recognition10K<n<100K0 likes141 downloads2mo agoHugging Face19rshahbaz /pragmatic-code-switch-blindspot Pragmatic Blind Spots Under Roman Urdu Framing Question 1: the blind spot The central blind spot in multilingual evaluation is that models frequently lose pragmatic and social meaning under Roman Urdu framing, even when they understand the individual words. I write in Roman Urdu myself, and I regularly see language models misunderstand the intended meaning in everyday exchanges. While standard multilingual benchmarks evaluate formal Perso-Arabic Urdu script or… See the full description on the dataset page: https://huggingface.co/datasets/rshahbaz/pragmatic-code-switch-blindspot.textn<1K0 likes137 downloads10d agoHugging Face20thetaone-ai /Korean-Japanese-Code-Switching-Speech Korean-Japanese-Code-Switching-Speech This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs. Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset. The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.audioautomatic-speech-recognitionn<1K4 likes124 downloads4mo agoHugging Face21Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_prepared_4_whisper_turbo_transcription1K<n<10K0 likes111 downloads1y agoHugging Face22georgechang8 /code_switch_yodas_zh Dataset Card for code-switching yodas This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon. Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.audio10K<n<100K4 likes109 downloads2y agoHugging Face23BrunoHays /fleurs_code_switching_test FLEURS Code-Switching Evaluation Set Dataset Summary This dataset is a synthetic code-switching evaluation set built from the google/fleurs corpus.Each sample is a single long-form audio sequence (minimum 5 minutes by default) composed by concatenating short utterances from multiple languages. The goal is to provide a controlled benchmark for testing ASR robustness when language switches happen frequently inside one recording. How The Dataset Was Curated… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/fleurs_code_switching_test.audio1K<n<10K0 likes107 downloads6mo agoHugging Face24ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K0 likes105 downloads2mo agoHugging Face25Tim2190 /kazakh-codeswitch-asr Kazakh Code-Switching ASR Benchmark A benchmark for evaluating ASR systems on natural Kazakh speech that code-switches with Russian — the everyday Kazakh–Russian mixing found in stand-up, interviews and vlogs, not scripted read speech. This is, to our knowledge, the first speech/ASR resource targeting the Kazakh–Russian code-switching pair (existing Kazakh–Russian NLP resources are text-only). Code, scoring harness and full analysis:… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kazakh-codeswitch-asr.audioautomatic-speech-recognitionn<1K0 likes100 downloads2mo agoHugging Face26mosesdaudu /switchboard-tierb-codeswitch SwitchBoard Tier B — African code-switched speech 87 consented utterances of intra-sentential code-switching — Nigerian Pidgin, Yorùbá, Hausa and Kiswahili each mixed with English inside a single sentence — recorded from 8 bilingual volunteers at the Deep Learning Indaba 2026, Lagos. Collected for the MLC (Africa) × Intron Agentic Voice AI Challenge as an evaluation set for telco/fintech voice agents. 8.75 minutes total. What this is for Measuring whether a speech… See the full description on the dataset page: https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch.audioautomatic-speech-recognitionn<1K0 likes95 downloads2mo agoHugging Face27Panhapich /khmer-english-codeswitch-tts Khmer–English Code-Switch Synthetic Speech 8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech, generated with VoxCPM2 from code-switch text manufactured by confirmed lexical substitution over a Khmer–English parallel corpus. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and the code-switch sentences were manufactured by word substitution — they are not transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.audioautomatic-speech-recognition1K<n<10K2 likes91 downloads2mo agoHugging Face28prokelly /neuromoyo-sahara-codeswitch-benchmark NEUROMOYO — Sahara CodeSwitch Africa Benchmark 🔗 Live Benchmark Results Interactive benchmark: https://www.neuromoyo.app/benchmark This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech. 🚀 Live NEUROMOYO Demo Live application: https://www.neuromoyo.app The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.tabularn<1K0 likes89 downloads16d agoHugging Face29Rabe3 /saudi-english-code-switching-datasetaudio10K<n<100K0 likes87 downloads8mo agoHugging Face30code-switching /question-answertext1K<n<10K0 likes86 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.