Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K5 likes629 downloads27d agoHugging Face02Kennethdot /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K4 likes315 downloads2mo agoHugging Face03MohamedRashad /arabic-english-code-switching Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨ The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning. Citation If you use this dataset, please cite it as follows: @misc{rashad2024arabic, author = {Mohamed Rashad}, title = {arabic-english-code-switching}, year = {2024}, publisher = {Hugging Face}, url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.audioautomatic-speech-recognition10K<n<100K35 likes286 downloads2y agoHugging Face04ServiceNow-AI /asr_codeswitchedaudio1K<n<10K6 likes166 downloads3mo agoHugging Face05abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes160 downloads2mo agoHugging Face06nlpai-lab /ko_commongen_v2_code_switching 🇰🇷🇺🇸🇯🇵🇨🇳🇪🇸 KoCommonGEN v2 Code-switching This KoCommonGEN v2 Code-switching dataset consists of 99 samples for numerical commonsense reasoning, which were created relying on machine translation. The dataset can be found on Hugging Face at: nlpai-lab/ko_commongen_v2_code_switching This dataset contains code-switching data for the following languages: Korean (korean) English (english) Japanese (japan) Chinese (china) Spanish (espanol) (The code-switching data relies on… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/ko_commongen_v2_code_switching.textn<1K1 likes159 downloads2y agoHugging Face07Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes155 downloads5mo agoHugging Face08Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes129 downloads1y agoHugging Face09thetaone-ai /Korean-Japanese-Code-Switching-Speech Korean-Japanese-Code-Switching-Speech This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs. Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset. The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.audioautomatic-speech-recognitionn<1K4 likes119 downloads4mo agoHugging Face10georgechang8 /code_switch_yodas_zh Dataset Card for code-switching yodas This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon. Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.audio10K<n<100K4 likes104 downloads2y agoHugging Face11BrunoHays /fleurs_code_switching_test FLEURS Code-Switching Evaluation Set Dataset Summary This dataset is a synthetic code-switching evaluation set built from the google/fleurs corpus.Each sample is a single long-form audio sequence (minimum 5 minutes by default) composed by concatenating short utterances from multiple languages. The goal is to provide a controlled benchmark for testing ASR robustness when language switches happen frequently inside one recording. How The Dataset Was Curated… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/fleurs_code_switching_test.audio1K<n<10K0 likes87 downloads6mo agoHugging Face12code-switching /text-summarizationtextsummarizationn<1K0 likes78 downloads1mo agoHugging Face13ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K0 likes76 downloads2mo agoHugging Face14oist /algerian-code-switched-asr-eval Dataset Description This dataset contains 335 manually transcribed audio segments (approximately 30 minutes total) extracted from Algerian political and misinformation-related social-media videos, collected from public interviews on TikTok and YouTube in September 2025. It was built as the evaluation set for What WER Hides: A Closer Look at Algerian Code-Switched ASR, and is intended for benchmarking automatic speech recognition (ASR) systems on Algerian dialectal and… See the full description on the dataset page: https://huggingface.co/datasets/oist/algerian-code-switched-asr-eval.audion<1K0 likes67 downloads24d agoHugging Face15ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech-ipa KasaSpeech English–Twi Code-Switching Speech — IPA A phonemised version of ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech (KasaSpeech) with one added column: ipa. Every other column — audio included — is carried over byte-for-byte, and row order is unchanged, so this dataset aligns one-to-one with the original. The ipa column Each transcript is converted to a phoneme sequence with ghanag2p-uni, the Twi-only grapheme-to-phoneme library built on ghana-g2p… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech-ipa.audiotext-to-speech10K<n<100K0 likes65 downloads2mo agoHugging Face16devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes61 downloads7mo agoHugging Face17Panhapich /khmer-english-codeswitch-tts Khmer–English Code-Switch Synthetic Speech 8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech, generated with VoxCPM2 from code-switch text manufactured by confirmed lexical substitution over a Khmer–English parallel corpus. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and the code-switch sentences were manufactured by word substitution — they are not transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.audioautomatic-speech-recognition1K<n<10K2 likes57 downloads2mo agoHugging Face18HuggingPanda /SDAIANCAI-Saudilang-Code-Switch-Corpusaudio1K<n<10K2 likes56 downloads2y agoHugging Face19BrunoHays /english-en-x-code-switching-main-lang English EN-X Code-Switching Main-Language This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.audioautomatic-speech-recognitionn<1K0 likes53 downloads5mo agoHugging Face20Moamen-dcp /arazn_codeSwitched_mp3_full_notLoweraudio1K<n<10K0 likes45 downloads1y agoHugging Face21lxyuan /nemo-codeswitch-reasoning-debate Overview This is a synthetic, multilingual code-switching dataset. Each record contains: a realistic user query a long-form reasoning section a debate / counterargument section a concise final_answer It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses. This snapshot contains 574,977 rows and 10 string columns. Data provenance Important: Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.texttext-generation100K<n<1M0 likes44 downloads7mo agoHugging Face22Atufa /codeswitch-fr-en-kyutai-stt Code-Switched French–English STT Probe Dataset Dataset Summary This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.audioautomatic-speech-recognitionn<1K0 likes44 downloads7mo agoHugging Face23BrunoHays /english-x-code-switching Synthetic English Code-Switching Evaluation Set This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.audioautomatic-speech-recognitionn<1K0 likes43 downloads5mo agoHugging Face24MINERVA-TEAM /minerva-ar-en-edu-codeswitch-datasetaudio1K<n<10K0 likes38 downloads8mo agoHugging Face25yangzhang33 /CEB_code_switchedtabular10K<n<100K0 likes35 downloads5mo agoHugging Face26lilgoose777 /nepali-podcast-code-switch-asr-datasetaudio1K<n<10K0 likes32 downloads5mo agoHugging Face27code-switching /topic-classificationtabulartext-classificationn<1K0 likes29 downloads1mo agoHugging Face28zenyn /Code-Switching-Testaudion<1K0 likes27 downloads2y agoHugging Face29gimmy256 /african-codeswitching Pan-African Code-Switching Dataset Built with Adaptive Data by Adaption | Crane AI Labs Submitted to the Uncharted Data Challenge 2026 by Adaption Labs Overview The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations. Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.texttext-classificationn<1K0 likes25 downloads6mo agoHugging Face30BrunoHays /english-x-code-switching-samples Synthetic English Code-Switching Evaluation Set Samples This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.audioautomatic-speech-recognition10K<n<100K0 likes25 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.