Team Ai
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes5.8k downloads5mo agoHugging Face02ghanaopenai /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K5 likes629 downloads27d agoHugging Face03Kennethdot /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K4 likes315 downloads2mo agoHugging Face04MohamedRashad /arabic-english-code-switching Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨ The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning. Citation If you use this dataset, please cite it as follows: @misc{rashad2024arabic, author = {Mohamed Rashad}, title = {arabic-english-code-switching}, year = {2024}, publisher = {Hugging Face}, url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.audioautomatic-speech-recognition10K<n<100K35 likes286 downloads2y agoHugging Face05abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes160 downloads2mo agoHugging Face06Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes155 downloads5mo agoHugging Face07thetaone-ai /Korean-Japanese-Code-Switching-Speech Korean-Japanese-Code-Switching-Speech This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs. Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset. The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.audioautomatic-speech-recognitionn<1K4 likes119 downloads4mo agoHugging Face08shangeth /mls-mimi-codes Multilingual LibriSpeech (MLS) — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for Multilingual LibriSpeech — LibriVox audiobooks in 7 non-English languages. English is intentionally excluded. For English Mimi codes, use: shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits) shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native) shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents) shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.tabulartext-to-speech1M<n<10M0 likes107 downloads5mo agoHugging Face09ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K0 likes76 downloads2mo agoHugging Face10ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech-ipa KasaSpeech English–Twi Code-Switching Speech — IPA A phonemised version of ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech (KasaSpeech) with one added column: ipa. Every other column — audio included — is carried over byte-for-byte, and row order is unchanged, so this dataset aligns one-to-one with the original. The ipa column Each transcript is converted to a phoneme sequence with ghanag2p-uni, the Twi-only grapheme-to-phoneme library built on ghana-g2p… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech-ipa.audiotext-to-speech10K<n<100K0 likes65 downloads2mo agoHugging Face11shangeth /librispeech-mimi-codes LibriSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LibriSpeech corpus — multi-speaker English audiobook readings from the LibriVox project. This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech mirrors (e.g. openslr/librispeech_asr); these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models. Schema One row per utterance: Column Type Notes id string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.tabulartext-to-speech100K<n<1M0 likes62 downloads5mo agoHugging Face12devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes61 downloads7mo agoHugging Face13shangeth /expresso-mimi-codes Expresso — Mimi Codes (k = 32) Pre-extracted Kyutai Mimi tokens (all 32 codebooks) for both the read and conversational subsets of Expresso. Source audio + transcripts live in shangeth/expresso; this dataset publishes the discrete-token version for training Mimi-based speech models without re-extracting. ⚠️ License: CC-BY-NC-4.0 — non-commercial use only. Why Expresso for Wren? Expresso is the most directly relevant dataset for speech disentanglement research — the… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso-mimi-codes.tabulartext-to-speech10K<n<100K1 likes59 downloads5mo agoHugging Face14Panhapich /khmer-english-codeswitch-tts Khmer–English Code-Switch Synthetic Speech 8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech, generated with VoxCPM2 from code-switch text manufactured by confirmed lexical substitution over a Khmer–English parallel corpus. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and the code-switch sentences were manufactured by word substitution — they are not transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.audioautomatic-speech-recognition1K<n<10K2 likes57 downloads2mo agoHugging Face15BrunoHays /english-en-x-code-switching-main-lang English EN-X Code-Switching Main-Language This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.audioautomatic-speech-recognitionn<1K0 likes53 downloads5mo agoHugging Face16shangeth /ljspeech-mimi-codes LJSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LJSpeech corpus — 13,100 English utterances from a single female speaker reading public-domain audiobook passages (~24 hours). This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech release; these codes are designed to be loaded alongside it for training Mimi-based speech models without paying the ~1 hour of GPU extraction cost. Schema One row per utterance:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.tabulartext-to-speech10K<n<100K0 likes52 downloads5mo agoHugging Face17Atufa /codeswitch-fr-en-kyutai-stt Code-Switched French–English STT Probe Dataset Dataset Summary This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.audioautomatic-speech-recognitionn<1K0 likes44 downloads7mo agoHugging Face18BrunoHays /english-x-code-switching Synthetic English Code-Switching Evaluation Set This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.audioautomatic-speech-recognitionn<1K0 likes43 downloads5mo agoHugging Face19BrunoHays /english-x-code-switching-samples Synthetic English Code-Switching Evaluation Set Samples This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.audioautomatic-speech-recognition10K<n<100K0 likes25 downloads5mo agoHugging Face20sajalmadan0909 /hindi_and_english_stt_tts_codemix_datagated Hindi and English STT/TTS Codemix Data Hinglish (Hindi-English code-mixed) speech dataset for automatic speech recognition (ASR) and text-to-speech (TTS) research. Dataset Description Each row is a timestamped speech segment clipped from conversational Hinglish audio recordings. Column Type Description text string Transcript of the speech segment (Hinglish) audio audio (16 kHz mono) Corresponding audio clip duration float32 Clip duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/hindi_and_english_stt_tts_codemix_data.audioautomatic-speech-recognition10K<n<100K0 likes19 downloads3mo agoHugging Face21yilele /synth-qa-taste-codec-chat Synthetic QA Taste-S Codec Chat 18571 single-turn Traditional Chinese QA utterances with synthesized speech, 21.6 hours of audio before codec extraction. Assistant speech is represented as: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> Each text token is followed by its 16 Taste-S FSQ codes (codebooks a..p). Configurations default — messages (user question + assistant <SAY> speech), audio, and answer text. Statistics Utterances: 18571… See the full description on the dataset page: https://huggingface.co/datasets/yilele/synth-qa-taste-codec-chat.audiotext-to-speech10K<n<100K0 likes18 downloads3mo agoHugging Face22zdm-code /england-phoneme-datasetgated British English Phonetic Dataset Introduction This dataset is an extension of Common Voice, from which 6 subsets were selected (Common Voice Corpus 1, Common Voice Corpus 2, Common Voice Corpus 3, Common Voice Corpus 4, Common Voice Corpus 18.0, Common Voice Corpus 19.0). All data containing the England accent from these 6 subsets were extracted and phonetically annotated accordingly. Description Key fields explanation: sentence: The English sentence… See the full description on the dataset page: https://huggingface.co/datasets/zdm-code/england-phoneme-dataset.audioautomatic-speech-recognition100K<n<1M5 likes14 downloads2y agoHugging Face23BrunoHays /english-en-x-code-switching-main-lang-samples English EN-X Code-Switching Main-Language Samples This dataset contains the individual full FLEURS utterance chunks used to build the paired mixed dataset. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples.audioautomatic-speech-recognition1K<n<10K0 likes14 downloads5mo agoHugging Face24BrunoHays /english-en-x-code-switching-main-lang-samples-merged English EN-X Code-Switching Main-Language Merged Samples This dataset contains contiguous same-language segments from the paired mixed dataset. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.audioautomatic-speech-recognitionn<1K0 likes14 downloads5mo agoHugging Face25pauljunsukhan /throatmic_coderedgated Throat Microphone Dataset 🤗 Hugging Face: pauljunsukhan/throatmic_codered📦 GitHub: pauljunsukhan/throatmicdata 🚀 Fine-tuned Model: 🤗 pauljunsukhan/throatmic_subvocalization_whisper 📦 pauljunsukhan/whisper_finetuning 🔐 Dataset Access: Downloading: The dataset is publicly available. Use download_dataset.py (see instructions below) Contributing: Contributions are very welcome! Request write access through the Hugging Face dataset page - I'd love to have more… See the full description on the dataset page: https://huggingface.co/datasets/pauljunsukhan/throatmic_codered.audioautomatic-speech-recognition1K<n<10K0 likes11 downloads2y agoHugging Face26WTFO /codeswitchinggated WTFO Code-Switching Speech Code-switching speech dataset prepared for automatic speech recognition training. Dataset fields audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature text: transcript duration: audio duration in seconds (float64) Split summary Split: train Examples: 98,662 Total duration: 567422.698 seconds (157.62 hours) The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.audioautomatic-speech-recognition10K<n<100K0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.