Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes200 downloads3mo agoHugging Face02devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes61 downloads7mo agoHugging Face03zenyn /Code-Switching-Testaudion<1K0 likes27 downloads2y agoHugging Face04Resvand /Code-Switching_English-Hindi_Syntheticaudio1K<n<10K0 likes23 downloads11mo agoHugging Face05DynamicSuperb /CodeSwitchingSpeechIdentification_ASCENDaudion<1K0 likes21 downloads2y agoHugging Face06DynamicSuperb /CodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-enaudion<1K0 likes19 downloads2y agoHugging Face07Nash-pAnDiTa /ASR_En_Ar_CodeSwitchingaudio10K<n<100K0 likes17 downloads2y agoHugging Face08WTFO /codeswitchinggated WTFO Code-Switching Speech Code-switching speech dataset prepared for automatic speech recognition training. Dataset fields audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature text: transcript duration: audio duration in seconds (float64) Split summary Split: train Examples: 98,662 Total duration: 567422.698 seconds (157.62 hours) The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.audioautomatic-speech-recognition10K<n<100K0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.