Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fixie-ai /common_voice_17_0audio10M<n<100M18 likes210k downloads2y agoHugging Face02espnet /yodas3 YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data. For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.audioaudio-to-audio1M<n<10M203 likes128k downloads5d agoHugging Face03ylacombe /cml-tts Dataset Card for CML-TTS Dataset Summary CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.audiotext-to-speech1M<n<10M36 likes119k downloads3y agoHugging Face04google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M473 likes113k downloads5mo agoHugging Face05hf-internal-testing /librispeech_asr_dummyaudion<1K12 likes109k downloads2y agoHugging Face06fixie-ai /covost2This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer. The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger. As such, not all the data is included: Only the validation and test subsets are available. From the XX_EN subsets, only fr, es, and zh-CN are included. audio1M<n<10M5 likes74k downloads2y agoHugging Face07google /svq Simple Voice Questions Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions. It serves as a core evaluation componenet for Massive Sound Embedding Benchmark (MSEB). Technical Specifications Feature Details Locales 26 Languages 17 Total Speakers ~700 (Capped at 250 recordings per speaker) Audio Conditions Clean, Background Speech, Media, Traffic Noise Gender… See the full description on the dataset page: https://huggingface.co/datasets/google/svq.audioquestion-answering1M<n<10M62 likes57k downloads11d agoHugging Face08openslr /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.audioautomatic-speech-recognition100K<n<1M245 likes54k downloads1y agoHugging Face09japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes50k downloads2y agoHugging Face10agkphysics /AudioSet Dataset Card for AudioSet Dataset Summary AudioSet is a dataset of 10-second clips from YouTube, annotated into one or more sound categories, following the AudioSet ontology. Supported Tasks and Leaderboards audio-classification: Classify audio clips into categories. The leaderboard is available here Languages The class labels in the dataset are in English. Dataset Structure Data Instances Example… See the full description on the dataset page: https://huggingface.co/datasets/agkphysics/AudioSet.audioaudio-classification1M<n<10M109 likes47k downloads1y agoHugging Face11facebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M191 likes44k downloads2y agoHugging Face12MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M286 likes35k downloads2y agoHugging Face13disco-eth /EuroSpeech EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.audioautomatic-speech-recognition10M<n<100M99 likes34k downloads5mo agoHugging Face14facebook /voxpopuli Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.audioautomatic-speech-recognition1M<n<10M166 likes26k downloads8mo agoHugging Face15ARTPARK-IISc /VaanigatedVAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.audioautomatic-speech-recognition1M<n<10M160 likes24k downloads19d agoHugging Face16Hothan /Omni-DuplexEval Omni-DuplexEval 📖 arXiv | GitHub Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal interaction. Unlike conventional offline video understanding benchmarks, Omni-DuplexEval focuses on streaming settings where models must continuously process evolving multimodal inputs and decide what to respond and when to respond. The benchmark contains two scenarios: Real-Time Description (RTD): evaluates continuous streaming description ability. Proactive Reminder (PR):… See the full description on the dataset page: https://huggingface.co/datasets/Hothan/Omni-DuplexEval.audion<1K0 likes22k downloads5mo agoHugging Face17ai4bharat /IndicVoicesgated IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates [23 December 2025] We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.audio1M<n<10M122 likes22k downloads4mo agoHugging Face18lmms-lab /EgoIT-99KCheckout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information. audio100K<n<1M9 likes21k downloads2y agoHugging Face19MERaLiON /Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching. ASR: Automatic Speech Recognition SQA: Speech Question Answering SDS: Spoken Dialogue Summarization PQA: Paralinguistic Question Answering from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.audio10M<n<100M22 likes20k downloads2y agoHugging Face20jkot /parliament_hearings_processed Preprocessed parliament hearings ASR dataset to truecased form. Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126 dataset_info: features: - name: id dtype: string - name: audio dtype: audio: sampling_rate: 16000 - name: transcription sequence: string splits: - name: train num_bytes: 53645064353.18 num_examples: 191455 - name: test num_bytes: 740331298.0 num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.audio100K<n<1M1 likes20k downloads3y agoHugging Face21qmeeus /vp-er-10l Dataset Card for "vp-er-10l" More Information needed audio100K<n<1M0 likes18k downloads3y agoHugging Face22EarthSpeciesProject /NatureLM-audio-training Dataset card for NatureLM-audio-training Overview NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.audioaudio-classification10M<n<100M18 likes18k downloads1y agoHugging Face23arsaporta /symile-m3 Dataset Card for Symile-M3 Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice. Paper: https://arxiv.org/abs/2411.01053 GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.audiozero-shot-classification10M<n<100M8 likes18k downloads2y agoHugging Face24simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M281 likes17k downloads1mo agoHugging Face25zxsddcs /wenetspeechaudio10M<n<100M1 likes16k downloads1y agoHugging Face26edinburghcstr /ami Dataset Card for AMI Dataset Description The AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals synchronized to a common timeline. These include close-talking and far-field microphones, individual and room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings, the participants also have unsynchronized pens available to them that record what is written. The meetings were… See the full description on the dataset page: https://huggingface.co/datasets/edinburghcstr/ami.audioautomatic-speech-recognition100K<n<1M98 likes15k downloads9mo agoHugging Face27japanese-asr /whisper_transcriptions.mls.wer_10.0audio1M<n<10M2 likes14k downloads2y agoHugging Face28PolyAI /minds14 MInDS-14 MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14 intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties. Example MInDS-14 can be downloaded and used as follows: from datasets import load_dataset minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French # to download all data for multi-lingual fine-tuning uncomment following… See the full description on the dataset page: https://huggingface.co/datasets/PolyAI/minds14.audioautomatic-speech-recognition10K<n<100K111 likes14k downloads1y agoHugging Face29huseyin-karaca /hit-asr HIT-ASR — data and results The data behind HIT-ASR: Hierarchical Transformer Routing for Adaptive ASR Expert Selection (Huseyin Karaca, A. Samil Namli, Suleyman S. Kozat — Bilkent University): what the pretrained ASR experts of the paper produce on its four English corpora, and the stored results every notebook of the code repository reads. Code and notebooks: github.com/huseyin-karaca/hit-asr Documentation: huseyin-karaca.github.io/hit-asr What is here… See the full description on the dataset page: https://huggingface.co/datasets/huseyin-karaca/hit-asr.audioautomatic-speech-recognition100K<n<1M0 likes14k downloads4d agoHugging Face30PedroDKE /LibriS2S LibriS2S This repo contains scripts and alignment data to create a dataset build further upon librivoxDeEn such that it contains (German audio, German transcription, English audio, English transcription) quadruplets and can be used for Speech-to-Speech translation research. Because of this, the alignments are released under the same Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License These alignments were collected by downloading the English audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/PedroDKE/LibriS2S.audiotext-to-speech10K<n<100K4 likes13k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.