Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01espnet /yodas3 YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data. For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.audioaudio-to-audio1M<n<10M207 likes128k downloads5d agoHugging Face02google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M473 likes113k downloads5mo agoHugging Face03espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M34 likes73k downloads1y agoHugging Face04google /svq Simple Voice Questions Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions. It serves as a core evaluation componenet for Massive Sound Embedding Benchmark (MSEB). Technical Specifications Feature Details Locales 26 Languages 17 Total Speakers ~700 (Capped at 250 recordings per speaker) Audio Conditions Clean, Background Speech, Media, Traffic Noise Gender… See the full description on the dataset page: https://huggingface.co/datasets/google/svq.audioquestion-answering1M<n<10M62 likes57k downloads12d agoHugging Face05openslr /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.audioautomatic-speech-recognition100K<n<1M245 likes54k downloads1y agoHugging Face06facebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M191 likes44k downloads2y agoHugging Face07fsicoli /common_voice_15_0 Dataset Card for Common Voice Corpus 15.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 15. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_15_0.automatic-speech-recognition100B<n<1T6 likes39k downloads3y agoHugging Face08MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M286 likes35k downloads2y agoHugging Face09disco-eth /EuroSpeech EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.audioautomatic-speech-recognition10M<n<100M99 likes34k downloads5mo agoHugging Face10amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M492 likes33k downloads2y agoHugging Face11aoxo /t2a-mommy t2a-mommy Female-voice ASMR corpus for the text2asmr project. Previously published as aoxo/audios2. Companion repos: aoxo/t2a-daddy (male voice), aoxo/t2a-audios-v1 (the original v1 corpus). Layout path what <creator>/<title>.m4a source audio, 48 kHz AAC, one folder per creator <creator>/<title>.json word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable) labels/qwen3omni.jsonl non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-mommy.audioaudio-classification0 likes30k downloads8d agoHugging Face12fsicoli /common_voice_22_0 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_22_0.automatic-speech-recognition100B<n<1T20 likes28k downloads1y agoHugging Face13facebook /voxpopuli Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.audioautomatic-speech-recognition1M<n<10M166 likes26k downloads8mo agoHugging Face14ARTPARK-IISc /VaanigatedVAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.audioautomatic-speech-recognition1M<n<10M160 likes24k downloads19d agoHugging Face15disco-eth /WorldSpeech WorldSpeech 🎉 WorldSpeech has been accepted to NeurIPS 2026! 🎉See the paper on arXiv. A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/WorldSpeech.audioautomatic-speech-recognition10M<n<100M68 likes19k downloads11d agoHugging Face16kapturecx /bolAIndiagated bolAIndia Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. Alongside it, open Indian-language speech corpora converted to the same schema (16 kHz mono FLAC, one utterance per row), each in a config of its own and tagged with where it came from.… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.audioautomatic-speech-recognition10M<n<100M2 likes17k downloads23h agoHugging Face17simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M281 likes17k downloads1mo agoHugging Face18MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes17k downloads2y agoHugging Face19davidscripka /MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html. The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation. The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.audioaudio-classificationn<1K9 likes17k downloads3y agoHugging Face20MC7ever /hf-training-corpus HF Training Corpus Bulk-scraped multimodal training corpus: ~92,000 Hugging Face datasets streamed, normalized and stored as per-dataset Parquet files under data/. Every modality is captured: text, images (PNG bytes), audio (mono 16-bit WAV bytes), video (bytes, capped) and tabular/timeseries (serialized to text). Pipeline Live crawl from a 92,458-ID corpus list (see https://github.com/shadyuwugurl/hf-datasets-trees/blob/main/datasets.txt) streaming=True reads… See the full description on the dataset page: https://huggingface.co/datasets/MC7ever/hf-training-corpus.audiotext-generation1M<n<10M0 likes15k downloads8m agoHugging Face21fsicoli /common_voice_17_0 Dataset Card for Common Voice Corpus 17.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_17_0.automatic-speech-recognition100B<n<1T19 likes15k downloads2y agoHugging Face22edinburghcstr /ami Dataset Card for AMI Dataset Description The AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals synchronized to a common timeline. These include close-talking and far-field microphones, individual and room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings, the participants also have unsynchronized pens available to them that record what is written. The meetings were… See the full description on the dataset page: https://huggingface.co/datasets/edinburghcstr/ami.audioautomatic-speech-recognition100K<n<1M98 likes15k downloads9mo agoHugging Face23PolyAI /minds14 MInDS-14 MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14 intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties. Example MInDS-14 can be downloaded and used as follows: from datasets import load_dataset minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French # to download all data for multi-lingual fine-tuning uncomment following… See the full description on the dataset page: https://huggingface.co/datasets/PolyAI/minds14.audioautomatic-speech-recognition10K<n<100K111 likes14k downloads1y agoHugging Face24zaibihassan /Quranic-Recitation-Data 🌟 Overview Quranic Recitation Dataset (Word-by-Word Sync) is a highly optimized, production-ready dataset containing high-quality audio recitations of the Holy Quran synchronized at the word-by-word level. This dataset features 135 world-renowned reciters, with every Surah (114 chapters) mapped precisely to millisecond-accurate word timestamps. It is designed for modern Islamic mobile and web applications — served via a Cloudflare Edge CDN with native… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data.audioautomatic-speech-recognition10K<n<100K6 likes14k downloads1h agoHugging Face25fsicoli /common_voice_16_0 Dataset Card for Common Voice Corpus 16.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 16. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_16_0.automatic-speech-recognition100B<n<1T4 likes14k downloads3y agoHugging Face26huseyin-karaca /hit-asr HIT-ASR — data and results The data behind HIT-ASR: Hierarchical Transformer Routing for Adaptive ASR Expert Selection (Huseyin Karaca, A. Samil Namli, Suleyman S. Kozat — Bilkent University): what the pretrained ASR experts of the paper produce on its four English corpora, and the stored results every notebook of the code repository reads. Code and notebooks: github.com/huseyin-karaca/hit-asr Documentation: huseyin-karaca.github.io/hit-asr What is here… See the full description on the dataset page: https://huggingface.co/datasets/huseyin-karaca/hit-asr.audioautomatic-speech-recognition100K<n<1M0 likes14k downloads5d agoHugging Face27PedroDKE /LibriS2S LibriS2S This repo contains scripts and alignment data to create a dataset build further upon librivoxDeEn such that it contains (German audio, German transcription, English audio, English transcription) quadruplets and can be used for Speech-to-Speech translation research. Because of this, the alignments are released under the same Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License These alignments were collected by downloading the English audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/PedroDKE/LibriS2S.audiotext-to-speech10K<n<100K4 likes13k downloads1y agoHugging Face28ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes12k downloads14h agoHugging Face29Sinoosoida /SpeechRu Russian Podcasts (unlabeled) ~186k unlabeled Russian-language podcast episodes scraped from the web, packaged as Parquet shards with the audio bytes embedded. The audio has no transcripts — this is an unsupervised / self-supervised audio corpus, suitable for ASR pre-training, speech-representation learning, TTS data mining, audio classification, and similar tasks. Each row contains: audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo), decoded on-the-fly via the… See the full description on the dataset page: https://huggingface.co/datasets/Sinoosoida/SpeechRu.audioautomatic-speech-recognition100K<n<1M4 likes11k downloads3mo agoHugging Face30speechcolab /gigaspeechgated Dataset Card for Gigaspeech Dataset Description GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. Example Usage The training split has several configurations of… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech.audioautomatic-speech-recognition10M<n<100M174 likes11k downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.