Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01espnet /yodas3 YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data. For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.audioaudio-to-audio1M<n<10M206 likes128k downloads5d agoHugging Face02huseyin-karaca /hit-asr HIT-ASR — data and results The data behind HIT-ASR: Hierarchical Transformer Routing for Adaptive ASR Expert Selection (Huseyin Karaca, A. Samil Namli, Suleyman S. Kozat — Bilkent University): what the pretrained ASR experts of the paper produce on its four English corpora, and the stored results every notebook of the code repository reads. Code and notebooks: github.com/huseyin-karaca/hit-asr Documentation: huseyin-karaca.github.io/hit-asr What is here… See the full description on the dataset page: https://huggingface.co/datasets/huseyin-karaca/hit-asr.audioautomatic-speech-recognition100K<n<1M0 likes14k downloads5d agoHugging Face03ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes12k downloads10h agoHugging Face04SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes6.8k downloads1y agoHugging Face05QUD-Technologies /quranic-universal-ayahs Qur'anic Universal Ayahs Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset. This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.audioautomatic-speech-recognition100K<n<1M7 likes5.6k downloads2d agoHugging Face06twangodev /librivox-mirror LibriVox Mirror Fast, structured, continuously updated LibriVox audio mirror. Current snapshot Metric Value Published books 21,766 Published sections 494,200 Audio hours 132,876.0 Audio languages 86 Quarantined books 577 Last updated (UTC) 2026-10-05T17:46:56.602294Z Audio by language Language Hours English 131,926.6 German 417.0 Spanish 160.9 French 103.8 Portuguese 37.4 Polish 34.1 Dutch 25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.audioautomatic-speech-recognition100K<n<1M0 likes4.4k downloads10h agoHugging Face07MohammadJRanjbar /ParsVoicegated ParsVoice A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis 📣 Accepted to the EMNLP 2026 Main Conference. ParsVoice is the largest publicly available Persian speech–text corpus tailored for training multi-speaker text-to-speech (TTS) systems. It is built from long-form Persian audiobook recordings using a fully automated pipeline combining sentence-aware segmentation, ASR transcription, a ParsBERT sentence-completion classifier, binary-search… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice.audiotext-to-speech1M<n<10M35 likes3.9k downloads1mo agoHugging Face08FaisaI /tadabur Tadabur: A Large-Scale Quran Audio Dataset The most comprehensive and richly annotated Qur'anic recitation corpus to date Faisal Alherran &nbsp; &nbsp; &nbsp; ✦ Overview Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.audioaudio-classification100K<n<1M23 likes3.1k downloads3mo agoHugging Face09RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M5 likes2.8k downloads6h agoHugging Face10retkowski /ytseg YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.audiotoken-classification100K<n<1M10 likes2.5k downloads3mo agoHugging Face11facebook /2M-Belebele 2M-Belebele Highly-Multilingual Speech and American Sign Language Comprehension Dataset We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL). The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.tabularquestion-answering10K<n<100K13 likes2.3k downloads2y agoHugging Face12syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes2.1k downloads20d agoHugging Face13Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes1.7k downloads3mo agoHugging Face14humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.2k downloads8mo agoHugging Face15RodolfoRizzi /MyMentorLLM-dataset Dataset Card for MyMentorLLM This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information). Dataset Summary MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.audiotext-generation10K<n<100K0 likes1.1k downloads15d agoHugging Face16anuj-inavlabs /kupe-asr-en-data kupe-asr-en-mini-150m — data Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly): raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this. mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this. Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state. from datasets import load_dataset ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train") tabularautomatic-speech-recognition1M<n<10M0 likes1k downloads26d agoHugging Face17Scicom-intl /Whisper-Hallucination Whisper Hallucination and Repetition Probes This is a benchmark. Every evaluation config is test — do not fine-tune on it. lexicon_synth is the exception: synthetic training material with its own train/test split, and not one of the eight benchmark arms. To build training data, exclude the items in benchmark/exclusions.json (546 FMA tracks, 1,168 FSD50K ids, 2,620 LibriSpeech utterances, the Malay stems). The benchmark draws FSD50K eval and FMA shards 0–1, so training can use… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.audioautomatic-speech-recognition100K<n<1M0 likes980 downloads11d agoHugging Face18risaleinur /risale-i-nur-sohbet Risale-i Nur Sohbet Prof. Dr. Şener Dilek’ten izin alındı. Türkçe Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri kümelerine karıştırılmaz. Kapsam 2095 sohbet, 954.66 saat 16 kHz mono FLAC ses Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.audioautomatic-speech-recognition1M<n<10M1 likes962 downloads1mo agoHugging Face19Borrison /hk-legicost HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation HK-LegiCoST is a three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and aligned at the sentence level. Paper: arXiv:2306.11252 Authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur Dataset Description The raw… See the full description on the dataset page: https://huggingface.co/datasets/Borrison/hk-legicost.audioautomatic-speech-recognition100K<n<1M3 likes876 downloads4mo agoHugging Face20raianand /TIE_shorts Dataset Card for TIE_Shorts Dataset Summary TIE_shorts is a derived version of the Technical Indian English (TIE) dataset, a large-scale speech dataset (~ 8K hours) originally consisting of approximately 750 GB of content sourced from the NPTEL platform. The original TIE dataset contains around 9.8K technical lectures in English delivered by instructors from various regions across India, with each lecture averaging about 50 minutes. These lectures cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/raianand/TIE_shorts.audioautomatic-speech-recognition1K<n<10K1 likes865 downloads2y agoHugging Face21Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes799 downloads4mo agoHugging Face22humair025 /hashed_data Munch Hashed Index - Lightweight Audio Reference Dataset 📖 Overview Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling: ✅ Fast duplicate detection across 4.17 million audio samples ✅ Efficient dataset exploration without downloading terabytes ✅ Quick metadata queries (voice… See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.tabulartext-generation1M<n<10M0 likes734 downloads10mo agoHugging Face23Baekpica /Inkling-Small-Multimodal-Calibration Inkling-Small Multimodal Calibration The exact 1,663 samples used for BF16 routed-expert importance collection for Inkling-Small Mixed Quant GGUF. This is calibration material, not a held-out evaluation benchmark. The primary balanced pass is: Category Samples Valid decoder tokens Share Text / reasoning 462 471,858 44.976% Code / tool-oriented source text 205 209,715 19.989% Real image / document 486 262,476 25.018% Real speech audio 309 105,080 10.016% Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.tabulartext-generation1K<n<10K1 likes686 downloads27d agoHugging Face24parler-tts /mls-eng-speaker-descriptions Dataset Card for Annotations of English MLS This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.tabularautomatic-speech-recognition10M<n<100M13 likes620 downloads2y agoHugging Face25Sanghyang00 /omniasr-molge OmniASR Molge Aligned Training-friendly re-segmentation of Meta’s facebook/omnilingual-asr-corpus: long utterances are segmented / aligned into ≤30s clips with transcripts, then packed as Parquet shards with embedded FLAC. Source facebook/omnilingual-asr-corpus Configs omniasr_aligned_v1, omniasr_aligned_v2 Splits train / validation (dev-*.parquet) / test Scale ~2.56M utts · ~839 shards · ~439GB If this dataset is useful for your work, we’d appreciate a… See the full description on the dataset page: https://huggingface.co/datasets/Sanghyang00/omniasr-molge.tabularautomatic-speech-recognition1M<n<10M0 likes528 downloads2mo agoHugging Face26deepghs /arknights_voices_zh ZH Voice-Text Dataset for Arknights Waifus This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models. Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset. 12431 records, 25.9 hours in total. Average duration is approximately 7.49s. id char_id voice_actor_name voice_title voice_text time sample_rate file_size filename mimetype file_url char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.tabularautomatic-speech-recognition10K<n<100K6 likes503 downloads2y agoHugging Face27ai-music4you3 /enhanced-audiosnippets-long-2-8M Enhanced Audiosnippets Long 2.8M Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis. Dataset Summary Metric Value Total samples 2,633,037 Total audio hours 4,932 h Duration range 3.0s - 1124.3s Mean duration 6.7s Audio format WAV, 48kHz mono Tar files 1,410 Processing Pipeline Each audio sample was processed through: Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.tabularaudio-classification1M<n<10M1 likes435 downloads7mo agoHugging Face28united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes415 downloads7mo agoHugging Face29mteb /svq Simple Voice Questions Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions. Data Collection Speakers were presented with recording instructions specifying the recording environment and text query to be recorded. They recorded using their own phones or tablets under four conditions: clean: Record in quiet environment background speech noise: Record while audio from sources like podcasts… See the full description on the dataset page: https://huggingface.co/datasets/mteb/svq.audioquestion-answering100K<n<1M0 likes407 downloads8mo agoHugging Face30risaleinur /risale-nur-audio Risale-i Nur Audio–Text Corpus Gerçek insan okumalarını, aynı satırdaki kaynak metinle birlikte sunan açık bir ses–metin veri kümesidir. Yeni varsayılan audio-text yapılandırması 15 kitaptan 91.792 oynatılabilir klip ve 203,02 saat ses içerir. Metinler kanonik kaynaktan değiştirilmeden alınır ve her kayıt byte-exact section_id alıntılarıyla bağlanır. An open speech corpus pairing human readings with their source text in the same row. The default audio-text config contains 91,792… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-audio.audioautomatic-speech-recognition100K<n<1M1 likes382 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.