Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes8.8k downloads8mo agoHugging Face02LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B90 likes6.5k downloads6mo agoHugging Face03nvidia /Granary Granary: Speech Recognition and Translation Dataset in 25 European Languages Granary is a large-scale, open-source multilingual speech dataset covering 25 European languages for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks. Overview Granary addresses the scarcity of high-quality speech data for low-resource languages by consolidating multiple datasets under a unified framework: 🗣️ ~1M hours of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Granary.tabularautomatic-speech-recognition100M<n<1B225 likes5.7k downloads4mo agoHugging Face04shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes4.4k downloads9mo agoHugging Face05apptek-com /apptek_callcenter_dialogues AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions. 128.6 hours of speech 14 English accent groups 16 service domains 5–15 minute conversations (long-form) Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.audioautomatic-speech-recognition1K<n<10K40 likes3.3k downloads2mo agoHugging Face06RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K3 likes3.3k downloads4mo agoHugging Face07zhifeixie /StreamAudio-2M StreamAudio-2M Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips are organised into six task subsets. Subsets Subset Rows Description Stream_Audio_Understanding 90,738 Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA Real_time_ASR 28,109 Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.tabularaudio-classification100K<n<1M30 likes2.8k downloads4mo agoHugging Face08besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K14 likes1.7k downloads1d agoHugging Face09MrSupW /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K38 likes1.4k downloads1y agoHugging Face10tutu0604 /UltraVoice UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset… See the full description on the dataset page: https://huggingface.co/datasets/tutu0604/UltraVoice.audiotext-to-speech100K<n<1M16 likes1.3k downloads11mo agoHugging Face11ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes1.3k downloads3mo agoHugging Face12OPPOer /HearInContextEnglish | 中文 HearInContext A Benchmark for Implicit Context in Speech Recognition Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative. Same audio. Different contexts. Different meanings. HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.audioautomatic-speech-recognition100K<n<1M2 likes1.3k downloads20d agoHugging Face13ArtificialAnalysis /VoxPopuli-Cleaned-AA VoxPopuli-Cleaned-AA Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models. This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes1.1k downloads8mo agoHugging Face14rushilrawat /garhwali-corpus Garhwali Language Lab Current developer status — 8 October 2026 The live Dataset Viewer reports 20 configurations, 31 config/split views, and 963,484 displayed rows. These views overlap and include source indexes; they are not 963,484 distinct training examples. The screened_meta_gbm view has 1,841 automatically screened sentence-length transcripts for exploratory text training, and short_utterances_meta_gbm has 110 context rows. Both reuse transcript values… See the full description on the dataset page: https://huggingface.co/datasets/rushilrawat/garhwali-corpus.tabularautomatic-speech-recognition100K<n<1M0 likes987 downloads2d agoHugging Face15Sheeba2026 /bharatvani-hindi-speech-corpusgated BharatVani Hindi Speech Corpus (205-Hour Studio Dataset) Proprietary Speech Asset • TheCreatorOS • BharatVani AI 1. Overview The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered for speech foundation models, acoustic research, and voice synthesis in Devanagari Hindi & Conversational Hinglish. Audio Clips: 99,475 Pristine, 1:1 Verified Audio Clips (24,000 Hz, 16-bit Mono PCM WAV) Duration: ~205.3… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.audiotext-to-speech10K<n<100K1 likes849 downloads10d agoHugging Face16Quran-Lab /quran-tajweed-phonetics The complete phonetic layer of the Quran in the riwaya of Hafs 'an 'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every phone carrying its tajweed attribution: madd class with its transmitted length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt, the seventeen sifat, and the rule that produced it. Built and maintained by Quran Lab, a waqf building open technology in the service of the Quran. How it was built and verified Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.tabularautomatic-speech-recognition10K<n<100K4 likes732 downloads9d agoHugging Face17FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes674 downloads6mo agoHugging Face18abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes518 downloads1mo agoHugging Face19wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes493 downloads2mo agoHugging Face20medkit /simsamu Simsamu dataset This repository contains recordings of simulated medical dispatch dialogs in the french language, annotated for diarization and transcription. It is published under the MIT license. These dialogs were recorded as part of the training of emergency medicine interns, which consisted in simulating a medical dispatch call where the interns took turns playing the caller and the regulating doctor. Each situation was decided randomly in advance, blind to who was playing the… See the full description on the dataset page: https://huggingface.co/datasets/medkit/simsamu.audioautomatic-speech-recognitionn<1K8 likes419 downloads11mo agoHugging Face21IMJONEZZ /star-wars-dataset Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21) One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of 15 titles with a forced-aligned time span, a character label and provenance. No audio or video is included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips from your own copies with the pipeline code (export_asr.py, export_diarization.py). The text layers contain… See the full description on the dataset page: https://huggingface.co/datasets/IMJONEZZ/star-wars-dataset.tabularautomatic-speech-recognition10K<n<100K3 likes345 downloads19d agoHugging Face22lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes306 downloads24d agoHugging Face23vnmoorthy /pavo-bench PAVO-Bench: 50K-Turn Benchmark for ASR-LLM-TTS Pipeline Routing Code: github.com/vnmoorthy/pavo-bench · Paper: TMLR 2026 (accepted) · Authors: NarasingaMoorthy VeiluKanthaPerumal (UPenn), Mohammed Imthathullah (Google) pip install git+https://github.com/vnmoorthy/pavo-bench.git Headline results (vs fixed-cloud baseline, 50,000 voice turns) Metric Result Significance P95 end-to-end latency (H100, LibriSpeech) −10.3% (−167 ms) — Median latency −34%… See the full description on the dataset page: https://huggingface.co/datasets/vnmoorthy/pavo-bench.documentautomatic-speech-recognition10K<n<100K0 likes287 downloads2mo agoHugging Face24panlr /teochew_wildgated Teochew-Wild:首个正字标注的野外潮州话数据集 中文 | English 本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正); Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。 文件说明 (File Structure Explanation) ├── annotation/ # 标注相关文件夹 │ ├── label_for_qwen_asr/ #… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.audiotext-to-speech10K<n<100K50 likes285 downloads2d agoHugging Face25ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes285 downloads6mo agoHugging Face26skylar-ai-hf /skylar-dataset Skylar Dataset A curated speech dataset built for automatic speech recognition (ASR) benchmarking. It is assembled by streaming samples from existing public audio datasets and keeping only the ones that pass a fixed set of quality and diversity rules, organized by target language and audio duration. This is not a single-source dataset: samples are pulled from multiple upstream datasets into one unified structure. Re-running the pipeline with a different source (same or different… See the full description on the dataset page: https://huggingface.co/datasets/skylar-ai-hf/skylar-dataset.audioautomatic-speech-recognition1K<n<10K0 likes244 downloads10d agoHugging Face27yunqi1766 /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.audioautomatic-speech-recognitionn<1K1 likes221 downloads3mo agoHugging Face28NightPrince /quran-asr-husary Quran ASR — Husary Muallim Dataset Description This dataset contains Quran recitation audio files by Sheikh Mahmoud Khalil Al-Husary at 16 kHz sampling rate, with Arabic transcriptions including diacritics. Dataset Structure Audio files: Stored in audio/ folder (e.g., audio/001_001.wav) Data file: manifest.json (NeMo format) Columns: audio_filepath: Path to audio file text: Arabic transcription with diacritics duration: Audio duration in seconds speaker:… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/quran-asr-husary.audioautomatic-speech-recognition1K<n<10K0 likes214 downloads7mo agoHugging Face29jm-vis /modelroom-catalog Open Model Catalog A small table of the models their publishers call current: text, vision, embedding, speech recognition, image and video generation. Not every build on the Hub. One row per model, with the publisher page that says it is current and the date that was checked. 247 models · 56 publisher accounts watched · one JSON file (183 KB) · schema 2 · CC BY 4.0 What a row carries Field Meaning hf_repo, publisher, family, kind the repository, its… See the full description on the dataset page: https://huggingface.co/datasets/jm-vis/modelroom-catalog.texttext-generationn<1K0 likes214 downloads2d agoHugging Face30lysanderism /FTAR TimeAudio: Bridging Temporal Gaps in Large Audio-Language Models Abstract Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perception, leading to constrained capabilities on fine-grained tasks. We identify three key aspects that limit… See the full description on the dataset page: https://huggingface.co/datasets/lysanderism/FTAR.audioaudio-classification100K<n<1M3 likes210 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.