datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voicehub-arena-seed-tts-eval
VoiceHub Arena — full English Seed-TTS-Eval
35,904 synthesized WAV files: 33 model families × the same 1,088 target texts.
The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB.
All 198 shards and every WAV SHA256 were verified after backup.
Interactive leaderboard and all audio samples
· Source repository (access required).
Contents
audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.LEMAS-Dataset-eval
Overview
This dataset is part of LEMAS-Project(lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-eval.nb-asr-eval-withwav-sorted
NB-ASR Eval with WAV Sorted
Hardest-first copy of NbAiLab/nb-asr-eval-withwav for targeted human cleanup.
Rows are intended to be sorted independently within each split by ASR/WER difficulty.
Audio paths and split metadata layout are preserved so downstream tools can switch
from the original repo to NbAiLab/nb-asr-eval-withwav-sorted without changing file lookup logic.
After scoring, each metadata row may include original_source_index, priority_rank,
asr_wer, asr_cer… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-eval-withwav-sorted.MLC-SLM-Eval
Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) Eval Groundtruth
🖥️ Overview
In the MLC-SLM challenge, we only provided the participants with the audio files of the Eval sets.
Now, we release the oracle segmentation, speaker labels, and transcriptions of the Eval sets to facilitate further research by all participants on the MLC-SLM dataset!
In addition, the MLC-SLM challenge summary paper "Summary on The Multilingual Conversational Speech… See the full description on the dataset page: https://huggingface.co/datasets/bsmu/MLC-SLM-Eval.agri-voice-eval
Agricultural Voice Evaluation Set — Hindi, Telugu, Odia
Human quality-checked FarmerChat field recordings with per-clip audio and human reference transcripts,
used to evaluate Digital Green's agricultural voice pipeline stage by stage. Companion to the larger
Agri STT Benchmarking Dataset,
focused on the multi-speaker / noisy conditions that motivate speaker selection and enhancement.
Each clip carries two references: the farmer's own (main-speaker) transcript — the… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/agri-voice-eval.eval-whatsapp
Dataset Card for ivrit.ai Whatsapp Eval
Evaluation dataset of Hebrew Whatsapp voice-messages, expert-transcribed.
Dataset Details
Dataset Description
This dataset containeד Whatsapp hebrew voice recordings gathered around April 2025.
The recordings are of volunteer native hebrew speakers using consumer devices in natural environments.
Each recording is by a single speaker about a random topic of their choice spoken in a non-scripted natural manner.
Each… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-whatsapp.bg-med-consultations-eval
Bulgarian Medical Consultations — evaluation set
156 synthesised Bulgarian doctor–patient consultations with exact ground
truth: 133.4 minutes, 2,544 turns, 11 native Bulgarian voices.
The reference is not an annotation. Every turn was placed on the timeline by the
generator, so the RTTM is a construction — correct by definition, with none of
the annotator disagreement that inflates published DER.
What is in it
directory
contents
audio/
16 kHz mono… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bg-med-consultations-eval.vistaar_small_asr_eval
Vistaar Small ASR Eval
Dataset Description
The Vistaar Small ASR Eval dataset is a multilingual automatic speech recognition evaluation dataset containing 9,486 audio samples across 12 Indian languages. This dataset represents a subset of the larger Vistaar dataset published by AI4Bharat, designed specifically for evaluating ASR model performance on diverse Indian language speech data. A smaller evaluation dataset was created for the use-cases where a quick benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/vistaar_small_asr_eval.open-ko-s2s-eval-artifacts
Open Ko-S2S 평가 산출물 (감사용)
⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다
KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의
ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과
채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다.
Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라
ref 원문을 그대로 담고 있습니다.
라이선스 보유자의 검증 절차
AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다.
리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다.
자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.bocalantics-wolof-eval
Bocalantics Wolof eval set
The held-out Wolof test rows from MOH749/Bocalantics-2.0, with audio attached, so a
candidate model can be scored without re-materialising anything.
3,557 rows, 5.38 hours.
Why it exists
"Beats unadapted Whisper" is not a result for Wolof. Whisper has seen 2 of the parent
corpus's 26 languages, so beating it is arithmetic rather than evidence. The bar is the
best published model for the language -… See the full description on the dataset page: https://huggingface.co/datasets/MOH749/bocalantics-wolof-eval.CoSHE-Eval
🎙️ CoSHE-Eval: A Code-Switching ASR Benchmark for Hindi–English Speech
🧠 Overview
CoSHE-Eval is an evaluation dataset curated for testing Automatic Speech Recognition (ASR) systems on Hindi-English code-mixed speech.It focuses on bilingual conversational contexts commonly found in India, where Hindi (in Devanagari) and English (in Latin script) co-occur naturally within the same utterance.
Detailed Blog: CoSHE-Eval Blog
Technical Specifications… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/CoSHE-Eval.eval-forced-alignment
Hebrew Forced Alignment Evaluation Dataset
Human-verified, word-level time-aligned Hebrew speech clips.
To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was
built. The system lets labelers fix the transcript and align each spoken word to the audio,
down to 1ms precision (though annotators typically work at ~10ms granularity).
The audio samples were gathered by randomly sampling from several of ivrit-ai's larger,
published open datasets. The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-forced-alignment.TTS-Evaluvation
Marathi Indic-Speak QLoRA Evaluation
This dataset stores reproducible evaluation outputs for
PrakashPask/marathi-indic-speak-qlora on the official Rasa Marathi test split.
Each uploaded run contains the selected text and conditioning metadata, the
exact Rasa reference WAV, adapter-generated WAV, Whisper transcription,
per-sample WER/CER, and measured full-waveform response time.
Latest packaged run: rasa_marathi_n015_seed3407 (15 rows)
Corpus WER: 66.9388%
Corpus CER: 20.7738%… See the full description on the dataset page: https://huggingface.co/datasets/PrakashPask/TTS-Evaluvation.Audio-Understanding-Bitrate-Eval-0426
Audio Understanding — MP3 Bitrate Evaluation (April 2026)
Empirical eval measuring how MP3 compression bitrate affects transcription accuracy across every audio-input LLM available on OpenRouter.
📝 Blog post: MP3 Bitrate Sensitivity in Audio-Multimodal LLMs
💻 Code & methodology: github.com/danielrosehill/Audio-Understanding-Bitrate-Eval-0426
TL;DR
Ran a benchmark across 12 OpenRouter audio-multimodal models × 4 dictation samples × 5 MP3 bitrates (16/24/32/48/64 kbps)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Bitrate-Eval-0426.Indic_ASR_Eval
Indic ASR Eval
A curated evaluation set for Indic-language automatic speech recognition.
100 samples are sampled (seed = 42) from each (source dataset × language)
cell of seven public Indic ASR corpora. Each source corpus is published
as its own dataset config with a single test split, at 16 kHz.
Rows: 6,169 across 7 configs
Total audio: ~13.3 hours
Sampling rate: 16 kHz (mono)
Split: test (single split in every config)
Configs
Config
Rows
Notes
kathbath… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/Indic_ASR_Eval.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.tiron-eval-meetings
Tiron evaluation meetings
The 17 held-out whole meetings used for the benchmarks on the
Trelis/tiron model card — far-field
single-channel audio (16 kHz mono WAV) with reference speaker-attributed
transcripts, packaged so results can be reproduced with the
Tiron harness.
split
meetings
source
ami
ES2004a, IS1009a, TS3003a, EN2002a
AMI Meeting Corpus, single distant microphone (Array1-01)
icsi
Bmr013, Bmr018, Bro021
ICSI Meeting Corpus, mean of 4 distant PZM room… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/tiron-eval-meetings.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.Small-STT-Eval-Audio-Dataset
Small STT Eval Audio Dataset
A small speech-to-text evaluation dataset containing 92 audio samples with ground truth transcriptions. Designed for evaluating STT systems on technical vocabulary, code-switching (English/Hebrew), and various speaking styles.
Dataset Description
This dataset contains audio recordings with accompanying transcriptions across multiple categories:
Category
Count
Description
tech_github
5
GitHub-related technical vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/hamzaiqbal590/Small-STT-Eval-Audio-Dataset.YouTube-Evaluation-Set
Awaaz se Alfaaz — YouTube Evaluation Set
This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.emova-asr-tts-eval
EMOVA-ASR-TTS-Eval
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-ASR-TTS-Eval is a dataset designed for evaluating the ASR and TTS performance of Omni-modal LLMs. It is derived from the test-clean set of the LibriSpeech dataset. This dataset is part of the EMOVA-Datasets collection. We extract the speech units using the EMOVA Speech Tokenizer.
Structure
This… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-asr-tts-eval.Indic-subtitler-audio_evals
Indic_audio_evals
As part of this project. We are evaluating our performance of various ASR models as well
in a benchmarking dataset, we have created in various languages. This benchmarking dataset
is more alligned to real-world use-cases rather than having any academic datasets.
About Dataset
Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals
This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.nvidia-brain-noise-evaluation-dataset
Nvidia Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 64 samples organized across multiple splits and 32 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test: 2 samples
noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.Small-STT-Eval-Audio-Dataset
Small STT Eval Audio Dataset
A small speech-to-text evaluation dataset containing 92 audio samples with ground truth transcriptions. Designed for evaluating STT systems on technical vocabulary, code-switching (English/Hebrew), and various speaking styles.
Dataset Description
This dataset contains audio recordings with accompanying transcriptions across multiple categories:
Category
Count
Description
tech_github
5
GitHub-related technical vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Small-STT-Eval-Audio-Dataset.lithuanian-dialect-asr-eval
Lithuanian dialect ASR evaluation bundle
Everything needed to check or extend the evaluation of the study Lithuanian speech recognition: effects of dialect training and transcript spelling: the test sets, every
reference layer, the hypotheses of every evaluated model on every test set, and all scores. Code
and scorer: github.com/Digisensus/lithuanian-dialect-asr.
Test sets
Config
Test set
Clips
Hours
Speakers
References
dial_test
LIEPA-3 dialect test:… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-dialect-asr-eval.ksponspeech-evalpaper link: https://www.mdpi.com/846876
kiswahili-asr-challenge-eval
Swahili ASR Zindi Evaluation Dataset
Presented by Sartify in partnership with AI for Good (ITU) on the Zindi platform.
Dataset Summary
This is the official evaluation (test) dataset for the "Your Voice, Your Device, Your Language" challenge — a Zindi competition focused on building lightweight, on-device Kiswahili Automatic Speech Recognition (ASR) systems.
The dataset contains Kiswahili audio samples intended for evaluating speech-to-text models under real-world… See the full description on the dataset page: https://huggingface.co/datasets/sartifyllc/kiswahili-asr-challenge-eval.atc-asr-eval
ATC ASR Evaluation Set
Human-verified air-traffic-control transmissions for evaluating ASR models.
Each clip is 16 kHz mono WAV with a corrected ground-truth transcript in
metadata.csv (columns: file_name, transcription, icao, city, region, source, feed).
Audio captured from LiveATC.net feeds. Private — not for redistribution
(LiveATC terms prohibit rebroadcasting).
Albayzin-2024-BBS-S2T-eval
Albayzin 2024 Bilingual Basque-Spanish Speech to Text (BBS-S2T) Challenge - Evaluation dataset
see Albayzin_2024_BBS-S2T_EvalPlan for a description of the challenge.
This is the evaluation data for the challenge.
The database consists of a single split:
eval : 12498 audio segments
How to download this database
1 - If you can handle yourself comfortably with Huggingface Datasets:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/Albayzin-2024-BBS-S2T-eval.
