datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VoiceIsolation-Benchmark-Dataset
Voice Isolation Benchmark Dataset
265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing.
265 recordings · 47 speakers · 3 scenarios · 65 scripts
Why this dataset exists
Modern STT engines handle noise well. They still fail when a second person talks near the microphone: they transcribe the wrong speaker, and voice… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset.quranic-asr-benchmark
Quranic ASR Benchmark - leakage-free, held-out
A small, leakage-free benchmark (600 clips) for evaluating Arabic ASR on Quranic recitation
(Hafs riwayah). Every clip is verified absent from our training data, so it measures
generalization, not memorization. Same clips + same scoring for every model.
📊 Live leaderboard: https://huggingface.co/spaces/Muno459/quranic-asr-leaderboard
The set (600 clips, 200 per source)
Source
n
What it is
everyayah_heldout… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-benchmark.amharic-asr-benchmark
Amharic ASR Benchmark
An evaluation of open speech recognition models for Amharic, on a test set with
certain labels and honest statistics.
16 models. 1,548 clips. 4.72 hours. Every hypothesis published.
Published by Dataset.ET.
Read this table first
Round 1 of this benchmark rested on a single clean claim: every model predated
our dataset, so none could have trained on it. That claim no longer holds.
Models trained on snapwre/amharic-speech now exist, and others… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-asr-benchmark.persian-accents-benchmark
Persian Accents Benchmark
Dataset Summary
A benchmark for Persian automatic speech recognition (ASR): 279 short
utterances of informal Persian (Farsi) dialect speech across 16 regional accents,
released as a fixed evaluation set. Total audio duration is approximately 4.4
hours. The primary label is the transcription; each utterance also carries an
accent label (usable for accent classification as a secondary task) and an
emotion label as auxiliary metadata.
This… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-accents-benchmark.quran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.entity-transcription-benchmark
Entity Transcription Benchmark
Measures whether a speech recognition system transcribes named entities
correctly — as distinct from word error rate.
WER weights every token equally. The tokens that matter for redaction, lookup,
routing and search are proper nouns, and they are a small fraction of any
transcript. A system can improve WER while getting worse at exactly the words a
downstream consumer needs, and nothing in the standard evaluation will show it.
2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for six language variants, built
from Common Voice 17.0 by coverage-driven selection rather than random sampling.
Every example pairs a reference clip of one speaker with a target text that
speaker never read, so a system is asked to clone a voice and produce new
speech, which is what zero-shot TTS is actually for.
Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.Agri_STT_Benchmarking_Dataset
Agri STT Benchmarking Dataset
10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository.
Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.commonvoice_benchmark_catalan_accentsA new presentation of the corpus Catalan Common Voice v17.0 - metadata annotated version with the splits redefined to benchmark ASR models with various Catalan accenta5sv2-asr-benchmark-dataset
A5Sv2 ASR Benchmark Dataset
Public references, saved predictions, scores, and provenance for the
A5Sv2 ASR benchmark. The benchmark evaluates
streaming English ASR on four fixed public corpora with approximately equal normalized reference
word counts.
Corpus
Fixed selection
Reference words
Audio in this repository
Mega-ASR / Voices-in-the-Wild-2M
1,250 utterances, 250 per acoustic condition
32,928
Yes
AMI
7 scenario-only unseen-evaluation meetings
32,928
Yes
DiPCo… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/a5sv2-asr-benchmark-dataset.Vaani-Benchmark-V1.0
Vaani-Benchmark-V1.0
A curated Hindi ASR evaluation set collected as part of the Vaani project at IISc Bengaluru. This is a separate, held-out collection — distinct from the publicly released Vaani dataset — built specifically for benchmarking. This benchmark contains 5,050 audio segments from 1,103 speakers across 104 Indian districts, each with three independent human transcriptions.
Dataset Summary
Property
Value
Language
Hindi (with code-switching)… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0.script-fidelity-benchmark
Script fidelity benchmark
Anonymous supplement for the paper "Script collapse in multilingual ASR:
A reference-free metric and 100-pair benchmark."
Script Fidelity Rate (SFR) measures the fraction of ASR hypothesis characters
that belong to the expected target script. WER measures word edits, while SFR
checks whether the output is written in the target orthography.
Related resources:
PyPI package: https://pypi.org/project/script-fidelity/
Hugging Face Evaluate metric:… See the full description on the dataset page: https://huggingface.co/datasets/themechanism/script-fidelity-benchmark.uzbek-asr-benchmark-spontaneous
Uzbek Spontaneous Speech ASR Benchmark
A 311-clip, 1.98-hour test set for Uzbek speech recognition, built from
spontaneous YouTube speech: podcasts, interviews and multi-speaker
conversation with overlapping turns, fillers, code-switching into Russian, and
dialect spelling. Every transcript that an automatic difficulty check flagged
as possibly wrong was corrected by hand — 118 of the 311 — and the protocol
below says exactly which ones and why.
It exists because the public… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-benchmark-spontaneous.Persian-ASR-BenchmarkThis dataset consists of 3 hours of 16kHz audio collected from diverse environments to better represent real-world scenarios. The recordings were sourced from audiobooks, YouTube, and other public sources, ensuring a wide variety of speech styles and acoustic conditions.
One key advantage of this dataset is that it was collected from recent sources within the last few months, ensuring no overlap with training data and fairness for evaluating other STT models.
To enable a robust and fair… See the full description on the dataset page: https://huggingface.co/datasets/C1Tech/Persian-ASR-Benchmark.turkish-asr-benchmark
TurkMedSTT Turkish ASR Benchmark
This repository publishes general-domain Turkish ASR benchmark metrics produced by
Muhammed Kumcu and Yagmur Tuncer.
Interactive table:
turkmedstt/turkish-asr-leaderboard
The release compares 22 ASR models on general-domain Turkish speech clips drawn from
three established Turkish speech sources. It contains metrics only. Source audio,
reference transcripts, generated hypotheses, local paths, and medical evaluation
results are not distributed.… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/turkish-asr-benchmark.vocal-money-codeswitch-asr-benchmark
Vocal Money — Yoruba–English Code-Switched ASR Benchmark
A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally
code-switched Yoruba–English speech, together with the reference transcriptions and the output of
every system on every clip, so that the published results can be recomputed or contradicted.
Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026.
Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.gradrai-viva-codeswitch-benchmark
GradrAI Viva Code-Switched Oral Benchmark
Consented, de-identified classroom-style oral answer clips used to benchmark GradrAI Viva for the Sahara CodeSwitch Africa challenge.
Contents
metadata.csv / metadata.jsonl: one row per clip.
audio/: 16 kHz mono WAV files for Hugging Face dataset preview and ASR reuse.
audio_original/: original submitted browser/Opus/WebM audio files.
benchmark/: benchmark outputs (results.md, results.json) and manifest used by GradrAI… See the full description on the dataset page: https://huggingface.co/datasets/fiewor/gradrai-viva-codeswitch-benchmark.humans-benchmark
HUMANS Benchmark Dataset
Authors: Woody Haosheng Gan¹, William Held²'³, Diyi Yang²
¹University of Southern California, ²Stanford University, ³OpenAthena
This dataset is part of the Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment paper.
HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark is designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/humans-benchmark.intron-stt_tts-benchmark
Code-switched benchmark audio
The exact 16 recordings used to benchmark speech models for the Sahara
CodeSwitch Africa Challenge. Published so the reported numbers can be checked
against the audio that produced them.
Every clip is intra-sentential code-switching — one speaker moving between a
Nigerian language and English inside a single utterance, which is the case the
challenge is judged on.
Language
Clips
Hausa
4
Igbo
4
Nigerian Pidgin
4
Yoruba
4… See the full description on the dataset page: https://huggingface.co/datasets/yusasif/intron-stt_tts-benchmark.asr-benchmark-outputs
SaarAI ASR Benchmark Outputs
Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets.
files: 570
utterances: 4615160
languages: 7
models: 50
Layout
data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl
index.jsonl # one record per file (language, split, model, rows, sha256, ...)
index.csv
Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.benchmark_eseu_testsets
Benchmark Test-sets for evaluations on Spanish and Basque
This test-sets are a reduced version of public available datasets. The datasets are balanced with more or less the same amount of hours in each dataset, for equal evaluation tasks.
Test splits:
mozilla-foundation/common_voice_18_0/es: a small split made from the official "test" split for spanish.
mozilla-foundation/common_voice_18_0/eu: a small split made from the official "test" split for basque.
openslr/es: a… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/benchmark_eseu_testsets.weha-health-benchmark-audio
Weha Health Benchmark Audio
Consented, de-identified voice recordings collected for benchmarking
speech-to-text engines as part of the Sahara CodeSwitch Africa Challenge
(Intron Health), Health category submission "Weha Health."
Contents
60 short simulated health-triage utterances across 4 languages, recorded
by the Weha Health team, each naturally code-switched with English
(except the Yoruba subset, which is monolingual Yoruba):
Yoruba (yo): 20 clips
Nigerian… See the full description on the dataset page: https://huggingface.co/datasets/techwithnel/weha-health-benchmark-audio.darija-asr-benchmark-6speaker
Darija ASR 6-Speaker Benchmark
A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3
female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus),
used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi)
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
Consent and anonymization
Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.Fasih-TTS-Benchmark
Fasih-TTS-Benchmark
Evaluation audio and objective scores for the
Fasih-TTS-V1 Arabic (MSA / Fusha)
text-to-speech model. Every clip is the model's own output, paired with its reference text, its
Whisper-large-v3 transcription, and per-clip WER / CER.
Contents
Split (test_set)
Clips
Purpose
Mean CER
silma_msa
10
SILMA open-source Arabic TTS benchmark (MSA sentences)
2.0%
samples
3
General showcase (greeting, fiqh, reflection)
0.6%
consistency
8… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/Fasih-TTS-Benchmark.LCAR-Hallucination-Benchmark
LCAR Hallucination Benchmark
LCAR Hallucination Benchmark is a manually reviewed speech benchmark for
studying acoustic-grounding failures in LLM-based ASR. It contains two
500-utterance suites: controlled speech synthesized with IndexTTS2 and speech
derived from openly released corpora. The benchmark covers translation or
transliteration, spoken or text-prompt instruction execution, unsupported
repetition, and catastrophic deletion.
The benchmark is a targeted stress set. It is… See the full description on the dataset page: https://huggingface.co/datasets/aguangguang/LCAR-Hallucination-Benchmark.gametime
Gametime Benchmark
The Gametime dataset provides lightweight, streaming-friendly splits for TTS/ASR/SpokenLM prototyping.For full details, please refer to the paper:👉 Game-Time: Evaluating Temporal Dynamics in Spoken Language Models
📦 Download Options
1️⃣ Recommended — Full ZIP Download
If you prefer the original folder layout you can download one of the ZIPs packaged in gametime/download/. There are two kinds available in this repository:… See the full description on the dataset page: https://huggingface.co/datasets/gametime-benchmark/gametime.faultbridge-asr-benchmark-evidence
FaultBridge Four-Model Code-Switched ASR Evidence
This evidence repository is publicly available for Sahara CodeSwitch Africa
Challenge evaluation, independent review, and reproducibility.
It contains the clip-level outputs and aggregate scores for FaultBridge's
four-model code-switched ASR benchmark. The frozen panel contains 400 natural
AfriSwitch clips: 100 each for Hausa-English, Igbo-English, Pidgin-English, and
Yoruba-English.
Project and Report
Source… See the full description on the dataset page: https://huggingface.co/datasets/Tiamz/faultbridge-asr-benchmark-evidence.nra-benchmarks
🧬 NRA Benchmark Datasets
All benchmark datasets for Neural Ready Archive (NRA) — the Rust-native streaming format for ML training.
Train on gigabytes of real data without downloading a single byte. NRA replaces tar.gz and zip for the AI era.
📦 Available Datasets
File
Domain
Source
Files
Size
food-101.nra
🖼️ Vision
ethz/food101
101,000 images
4.7 GB
wikitext.nra
📝 Text
Salesforce/wikitext
23,767 text files
7.6 MB
pokemon.nra
🎨 Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/zevatov/nra-benchmarks.kurdish-multidialect-asr-benchmark
Kurdish Dialect Speech Corpus
This project aims to provide a multi-dialect speech recognition benchmark for the Kurdish language. The Central Kurdish portion is the same as the Asosoft benchmark. The sentences were originally written in Central Kurdish (CKB), translated into other Kurdish dialects, and then recorded by native speakers.
The current version includes three Kurdish dialects: Central Kurdish, Northern Kurdish, and Southern Kurdish. A Hawrami version and the Badini… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/kurdish-multidialect-asr-benchmark.
