datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GAIA
GAIA dataset
GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc).
We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format.
Data and leaderboard
GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for 8 language subsets,
with 10,100 examples selected from Common Voice 17.0. Each example supplies
a speaker reference and an independently selected target text, with human
recordings as WER/CER and speaker-similarity anchors when the corresponding
audio is available. Corpus WER and the WavLM-FT evaluation follow
seed-tts-eval.
Version 3.1. Adds ru, kk through the same S1–S4 selection… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.stt-benchmark-dataDataset for Pipecat Speech-to-Text benchmarks:
https://github.com/pipecat-ai/stt-benchmark
VoiceIsolation-Benchmark-Dataset
Voice Isolation Benchmark Dataset
265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing.
265 recordings · 47 speakers · 3 scenarios · 65 scripts
Why this dataset exists
Modern STT engines handle noise well. They still fail when a second person talks near the microphone: they transcribe the wrong speaker, and voice… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset.turn-benchmark-dev
TurnBench - Dev Set
TurnBench is a benchmark for evaluating
conversational turn-taking: end-of-turn and interruption detection on real
annotated two-speaker conversations.
This repository contains the development split: 38 English conversations,
about 7.3 hours of audio, packaged as one row per conversation. Each row contains
two time-aligned per-speaker audio streams plus three independent annotator
tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.nsanku-tts-benchmark-audioMSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.musicai-background-music-audio-llm-benchmark
Does Background Music Matter to Speech in Pre-trained Language Models
The completed September 2026 study covers 8 model families, 55 instrumental recordings, and 10 evaluation settings. It studies how adding background music to the same spoken question changes model responses.
Latest release and artifact guide
Technical report PDF
Complete LaTeX project
LaTeX GitHub repository
Matrices, figures, and supporting data
Regenerated speech and mixtures: 550 archives / 250,800… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/musicai-background-music-audio-llm-benchmark.video-full-duplex-benchmark
VideoFDB: Video-Full-Duplex-Benchmark
Project Page · HuggingFace · Paper (arXiv)
Dataset Description
A benchmark dataset of annotated, two-person video conference recordings designed to support the evaluation of multimodal AI agents in conversational settings. The dataset covers 11 distinct conversational dynamics — spanning verbal, nonverbal, and mixed-modality behavior — annotated through a three-pass human-in-the-loop pipeline.
The benchmark consists of trimmed… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-full-duplex-benchmark.quranic-asr-benchmark
Quranic ASR Benchmark - leakage-free, held-out
A small, leakage-free benchmark (600 clips) for evaluating Arabic ASR on Quranic recitation
(Hafs riwayah). Every clip is verified absent from our training data, so it measures
generalization, not memorization. Same clips + same scoring for every model.
📊 Live leaderboard: https://huggingface.co/spaces/Muno459/quranic-asr-leaderboard
The set (600 clips, 200 per source)
Source
n
What it is
everyayah_heldout… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-benchmark.audio_benchmarksYingMusic-SVC_Difficulty-Graded_Benchmark
YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases
github:YingMusic-SVC
The difficulty grading benchmark for SVC. Each sample provides a clean vocalist (lead)/harmony (back)/ full song (mix)/ full vocal (mix_vocal) and the lead vocalist obtained using our self-developed separation model (ourlead).
The metadata records the gender of the singer for each sample, as well as the presence of echo and reverberation in… See the full description on the dataset page: https://huggingface.co/datasets/GiantAILab/YingMusic-SVC_Difficulty-Graded_Benchmark.persian-accents-benchmark
Persian Accents Benchmark
Dataset Summary
A benchmark for Persian automatic speech recognition (ASR): 279 short
utterances of informal Persian (Farsi) dialect speech across 16 regional accents,
released as a fixed evaluation set. Total audio duration is approximately 4.4
hours. The primary label is the transcription; each utterance also carries an
accent label (usable for accent classification as a secondary task) and an
emotion label as auxiliary metadata.
This… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-accents-benchmark.turn-benchmark-test
TurnBench - Test Set
TurnBench is a benchmark for evaluating
conversational turn-taking: end-of-turn and interruption detection on real
annotated two-speaker conversations.
This repository contains the test split: 116 English conversations, packaged
as one row per conversation. Each row contains two time-aligned per-speaker audio
streams. The six annotation columns follow the same schema as the dev set but are
intentionally… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-test.quran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.Ming-Freeform-Audio-Edit-Benchmark
README
Introduction
This repository hosts Ming-Freeform-Audio-Edit, the benchmark test set for evaluating the downstream editing tasks of the Ming-UniAudio model.
This test set covers 7 distinct editing tasks, categorized as follows:
Semantic Editing (3 tasks):
Free-form Deletion
Free-form Insertion
Free-form Substitution
Acoustic Editing (5 tasks):
Time-stretching
Pitch Shifting
Dialect Conversion
Emotion Conversion
Volume Conversion
The audio samples are sourced… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ming-Freeform-Audio-Edit-Benchmark.entity-transcription-benchmark
Entity Transcription Benchmark
Measures whether a speech recognition system transcribes named entities
correctly — as distinct from word error rate.
WER weights every token equally. The tokens that matter for redaction, lookup,
routing and search are proper nouns, and they are a small fraction of any
transcript. A system can improve WER while getting worse at exactly the words a
downstream consumer needs, and nothing in the standard evaluation will show it.
2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.benchmark_DEMAND_noise
benchmark_DEMAND_noise
This dataset is a segmented subset derived from DEMAND: Diverse Environments Multichannel Acoustic Noise Database.
It is prepared for the SPARCO noise ablation benchmark. The intended use is to provide fixed 4-second environmental noise segments for:
AUROC-based SAE noise-related feature selection
binary noise-presence scorer training
scorer threshold calibration
final held-out benchmark evaluation
Source
Original source:
DEMAND: Diverse… See the full description on the dataset page: https://huggingface.co/datasets/SPARCO-project/benchmark_DEMAND_noise.VoiceIsolation-Benchmark-Dataset-Processed
Voice Isolation Benchmark – Processed Audios
Audio samples processed by four Krisp Voice Isolation models. Each scenario folder contains subfolders for every model, with one processed .wav file per original sample.
Voice Isolation Models
VI 2.5 Default (vi_2_5_default)
The main Voice Isolation model. Strongest at removing noise and other speakers. Goes fully silent when only a bystander is talking.
VI 2.5 Balanced (vi_2_5_balanced)… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset-Processed.TTS-Voice-Direction-Benchmark
TTS Voice Direction Benchmark
🏆 Leaderboard | 🛠️ Evaluation Suite
TTS Voice Direction is a benchmark of 700 reference-conditioned speech
generation tasks. It evaluates whether a text-to-speech model can preserve a
reference speaker while following a natural-language direction that controls
how a new transcript is performed.
The benchmark emphasizes practical voice acting beyond basic emotion control.
It covers accent, acoustic delivery, vocal events, emotion, physiological… See the full description on the dataset page: https://huggingface.co/datasets/BreezeBlue/TTS-Voice-Direction-Benchmark.amharic-tts-benchmark
Amharic TTS Benchmark
Seven text-to-speech systems and the original human recordings, evaluated on 100
Amharic prompts from three open datasets. Run date 2026-08-12.
Published results: addisassistant.com/benchmarks
Reproduce the CER/WER results
python score.py
No arguments. It reads data/judge_rows.jsonl, recomputes every character and
word edit count from the transcripts and writes data/summary.json.
This covers the CER/WER results only. Listening scores… See the full description on the dataset page: https://huggingface.co/datasets/addisai/amharic-tts-benchmark.MMAG-Benchmark
MMAG: A Multi‑Control Mixed Audio Generation Benchmark
MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
Dataset Structure
The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.Agri_STT_Benchmarking_Dataset
Agri STT Benchmarking Dataset
10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository.
Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.gtts-benchmark-audioVaani-Benchmark-V1.0
Vaani-Benchmark-V1.0
A curated Hindi ASR evaluation set collected as part of the Vaani project at IISc Bengaluru. This is a separate, held-out collection — distinct from the publicly released Vaani dataset — built specifically for benchmarking. This benchmark contains 5,050 audio segments from 1,103 speakers across 104 Indian districts, each with three independent human transcriptions.
Evaluation Toolkit
A standalone toolkit implementing this benchmark's scoring… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0.nexa-audiolm-instuct-benchmarkMSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.ParlaSpeech-HR-benchmark_v3
ParlaSpeechHR Benchmark v3
A curated benchmark dataset of 22,008 Croatian parliamentary speech clips extracted from ParlaSpeech-HR v3. Each clip includes aligned audio (WAV) and TextGrid annotations for linguistic analysis.
Contents
22,008 audio segments (various durations)
17,622 clips with complete TextGrid triplets:
.align (word-level boundaries via WordAlign tier)
.stress (primary stress frame labels; derivative of .align)
.pause (filled pause annotations… See the full description on the dataset page: https://huggingface.co/datasets/porupski/ParlaSpeech-HR-benchmark_v3.SEAR
SEAR: Spoofing Evidence-Grounded Audio Reasoning
SEAR is an audio question-answering benchmark for testing whether audio language models
can identify and quantify signal-level acoustic anomalies and use them as evidence for
audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR
separates deepfake detection, forgery-cue identification, acoustic measurement, and
forensic rationale generation.
SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.vocal-money-codeswitch-asr-benchmark
Vocal Money — Yoruba–English Code-Switched ASR Benchmark
A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally
code-switched Yoruba–English speech, together with the reference transcriptions and the output of
every system on every clip, so that the published results can be recomputed or contradicted.
Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026.
Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.
