datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GTSinger
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University
Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.VietSuperSpeech
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
Total samples: 32,267
Train samples: 29,041
Dev samples: 3,226
Total duration: 103.18 hours
Sample rate: 16000 Hz
Average segment length: ~12 seconds
Source Datasets
asr_dataset_nguoivietdailynews
asr_dataset_nguyenkhangofficial
asr_dataset_trinhlieu
Format
The dataset follows Icefall format:
train.json: Training samples
dev.json: Development samples
manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.tlott-digital-products
T. Lott Digital Products
Digital product files for T. Lott's online store.
Products
Audiobooks (MP3)
eBooks (PDF)
Software (ZIP)
Cover images (PNG)
Download URLs
Files can be downloaded directly:
https://huggingface.co/datasets/ziggylott/tlott-digital-products/resolve/main/{filepath}
vibevoice-quran_persian-single-speakerAudio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.AudioMCQ-StrongAC-GeminiCoT
[ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT
This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly.
Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis.
🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.AVQA
Summary | 摘要
This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys.
The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds).
Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.vibevoice-gptinformal_persian-single-speakerNeko_Audio-30K_LongStreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.apptek_callcenter_dialogues
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents
across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions.
128.6 hours of speech
14 English accent groups
16 service domains
5–15 minute conversations (long-form)
Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.H3-Character-Swap-v1
H3 Character Swap v1
A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos.
134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately.
Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.TalkVid
TalkVid Dataset
This repository hosts the TalkVid dataset.
Paper: TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
Arxiv paper: https://arxiv.org/abs/2508.13618
Project Page: https://freedomintelligence.github.io/talk-vid
GitHub: https://github.com/FreedomIntelligence/TalkVid
Abstract
Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TalkVid.gdpval_preference_rubricsAudioJailbreak
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly.
📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.cmi-pref
CMI-Pref Dataset
CMI-Pref is a music preference comparison dataset for multimodal music generation research. Each record represents a single human vote comparing two generated audio samples, with preferences annotated along two dimensions (musicality and alignment) and a confidence score for each preference.
⚠️ Important Notes
The modality distribution below is computed over train + test combined.
The dataset contains overlapping votes by design: multiple users may vote… See the full description on the dataset page: https://huggingface.co/datasets/HaiwenXia/cmi-pref.GSmeow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.daily-bio-newsvoice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.ContextASR-Bench
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.OpenGameArt-CC0
Dataset Card for OpenGameArt-CC0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.UltraVoice
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
📝 Abstract
Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset… See the full description on the dataset page: https://huggingface.co/datasets/tutu0604/UltraVoice.dahih-tts2-demucs-cleanedHearInContextEnglish | 中文
HearInContext
A Benchmark for Implicit Context in Speech Recognition
Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative.
Same audio. Different contexts. Different meanings.
HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.locomomicroduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.
