Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AaronZ345 /GTSinger GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.audiotext-to-audio10K<n<100K17 likes36k downloads1y agoHugging Face02ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes7.9k downloads8mo agoHugging Face03thanhnew2001 /VietSuperSpeech VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.audio10K<n<100K6 likes7.7k downloads8mo agoHugging Face04ziggylott /tlott-digital-products T. Lott Digital Products Digital product files for T. Lott's online store. Products Audiobooks (MP3) eBooks (PDF) Software (ZIP) Cover images (PNG) Download URLs Files can be downloaded directly: https://huggingface.co/datasets/ziggylott/tlott-digital-products/resolve/main/{filepath} audion<1K0 likes6.9k downloads1mo agoHugging Face05Dorsaasgari /vibevoice-quran_persian-single-speakeraudio1K<n<10K1 likes5.1k downloads22d agoHugging Face06RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K3 likes4.9k downloads4mo agoHugging Face07Harland /AudioMCQ-StrongAC-GeminiCoT [ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly. Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis. 🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.audio10K<n<100K7 likes3.5k downloads3mo agoHugging Face08Joysw909 /AVQA Summary | 摘要 This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.audioquestion-answering10K<n<100K2 likes3.5k downloads11mo agoHugging Face09shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes3.4k downloads9mo agoHugging Face10Dorsaasgari /vibevoice-gptinformal_persian-single-speakeraudio1K<n<10K0 likes3.3k downloads22d agoHugging Face11liumindmind /Neko_Audio-30K_Longaudio10K<n<100K7 likes3k downloads4mo agoHugging Face12zhifeixie /StreamAudio-2M StreamAudio-2M Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips are organised into six task subsets. Subsets Subset Rows Description Stream_Audio_Understanding 90,738 Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA Real_time_ASR 28,109 Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.tabularaudio-classification100K<n<1M30 likes2.9k downloads4mo agoHugging Face13apptek-com /apptek_callcenter_dialogues AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions. 128.6 hours of speech 14 English accent groups 16 service domains 5–15 minute conversations (long-form) Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.audioautomatic-speech-recognition1K<n<10K40 likes2.8k downloads1mo agoHugging Face14akatz-ai /H3-Character-Swap-v1 H3 Character Swap v1 A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos. 134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately. Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.imagen<1K4 likes2.8k downloads11d agoHugging Face15FreedomIntelligence /TalkVid TalkVid Dataset This repository hosts the TalkVid dataset. Paper: TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis Arxiv paper: https://arxiv.org/abs/2508.13618 Project Page: https://freedomintelligence.github.io/talk-vid GitHub: https://github.com/FreedomIntelligence/TalkVid Abstract Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TalkVid.audioimage-to-videon<1K22 likes2.8k downloads1y agoHugging Face16cm2435-new /gdpval_preference_rubricsaudion<1K0 likes2.7k downloads6mo agoHugging Face17MBZUAI /AudioJailbreak Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly. 📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.audioquestion-answering1K<n<10K9 likes2.2k downloads1y agoHugging Face18HaiwenXia /cmi-pref CMI-Pref Dataset CMI-Pref is a music preference comparison dataset for multimodal music generation research. Each record represents a single human vote comparing two generated audio samples, with preferences annotated along two dimensions (musicality and alignment) and a confidence score for each preference. ⚠️ Important Notes The modality distribution below is computed over train + test combined. The dataset contains overlapping votes by design: multiple users may vote… See the full description on the dataset page: https://huggingface.co/datasets/HaiwenXia/cmi-pref.audiotext-to-audio1K<n<10K5 likes1.8k downloads7mo agoHugging Face19m-a-p /GSaudio1K<n<10K1 likes1.8k downloads1y agoHugging Face20smgjch /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.audio10K<n<100K3 likes1.7k downloads5mo agoHugging Face21titasmallick96 /daily-bio-newsaudion<1K0 likes1.5k downloads5h agoHugging Face22besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K14 likes1.4k downloads6d agoHugging Face23MrSupW /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K38 likes1.3k downloads1y agoHugging Face24nyuuzyou /OpenGameArt-CC0 Dataset Card for OpenGameArt-CC0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.audioimage-classification10K<n<100K10 likes1.2k downloads1y agoHugging Face25tutu0604 /UltraVoice UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset… See the full description on the dataset page: https://huggingface.co/datasets/tutu0604/UltraVoice.audiotext-to-speech100K<n<1M16 likes1.2k downloads11mo agoHugging Face26YomnaGharib /dahih-tts2-demucs-cleanedaudio10K<n<100K1 likes1.2k downloads4mo agoHugging Face27OPPOer /HearInContextEnglish | 中文 HearInContext A Benchmark for Implicit Context in Speech Recognition Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative. Same audio. Different contexts. Different meanings. HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.audioautomatic-speech-recognition100K<n<1M1 likes1.1k downloads15d agoHugging Face28insight /locomoaudio100K<n<1M1 likes1.1k downloads6mo agoHugging Face29pollen-robotics /microduck-emotions Microduck Emotions A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.audioroboticsn<1K6 likes1.1k downloads29d agoHugging Face30jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.