Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01H-Liu1997 /BEAT2audio1K<n<10K11 likes17k downloads3y agoHugging Face02verstar /MRSAudio MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audio, which limits the development of spatial audio generation and understanding. To address… See the full description on the dataset page: https://huggingface.co/datasets/verstar/MRSAudio.audio100K<n<1M7 likes10k downloads1y agoHugging Face03plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes3.8k downloads10mo agoHugging Face04Hui519 /WildElder WILDELDER: A CHINESE ELDERLY SPEECH DATASET FROM THE WILD WITH FINE-GRAINED MANUAL ANNOTATIONS Paper: https://huggingface.co/papers/2510.09344Code: https://github.com/NKU-HLT/WildElder WildElder is a speech dataset focused on elderly scenarios. It contains raw audio and corresponding text annotations and can be used for ASR, speaker-related tasks, and front-/back-end speech processing research. The data was collected and cleaned from real-world environments to preserve diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Hui519/WildElder.audioautomatic-speech-recognition10K<n<100K3 likes3.8k downloads5mo agoHugging Face05m-hamza-mughal /beat2-additional-annotations BEAT2 Official Release + Additional Annotations This is a fork of H-Liu1997/BEAT2 that adds annotations contributed by the RAG-Gesture (CVPR 2025) and MIBURI (CVPR 2026) projects. The base BEAT2-English data (motion, audio, TextGrids, semantic labels, pretrained motion-autoencoder weights) is inherited verbatim from upstream; the additional annotations from RAG-Gesture and MIBURI are pushed on top. Citations If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.audio1K<n<10K0 likes3.3k downloads4mo agoHugging Face06Hezep /AudioMarathon 🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audioaudio-classification1K<n<10K4 likes3.2k downloads11mo agoHugging Face07softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.5k downloads2mo agoHugging Face08MahiA /UrbanSound8K UrbanSound8K This is an audio classification dataset for Sound Event Classification. Classes = 10   ,   Split = Ten-Fold Structure audios folder contains audio files. csv_files folder contains CSV files for ten-fold cross-validation. To perform cross-validation on fold 1, train_1.csv will be used for the training split and test_1.csv for the testing split, with the same pattern followed for the other folds. To perform training and testing witout cross-validation, use… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/UrbanSound8K.audio10K<n<100K2 likes2.3k downloads2y agoHugging Face09KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes2.3k downloads5mo agoHugging Face10hzhongresearch /ahead_ds Another HEaring AiD DataSet (AHEAD-DS) Another HEaring AiD DataSet (AHEAD-DS) is an audio dataset labelled with audiologically relevant scene categories for hearing aids. Website Paper Code Dataset AHEAD-DS Dataset AHEAD-DS unmixed Models Description of data All files are encoded as single channel WAV, 16 bit signed, sampled at 16 kHz with 10 seconds per recording. Category Training Validation Testing All cocktail_party 934 134 266 1334 interfering_speakers… See the full description on the dataset page: https://huggingface.co/datasets/hzhongresearch/ahead_ds.audioaudio-classification1K<n<10K0 likes2k downloads9mo agoHugging Face11Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K28 likes1.8k downloads4mo agoHugging Face12ai4bharat /MANGO MANGO: A Corpus of Human Ratings for Speech MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages. Key Features: 255,150 human ratings of TTS-generated outputs and ground-truth human speech. Covers two major Indian languages: Hindi & Tamil, and English. Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.audiotext-to-speech10K<n<100K6 likes1.4k downloads1y agoHugging Face13fawzanaramam /the-truthaudio100K<n<1M0 likes1.1k downloads2y agoHugging Face14BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.1k downloads11mo agoHugging Face15PianoVAM /PianoVAM_v1 PianoVAM v1.2: A Multimodal Piano Performance Dataset Version History v1.2 (current). Adds Fingering/ (per-note fingering labels for 106 recordings) and Fingering_GT/ (manual fingering annotations for 11 recordings). All other files are unchanged from v1.1. v1.1. metadata.json is the canonical split file. Video files for the 'Sep 04-05' recordings are the sync-corrected versions, and files previously found to have video-MIDI synchronization issues have been… See the full description on the dataset page: https://huggingface.co/datasets/PianoVAM/PianoVAM_v1.audio1M<n<10M7 likes1k downloads9d agoHugging Face16nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes1k downloads10mo agoHugging Face17khaledalganem /sada2022 Dataset Card for SADA صدى Dataset Summary يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر. ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.audio100K<n<1M4 likes896 downloads2y agoHugging Face18egcortes /asr-jargon-specialized-vocabulary A Dataset for Evaluating ASR on Specialized Vocabulary Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026). Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code Configs Config Language Description synthetic_terms_en English Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms synthetic_terms_pt Portuguese Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes834 downloads3mo agoHugging Face19AudioMarathon /AudioMarathon AudioMarathon AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings. The release package in this directory is organized around 11 benchmark tasks spanning meeting summarization, automatic speech recognition, reading comprehension, authenticity detection, music genre classification, acoustic scene classification, emotion recognition, spoken named entity reasoning, sound event detection, speaker gender… See the full description on the dataset page: https://huggingface.co/datasets/AudioMarathon/AudioMarathon.audioaudio-classificationn<1K0 likes713 downloads5mo agoHugging Face20Mohammed01 /ArFakegated ArFake-Dataset ARFAKE: A Robust Framework for Multi-Dialect Arabic Speech Spoofing Detection Benchmark ARFAKE is the first end-to-end benchmark for Arabic speech spoofing detection across multiple dialects. The framework systematically generates synthetic Arabic speech, evaluates its intelligibility and realism, constructs a large-scale spoofing dataset, trains robust detectors, and evaluates generalization across both unseen generators and unseen dialects.… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed01/ArFake.audio10K<n<100K3 likes704 downloads4mo agoHugging Face21tsinghua-ee /QualiSpeech QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions 📄 Paper: https://arxiv.org/abs/2503.20290 QualiSpeech is a comprehensive English-language speech quality assessment dataset designed to go beyond traditional numerical scores. It introduces detailed natural language comments with reasoning, capturing low-level speech perception aspects such as noise, distortion, continuity, speed, naturalness, listening effort, and overall… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.audioaudio-text-to-text10K<n<100K26 likes612 downloads1y agoHugging Face22Silasimo /SynthGT SynthGT A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment Authors Silas Antonisen, Iván López-Espejo Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing. Overview SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations. The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.audioautomatic-speech-recognition1K<n<10K1 likes608 downloads2mo agoHugging Face23MVRL /TaxaBench-8k Paper: TaxaBind: A Unified Embedding Space for Ecological Applications Venue: WACV 2025 Github: https://github.com/mvrl/TaxaBind Dataset Name: TaxaBench-8k Dataset Description: TaxaBench-8k is a multimodal dataset containing six modalities - image, text, satellite image, audio, geographic location, and environmental features for evaluating large ecological models. Usage: Please use the test_df.csv for reading data which… See the full description on the dataset page: https://huggingface.co/datasets/MVRL/TaxaBench-8k.audiozero-shot-classification1K<n<10K1 likes440 downloads2y agoHugging Face24elliottash /doppelganger Doppelganger: Sound Effects and Their Synthetic Twins Benchmark for matching a synthetic sound effect to the real recording it was generated from. Paper: https://arxiv.org/abs/2607.04337 · Code: https://github.com/elliottash/doppelganger · models: https://huggingface.co/elliottash/doppelganger Contents sao_pairs/<CatID>/<instance_id>.wav — Stable-Audio-Open audio-conditioned synthetic twins (the main UCS corpus, one twin per verified real clip).… See the full description on the dataset page: https://huggingface.co/datasets/elliottash/doppelganger.audioaudio-classification1K<n<10K0 likes439 downloads2mo agoHugging Face25strikersoft /strikerData 🎧 StrikerData Overview StrikerData is an audio dataset developed by Strikersoft for research and development in audio and speech technologies.It contains human speech, environmental noise, and other sound types. The dataset is available for non-commercial use only, except for the company Strikersoft. Category Percentage of Total Dataset Clean human speech 20% Distorted speech 15% Human-made noise 15% Non-human noise 50% ⚖️ License… See the full description on the dataset page: https://huggingface.co/datasets/strikersoft/strikerData.audio10K<n<100K2 likes424 downloads9mo agoHugging Face26michaelcacioli /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes419 downloads3mo agoHugging Face27committa /serena-synthetic-it-28h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.audiotext-to-speech10K<n<100K1 likes415 downloads2mo agoHugging Face28MahiA /CREMA-D CREMA-D This is an audio classification dataset for Emotion Recognition. Classes = 6   ,   Split = Train-Test Structure audios folder contains audio files. train.csv for training split and test.csv for the testing split. Download import os import huggingface_hub audio_datasets_path = "DATASET_PATH/Audio-Datasets" if not os.path.exists(audio_datasets_path): print(f"Given {audio_datasets_path=} does not exist. Specify a valid path ending with 'Audio-Datasets'… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/CREMA-D.audio1K<n<10K0 likes390 downloads2y agoHugging Face29itruonghai /EK100 Motivation The actual download link is very slow, including the academic torrent. Therefore, to spare fellow community members from this misery, I am uploading the dataset here. Source You can fnd the original source to download the dataset: https://github.com/epic-kitchens/epic-kitchens-download-scripts Citation @INPROCEEDINGS{Damen2018EPICKITCHENS, title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset}, author={Damen, Dima and Doughty, Hazel and… See the full description on the dataset page: https://huggingface.co/datasets/itruonghai/EK100.tabularvoice-activity-detection100K<n<1M0 likes368 downloads5mo agoHugging Face30sirui1 /MADB-Dataset MADB: Music Aesthetics Dataset and Benchmark Dataset Description MADB is a large-scale dataset for music aesthetic evaluation, designed to support research on multi-dimensional and subjective music perception. The dataset contains approximately 10,000 music tracks, each annotated by multiple trained annotators across 10 perceptual dimensions and one overall score. In addition, each track includes textual comments and semantic tags (genre and mood), enabling… See the full description on the dataset page: https://huggingface.co/datasets/sirui1/MADB-Dataset.audio1K<n<10K0 likes354 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.