Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-audio /open-asr-leaderboard-resultstabularn<1K2 likes13k downloads1d agoHugging Face02Hezep /AudioMarathon 🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audioaudio-classification1K<n<10K4 likes3.2k downloads11mo agoHugging Face03hf-audio /leaderboard_longformtabularn<1K0 likes2.2k downloads11d agoHugging Face04softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes1.7k downloads2mo agoHugging Face05nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes1.6k downloads10mo agoHugging Face06akazemian /audio-htmltabular10K<n<100K0 likes745 downloads1y agoHugging Face07AudioMarathon /AudioMarathon AudioMarathon AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings. The release package in this directory is organized around 11 benchmark tasks spanning meeting summarization, automatic speech recognition, reading comprehension, authenticity detection, music genre classification, acoustic scene classification, emotion recognition, spoken named entity reasoning, sound event detection, speaker gender… See the full description on the dataset page: https://huggingface.co/datasets/AudioMarathon/AudioMarathon.audioaudio-classificationn<1K0 likes727 downloads5mo agoHugging Face08paodigitalhub /pao-audio-dataset 🎙️ Pa'O Audio Dataset ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ 📌 Project Summary The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ). Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.audioautomatic-speech-recognitionn<1K1 likes161 downloads17d agoHugging Face09liuhuadai /AudioCoT AudioCoT AudioCoT is an audio-visual Chain-of-Thought (CoT) correspondent dataset for multimodal large language models in audio generation and editing. Homepage: ThinkSound Project Paper: arXiv:2506.21448 GitHub: FunAudioLLM/ThinkSound Dataset Overview Each CSV file contains three fields: id — Unique identifier for the sample caption — Simple audio description prompt caption_cot — Chain-of-Thought prompt for audio generation This dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/liuhuadai/AudioCoT.text100K<n<1M4 likes144 downloads1y agoHugging Face10Nielzac /CoM_Audio_Image_LLM_Generation This dataset is a Mixture of DIBT/10k_prompts_ranked, lj_speech and Falah/image_generation_prompts_SDXL Repartition Why this dataset ? Training a multimodal router holds crucial significance in the realm of artificial intelligence. By harmonizing different specialized models within a constellation, the router plays a central role in intelligently orchestrating tasks. This approach not only enables precise classification but also paves the way for diverse… See the full description on the dataset page: https://huggingface.co/datasets/Nielzac/CoM_Audio_Image_LLM_Generation.text10K<n<100K9 likes95 downloads3y agoHugging Face11mueller91 /human-perception-audio-deepfake-2026 Human Audio Deepfake Perception 2026 A large-scale listening study evaluating how well humans detect modern audio deepfakes. The dataset contains 35,532 deepfake-detection judgments from 1,768 anonymous participants across 138 TTS and voice-conversion systems, collected via a publicly accessible online listening game in 2025–2026. This is the successor to the 2021 ASVspoof-2019 perception study (Müller, Pizzi & Williams, 2022) and extends the same paradigm to modern systems… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/human-perception-audio-deepfake-2026.textaudio-classification10K<n<100K4 likes94 downloads5mo agoHugging Face12Sodkhuu /Mongolian_audiosaudio1K<n<10K1 likes71 downloads1y agoHugging Face13Reihaneh /audio_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Reihaneh/audio_dataset.audion<1K0 likes69 downloads3y agoHugging Face14recommendedaudiobooks /audiobook-listener-fit-framework Audiobook Listener-Fit Framework The Audiobook Listener-Fit Framework is an open structured taxonomy developed by Recommended Audiobooks for describing characteristics that influence the audiobook listening experience. Traditional ratings mostly describe whether listeners liked a title. This framework is designed to describe how an audiobook listens and which types of listeners may be better suited to it. Purpose The framework organizes audiobook characteristics… See the full description on the dataset page: https://huggingface.co/datasets/recommendedaudiobooks/audiobook-listener-fit-framework.textn<1K0 likes69 downloads23d agoHugging Face15igorriti /ambience-audio Ambience audio dataset Overview This dataset was generated by scraping videos from prominent YouTube channels focused on ambient audio. The dataset includes a collection of videos that feature various ambient sounds, such as nature sounds, relaxing music, and environmental noises. For each video, essential metadata was extracted, and a caption was generated using an AI model to enhance the discoverability of the content. This dataset can be useful in various applications… See the full description on the dataset page: https://huggingface.co/datasets/igorriti/ambience-audio.image1K<n<10K6 likes65 downloads2y agoHugging Face16plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes61 downloads1mo agoHugging Face17FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes60 downloads5mo agoHugging Face18Chengxiang1122 /mcl-mmcl-audiocapsaudio10K<n<100K0 likes43 downloads8mo agoHugging Face19vancenceho /youtube-spotify-audio-features Spotify–YouTube Audio Features Tabular librosa audio features for tracks aligned with the Spotify / YouTube pipeline in the viral-content-predictor project. Each row is one Spotify track_id matched to a downloaded YouTube audio clip; features are aggregated statistics (mean / std) computed on the decoded waveform. Files File Description audio_features.csv One row per track: track_id, 89 derived feature dimensions (means/stds), extraction_success, error_message.… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/youtube-spotify-audio-features.tabular10K<n<100K0 likes41 downloads6mo agoHugging Face20JavisVerse /JavisData-Audioaudio100K<n<1M0 likes40 downloads1y agoHugging Face21Olivia714 /audiocapsaudio1K<n<10K0 likes36 downloads2y agoHugging Face22cdactvm /save_audio_punjabi Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cdactvm/save_audio_punjabi.audion<1K0 likes35 downloads2y agoHugging Face23ilyaslbern7347 /darija-hotel-audioaudion<1K0 likes35 downloads6mo agoHugging Face24freococo /9000hours_voa_burmese_audio Overview VOA Burmese radio news archive covering Morning (နံနက် ၅:၃၀ – ၆:၃၀) and Evening (ညပိုင်း ၉:၀၀ – ၁၀:၀၀) programmes for every calendar day from 2012-09-16 → 2025-06-09. Metric Value Hours / rows 9 159 Files per day 2 (morning, evening) Typical file size 15 – 50 MB Licence Public-domain (VOA staff recordings, U.S. 17 U.S.C. § 105) This dataset upgrades Burmese from low-resource to mid-resource status for speech research, enabling self-supervised… See the full description on the dataset page: https://huggingface.co/datasets/freococo/9000hours_voa_burmese_audio.textautomatic-speech-recognition1K<n<10K1 likes34 downloads1y agoHugging Face25Dannynis /audioset-tf-boxes TF-SED AudioSet Time-Frequency Boxes Weak 2-D time-frequency bounding boxes (time and frequency extent) for a subset of AudioSet. Annotations only — no audio is included. Boxes for a clip labelled Speech / Bicycle / Music / Vehicle, drawn over its spectrogram. Each box localizes an event in both time and frequency. Using the data Load the manifest and parse the per-clip boxes: from datasets import load_dataset import json ds =… See the full description on the dataset page: https://huggingface.co/datasets/Dannynis/audioset-tf-boxes.tabularaudio-classification10K<n<100K0 likes32 downloads4mo agoHugging Face26SumitY22 /boat-vs-noise-audio-productstabularn<1K0 likes29 downloads9d agoHugging Face27AlexandrVictorov /audio_mistakes_examplestext1K<n<10K0 likes26 downloads13d agoHugging Face28abdelhaqueidali /Tizuzaf-Audio-Datasetaudioautomatic-speech-recognitionn<1K1 likes24 downloads4mo agoHugging Face29multimodalart /test-audio-datasettextn<1K0 likes22 downloads2y agoHugging Face30cdactvm /save_audioaudion<1K0 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.