Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Harland /AudioMCQ-StrongAC-GeminiCoT [ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly. Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis. 🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.audio10K<n<100K7 likes4.8k downloads3mo agoHugging Face02RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K3 likes3.3k downloads4mo agoHugging Face03yangwang825 /audioset AudioSet AudioSet[1] consists of an expanding ontology of 527 audio event classes and a collection of 2M human-labelled 10-second sound clips drawn from YouTube. Some clips are missing on YouTube, so the number of files downloaded is different from time to time. This repository contains 20550 / 22160 of the balanced train set, 1913637 / 2041789 of the unbalanced train set (separated into 41 parts), and 18887 / 20371 of the evaluation set. The pre-process script can be found at… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/audioset.textaudio-classification1M<n<10M5 likes2.4k downloads3y agoHugging Face04AudioVisual-Caption /ASID-1M ASID-1M: Attribute-Structured and Quality-Verified Audiovisual Instructions [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce ASID-1M, a large-scale audiovisual instruction dataset built to support universal video understanding with fine-grained, controllable supervision. Most existing video-instruction data represents complex audiovisual content as a single, monolithic caption. This often leads to incomplete coverage (missing audio… See the full description on the dataset page: https://huggingface.co/datasets/AudioVisual-Caption/ASID-1M.textimage-text-to-text100K<n<1M85 likes1.3k downloads7mo agoHugging Face05MBZUAI /AudioJailbreak Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly. 📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.audioquestion-answering1K<n<10K9 likes764 downloads1y agoHugging Face06inclusionAI /AudioMCQ [ICLR 2026] AudioMCQ: Audio Multiple-Choice Question Dataset Also the official repository for the paper "Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models" News [2026.04] Update on MMSU Metric of released models: Based on community feedback, we identified a flaw in our evaluation script that artificially inflated the MMSU scores of our released models by ignoring sequence order. We sincerely apologize for… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AudioMCQ.text100K<n<1M15 likes606 downloads6mo agoHugging Face07liumindmind /Neko_Audio-30K_Longaudio10K<n<100K7 likes485 downloads4mo agoHugging Face08Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16ktextn<1K0 likes460 downloads2y agoHugging Face09codaco /audio CoDaCo - audios dataset This dataset was created using codaco.app. Description All data contributed to this campaign goes to the global CoDaCo datasets. Labels This dataset includes the following labels: Spoken text Tags Emotions AI generated Quality rating License This dataset is licensed under CC BY 4.0. You are free to share and adapt it for any purpose, including commercially, as long as you give appropriate credit. See LICENSE… See the full description on the dataset page: https://huggingface.co/datasets/codaco/audio.textn<1K0 likes362 downloads16h agoHugging Face10zihan-audio /seamless-bg Seamless Background Robustness Pilot V3 expansion available: v3/README.md documents the expanded 1,985-event pool. Use v3/events_all.jsonl and v3/clips_all.jsonl for combined manifests. The original pilot statistics and files below remain unchanged. A compact, paired-audio candidate pool for incremental full-duplex interaction alignment and later background-speech augmentation. Derived from Meta's Seamless Interaction, by selecting events from the original train split only. This… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/seamless-bg.audion<1K0 likes356 downloads1mo agoHugging Face11vhands /audio-event-classification-post-public audio-event-classification-post-public Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.textaudio-classification100K<n<1M1 likes205 downloads3mo agoHugging Face12nymtheescobar /bengali-talkshow-audio Bengali Talkshow Audio Dataset A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs. Dataset Description This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.audioaudio-classification1K<n<10K0 likes168 downloads8mo agoHugging Face13liumindmind /Neko_Audio-80K_Short audio10K<n<100K30 likes168 downloads4mo agoHugging Face14aseth125 /audio-hallucination-attack Audio Hallucination Attacks (AHA) Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models" It contains two subsets: AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training Audio Files The audio files are provided as compressed archives in this repository: File Contents Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.audioaudio-classification100K<n<1M2 likes163 downloads6mo agoHugging Face15AudioCC-Lab /PICSAFEv1 Speech Quality Test Labels PICSAFEv1 is a multi-source annotated test dataset for evaluating speech quality assessment and audio data filtering methods. It contains 10,728 audio samples drawn from 14 source datasets, with annotations from a vocabulary of 33 tags. These tags describe recording provenance, speech styles, speaking rate and pitch, speaker attributes, noise, reverberation, distortion, and transcript errors. These annotations support benchmarking quality metrics and… See the full description on the dataset page: https://huggingface.co/datasets/AudioCC-Lab/PICSAFEv1.textaudio-classification10K<n<100K0 likes135 downloads16d agoHugging Face16zed-m97 /nano4m-Audio nano4M-Audio — Team (week-1) Week-1 data preparation for nano4M-Audio, an extension of EPFL's nano4M (the educational nano version of 4M / 4M-21) that adds audio as a fifth modality alongside RGB, depth, surface normals and captions. This dataset covers all 12 VGGSound classes assigned to the three-person team: person classes 1 (Hassan) lions roaring, horse neighing, pig oinking, cow lowing 2 (Ziyad) dog barking, cat meowing, coyote howling, elephant trumpeting 3… See the full description on the dataset page: https://huggingface.co/datasets/zed-m97/nano4m-Audio.textaudio-classification1K<n<10K0 likes123 downloads5mo agoHugging Face17umd-zhou-lab /AVQA-Audio-Rubrics AVQA Audio-Reasoning Rubrics Project Page | Paper | Code Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward). Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.textaudio-classification10K<n<100K1 likes91 downloads2mo agoHugging Face18BUT-FIT /orca-audio-qa-annotations ORCA Audio QA Annotations Annotation data for training and evaluating ORCA (Open-ended Response Correctness Assessment), a scoring model for audio question-answering tasks. Paper: ORCA: Open-ended Response Correctness Assessment for Audio Question Answering — accepted to TACL 2026 Code & usage: github.com/BUTSpeechFIT/ORCA Pretrained Models: orca-olmo-2-1b-multinomial orca-gemma-3-4b-it-multinomial orca-llama-3.2-3b-it-multinomial Dataset overview ORCA is… See the full description on the dataset page: https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations.texttext-classification100K<n<1M0 likes88 downloads3mo agoHugging Face19Aarjanm /youtube_audio_processed_dataset YouTube Audio Processed Dataset This dataset contains high-quality segmented speech datasets preprocessed from YouTube videos using the Emilia preprocessor framework. Dataset Structure Each row in the dataset contains a segmented audio clip, its aligned high-fidelity transcript, speaker labeling, and objective speech quality assessment scores. Features file_name: Audio column containing the relative path to the segmented .mp3 clip. text: The… See the full description on the dataset page: https://huggingface.co/datasets/Aarjanm/youtube_audio_processed_dataset.audion<1K0 likes88 downloads7d agoHugging Face20norwooodsystems /audio-files-diarisationaudion<1K0 likes76 downloads1y agoHugging Face21lilonghao /Audio-Cogito Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Audio-Cogito is a large-scale audio reasoning dataset introduced in the paper Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models. The released data contains 545k high-quality audio reasoning samples spanning sound, speech, and music domains. Each sample includes label annotations, Chain-of-Thought (CoT) annotations, and final answers. Links… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/Audio-Cogito.text100K<n<1M4 likes75 downloads3mo agoHugging Face22vhands /audio-music-mir-post-public audio-music-mir-post-public Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.textaudio-classification100K<n<1M0 likes67 downloads3mo agoHugging Face23BubuDavid /Selena-Gomez-With-Lyrics-And-Spotify-Audio-Featurestabularn<1K0 likes66 downloads3y agoHugging Face24sonalkum /AudioSkills-Llama3 AudioSkills-XL Dataset To promote the development of open source models, we have released AudioSkills using the exact same method generated with Llama 3.1-8B Instruct instead of GPT4o in the original. Project page | Paper | Code Dataset Description AudioSkills-XL is a large-scale audio question-answering (AQA) dataset designed to develop (large) audio-language models on expert-level reasoning and problem-solving tasks over short audio clips (≤30 seconds). It… See the full description on the dataset page: https://huggingface.co/datasets/sonalkum/AudioSkills-Llama3.textaudio-text-to-text100K<n<1M0 likes57 downloads1y agoHugging Face25RKB109 /audio-event-triage-20260912-dataset Audio Event Triage Baseline Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260912-dataset.textaudio-classificationn<1K0 likes57 downloads28d agoHugging Face26CentificAIResearch /audio-loop Audio Loop: Long-Audio Reasoning Benchmarks An evaluation set for long-form audio reasoning. It accompanies the paper AudioLoop: Program-Guided Iterative Reasoning Over Long-Form Speech (NeurIPS 2026 Workshop on Long Context Foundation Models), where it is referred to as IQ2-QA. Config Task Audio source Items iq2_qa Multi-hop QA over debates Recorded live debates 13 Load annotations from datasets import load_dataset iq2 =… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/audio-loop.textquestion-answeringn<1K0 likes56 downloads6d agoHugging Face27vhands /audio-reasoning-qa-post-public audio-reasoning-qa-post-public Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.textquestion-answering100K<n<1M0 likes52 downloads3mo agoHugging Face28RKB109 /audio-event-triage-20260922-dataset Audio Event Triage Baseline Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260922-dataset.textaudio-classificationn<1K0 likes52 downloads18d agoHugging Face29RKB109 /audio-event-triage-20261002-dataset Audio Event Triage Baseline Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20261002-dataset.textaudio-classificationn<1K0 likes51 downloads8d agoHugging Face30bnovikov /gemma-4-e4b-audio-qa Gemma-4 E4B Audio-QA Training Mix A 91k-row audio question-answering dataset assembled from four public upstream datasets, formatted as ChatML-style conversations for instruction-tuning an audio-language model. This is the exact training data used for bnovikov/gemma-4-e4b-audio-v3. Important: this repository contains only the metadata and prompts/answers. The audio files are NOT hosted here. Each audio_path is a source-tagged ID like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.textaudio-classification10K<n<100K0 likes45 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.