datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVQA
Summary | 摘要
This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys.
The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds).
Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.AudioJailbreak
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly.
📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.LongAudioSpan
LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension
Introduction
LongAudioSpan is a benchmark for long-form audio comprehension, spanning
diverse durations and cognitive depths.
Questions come from two complementary paths:
Native QA: questions drawn from the audio's natural content.
Anchor QA: questions built around acoustic anchors planted into the
audio.
Each path is scored in its own mode:
Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/holvan/LongAudioSpan.MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.SpokenNativQA
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.RAIL
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
NeurIPS 2026
Hongyu Jin1,*, Siyi Wang1,*, Yang Xiao1,*, Jiaheng Dong1,*, Shihong Tan4, Kaiyuan Peng1, Georgiana Juravle2, Shanquan Chen3, Gongping Huang4, Hong Jia5, Eun-Jung Holden1, James Bailey6, Ting Dang1,†
1The University of Melbourne, 2Alexandru Ioan Cuza University of Iași, 3The University of Hong Kong, 4Wuhan University, 5The University of Auckland, 6Monash University… See the full description on the dataset page: https://huggingface.co/datasets/AIMS-RAIL/RAIL.valor32k-avqa-v2
Valor32k-AVQA v2.0
Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position.
Links
Paper: ACM Digital Library
Project page: inesriahi.github.io/valor32k-avqa-2
Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.FTAR
TimeAudio: Bridging Temporal Gaps in Large Audio-Language Models
Abstract
Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perception, leading to constrained capabilities on fine-grained tasks. We identify three key aspects that limit… See the full description on the dataset page: https://huggingface.co/datasets/lysanderism/FTAR.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.SEAR
SEAR: Spoofing Evidence-Grounded Audio Reasoning
SEAR is an audio question-answering benchmark for testing whether audio language models
can identify and quantify signal-level acoustic anomalies and use them as evidence for
audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR
separates deepfake detection, forgery-cue identification, acoustic measurement, and
forensic rationale generation.
SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.OmniCoding
OmniCoding
A multimodal terminal-tool-use SFT/RL dataset. Each record is a
question + verifiable answer + media (video/audio/image) — the target
agent is expected to operate on the media via a Linux terminal (ffmpeg,
ffprobe, whisper, python, etc.) rather than a GUI.
Aggregated and filtered from four upstream sources, with a single unified
schema, global dedup, and category-balanced sampling.
Records
Source
n
Omnimodal-Agent-SFT-2K (RUC-NLPIR) — agentic… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/OmniCoding.StethoBench
StethoBench
StethoBench is a comprehensive benchmark for cardiopulmonary auscultation, comprising 77,027 instruction–response pairs synthesized from 16,125 labeled recordings across 11 public datasets. It is the training and evaluation benchmark for StethoLM, published in the Transactions on Machine Learning Research (TMLR).
Dataset Description
StethoBench was constructed by synthesizing instruction–response pairs from existing labeled cardiopulmonary audio datasets… See the full description on the dataset page: https://huggingface.co/datasets/askyishan/StethoBench.PhoneticQA-SO762
PhoneticQA-SO762 v0.1
PhoneticQA-SO762 is a small, word-level multiple-choice AudioQA benchmark derived from SpeechOcean762. It was created and released by Haopeng Geng as an early benchmark for comparing human and speech-language-model sensitivity to salient mispronunciations.
Benchmark design
160 full-utterance audio questions: 32 dev and 128 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-SO762.AVQA-Audio-Rubrics
AVQA Audio-Reasoning Rubrics
Project Page | Paper | Code
Audio-grounded, binary-evaluable evaluation rubrics for the full
AVQA training set, generated for
process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with
rubric-as-reward).
Each training question is annotated with 5 rubrics, one per evaluation
facet, that judge the quality of an audio-reasoning response — not just final
answer correctness. The rubrics are designed to be scored Yes/No by an
LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/Weisiqing123/ONOTE.voices-of-civilizations
Voices of Civilizations (VoC)
Voices of Civilizations (VoC) is the first multilingual QA benchmark designed to assess audio LLMs’ cultural comprehension using full-length music recordings. VoC spans:
38 languages 🇸🇦 Arabic (ar), 🇧🇩 Bengali (bn), 🇧🇬 Bulgarian (bg), 🇨🇳 Chinese (zh), 🇭🇷 Croatian (hr), 🇨🇿 Czech (cs), 🇩🇰 Danish (da), 🇳🇱 Dutch (nl), 🇬🇧 English (en), 🇪🇪 Estonian (et), 🇫🇮 Finnish (fi), 🇫🇷 French (fr), 🇩🇪 German (de), 🇬🇷 Greek (el), 🇮🇱 Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/voices-of-civilizations.TRIAD
TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models
TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio.
Overview
TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is needed… See the full description on the dataset page: https://huggingface.co/datasets/triad-26/TRIAD.hans-10k
Hans-10K · DPO recipe for the audio-visual Clever Hans
DPO training data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans 🐎 —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-10K is the 10,383-sample best-recipe preference-pair dataset
that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/chengxin666/ONOTE.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.VoiceGiraffe
VoiceGiraffe (Benchmark)
VoiceGiraffe is a benchmark for evaluating large audio language models (LALMs) on hour-level, long-context audio understanding. It contains 1,500 curated question-answer triplets over real-world recordings central to real-world long-form audio understanding — broadcast, sports/esports commentary, news, and TV drama — organized into a dual-level taxonomy of single-hop perception and multi-hop reasoning.
This repo is public and holds the annotations… See the full description on the dataset page: https://huggingface.co/datasets/Jashin-Yeah/VoiceGiraffe.LongAudioSpan
LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension
Introduction
LongAudioSpan is a benchmark for long-form audio comprehension, spanning
diverse durations and cognitive depths.
Questions come from two complementary paths:
Native QA: questions drawn from the audio's natural content.
Anchor QA: questions built around acoustic anchors planted into the
audio.
Each path is scored in its own mode:
Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/audiospan/LongAudioSpan.PhoneticQA-L2ARCTIC
PhoneticQA-L2ARCTIC v0.1
PhoneticQA-L2ARCTIC is a small, word-level multiple-choice AudioQA probe derived from L2-ARCTIC. It was created and released by Haopeng Geng to study the gap between phoneme-level error annotations and perceptually salient mispronunciations.
Benchmark design
120 full-utterance audio questions: 24 dev and 96 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most clearly… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-L2ARCTIC.hans-sft-4k
Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans
Supervised fine-tuning (SFT) data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-SFT-4K is the 3,834-sample SFT mix that teaches models to… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.vggsync-3k
VGGSync-3K · out-of-domain audio-visual sync benchmark
Out-of-domain evaluation set used in the paper
When Vision Speaks for Sound.
Derived from VGGSoundSync,
this 3,000-clip slice tests whether a video-capable MLLM can detect
audio temporal offsets on everyday sound events outside the THUD
in-domain training distribution.
Each item is one VGGSound clip in one of three conditions:
Condition
Count
gt_synced
gt_direction
gt_offset_sec
Audio aligned (no shift)
1,000
true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.StethoBench
StethoBench
StethoBench is a comprehensive benchmark for cardiopulmonary auscultation, comprising 77,027 instruction–response pairs synthesized from 16,125 labeled recordings across 11 public datasets. It is the training and evaluation benchmark for StethoLM, published in the Transactions on Machine Learning Research (TMLR).
Dataset Description
StethoBench was constructed by synthesizing instruction–response pairs from existing labeled cardiopulmonary audio datasets… See the full description on the dataset page: https://huggingface.co/datasets/hamedfrogh/StethoBench.
