Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Joysw909 /AVQA Summary | 摘要 This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.audioquestion-answering10K<n<100K2 likes3.5k downloads11mo agoHugging Face02MBZUAI /AudioJailbreak Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly. 📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.audioquestion-answering1K<n<10K9 likes2.2k downloads1y agoHugging Face03jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face04holvan /LongAudioSpan LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension Introduction LongAudioSpan is a benchmark for long-form audio comprehension, spanning diverse durations and cognitive depths. Questions come from two complementary paths: Native QA: questions drawn from the audio's natural content. Anchor QA: questions built around acoustic anchors planted into the audio. Each path is scored in its own mode: Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/holvan/LongAudioSpan.textaudio-text-to-text1K<n<10K9 likes957 downloads21d agoHugging Face05ASLP-lab /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.audioaudio-classification1K<n<10K1 likes740 downloads3mo agoHugging Face06QCRI /SpokenNativQA SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs The SpokenNativQA dataset consists of question-answer (QA) pairs, where queries are sourced from real users and answers are manually reviewed and edited. The dataset covers a diverse range of 18 topics that reflect culturally and regionally specific knowledge, as well as everyday queries. These topics include animals, business, clothing, education, events, food and drinks, general knowledge, geography, immigration… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/SpokenNativQA.audioquestion-answering10K<n<100K3 likes479 downloads1y agoHugging Face07AIMS-RAIL /RAIL RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark NeurIPS 2026 Hongyu Jin1,*, Siyi Wang1,*, Yang Xiao1,*, Jiaheng Dong1,*, Shihong Tan4, Kaiyuan Peng1, Georgiana Juravle2, Shanquan Chen3, Gongping Huang4, Hong Jia5, Eun-Jung Holden1, James Bailey6, Ting Dang1,† 1The University of Melbourne, 2Alexandru Ioan Cuza University of Iași, 3The University of Hong Kong, 4Wuhan University, 5The University of Auckland, 6Monash University… See the full description on the dataset page: https://huggingface.co/datasets/AIMS-RAIL/RAIL.audioaudio-classification10K<n<100K6 likes373 downloads6d agoHugging Face08inesriahi /valor32k-avqa-v2 Valor32k-AVQA v2.0 Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position. Links Paper: ACM Digital Library Project page: inesriahi.github.io/valor32k-avqa-2 Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.tabularquestion-answering100K<n<1M0 likes215 downloads4mo agoHugging Face09lysanderism /FTAR TimeAudio: Bridging Temporal Gaps in Large Audio-Language Models Abstract Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perception, leading to constrained capabilities on fine-grained tasks. We identify three key aspects that limit… See the full description on the dataset page: https://huggingface.co/datasets/lysanderism/FTAR.audioaudio-classification100K<n<1M3 likes213 downloads11mo agoHugging Face10rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes212 downloads1y agoHugging Face11Khalilah-Shields /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.audioaudio-classification1K<n<10K0 likes208 downloads3mo agoHugging Face12aseth125 /audio-hallucination-attack Audio Hallucination Attacks (AHA) Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models" It contains two subsets: AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training Audio Files The audio files are provided as compressed archives in this repository: File Contents Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.audioaudio-classification100K<n<1M2 likes175 downloads6mo agoHugging Face13SEAR-benchmark /SEAR SEAR: Spoofing Evidence-Grounded Audio Reasoning SEAR is an audio question-answering benchmark for testing whether audio language models can identify and quantify signal-level acoustic anomalies and use them as evidence for audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR separates deepfake detection, forgery-cue identification, acoustic measurement, and forensic rationale generation. SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.textaudio-classification100K<n<1M1 likes175 downloads19d agoHugging Face14shuaishuaicdp /OmniCoding OmniCoding A multimodal terminal-tool-use SFT/RL dataset. Each record is a question + verifiable answer + media (video/audio/image) — the target agent is expected to operate on the media via a Linux terminal (ffmpeg, ffprobe, whisper, python, etc.) rather than a GUI. Aggregated and filtered from four upstream sources, with a single unified schema, global dedup, and category-balanced sampling. Records Source n Omnimodal-Agent-SFT-2K (RUC-NLPIR) — agentic… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/OmniCoding.textquestion-answering10K<n<100K0 likes172 downloads2mo agoHugging Face15askyishan /StethoBench StethoBench StethoBench is a comprehensive benchmark for cardiopulmonary auscultation, comprising 77,027 instruction–response pairs synthesized from 16,125 labeled recordings across 11 public datasets. It is the training and evaluation benchmark for StethoLM, published in the Transactions on Machine Learning Research (TMLR). Dataset Description StethoBench was constructed by synthesizing instruction–response pairs from existing labeled cardiopulmonary audio datasets… See the full description on the dataset page: https://huggingface.co/datasets/askyishan/StethoBench.textaudio-classification10K<n<100K4 likes113 downloads7mo agoHugging Face16Haopeng /PhoneticQA-SO762 PhoneticQA-SO762 v0.1 PhoneticQA-SO762 is a small, word-level multiple-choice AudioQA benchmark derived from SpeechOcean762. It was created and released by Haopeng Geng as an early benchmark for comparing human and speech-language-model sensitivity to salient mispronunciations. Benchmark design 160 full-utterance audio questions: 32 dev and 128 test. Each item shows the canonical transcript and four candidate words. The task is to select the word that sounds most… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-SO762.audioaudio-classificationn<1K0 likes108 downloads1mo agoHugging Face17umd-zhou-lab /AVQA-Audio-Rubrics AVQA Audio-Reasoning Rubrics Project Page | Paper | Code Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward). Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.textaudio-classification10K<n<100K1 likes104 downloads2mo agoHugging Face18Weisiqing123 /ONOTE ONOTE: Omnimodal Notation Objective Topology Examination ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature. 📂 Dataset Structure The dataset is organized into two primary sub-directories based on the notation and instrument type: 1. pitch_Jianpu_dataset (Staff & Numbered Notation) This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/Weisiqing123/ONOTE.audioimage-to-text1K<n<10K1 likes96 downloads6mo agoHugging Face19sander-wood /voices-of-civilizations Voices of Civilizations (VoC) Voices of Civilizations (VoC) is the first multilingual QA benchmark designed to assess audio LLMs’ cultural comprehension using full-length music recordings. VoC spans: 38 languages 🇸🇦 Arabic (ar), 🇧🇩 Bengali (bn), 🇧🇬 Bulgarian (bg), 🇨🇳 Chinese (zh), 🇭🇷 Croatian (hr), 🇨🇿 Czech (cs), 🇩🇰 Danish (da), 🇳🇱 Dutch (nl), 🇬🇧 English (en), 🇪🇪 Estonian (et), 🇫🇮 Finnish (fi), 🇫🇷 French (fr), 🇩🇪 German (de), 🇬🇷 Greek (el), 🇮🇱 Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/voices-of-civilizations.textquestion-answeringn<1K1 likes71 downloads1y agoHugging Face20triad-26 /TRIAD TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio. Overview TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is needed… See the full description on the dataset page: https://huggingface.co/datasets/triad-26/TRIAD.audiovisual-question-answeringn<1K1 likes68 downloads5mo agoHugging Face21Rakancorle1 /hans-10k Hans-10K · DPO recipe for the audio-visual Clever Hans DPO training data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans 🐎 — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-10K is the 10,383-sample best-recipe preference-pair dataset that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.audioaudio-classification10K<n<100K0 likes62 downloads5mo agoHugging Face22vhands /audio-reasoning-qa-post-public audio-reasoning-qa-post-public Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.textquestion-answering100K<n<1M0 likes62 downloads3mo agoHugging Face23chengxin666 /ONOTE ONOTE: Omnimodal Notation Objective Topology Examination ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature. 📂 Dataset Structure The dataset is organized into two primary sub-directories based on the notation and instrument type: 1. pitch_Jianpu_dataset (Staff & Numbered Notation) This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/chengxin666/ONOTE.audioimage-to-text1K<n<10K0 likes53 downloads4mo agoHugging Face24bnovikov /gemma-4-e4b-audio-qa Gemma-4 E4B Audio-QA Training Mix A 91k-row audio question-answering dataset assembled from four public upstream datasets, formatted as ChatML-style conversations for instruction-tuning an audio-language model. This is the exact training data used for bnovikov/gemma-4-e4b-audio-v3. Important: this repository contains only the metadata and prompts/answers. The audio files are NOT hosted here. Each audio_path is a source-tagged ID like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.textaudio-classification10K<n<100K0 likes49 downloads6mo agoHugging Face25Jashin-Yeah /VoiceGiraffe VoiceGiraffe (Benchmark) VoiceGiraffe is a benchmark for evaluating large audio language models (LALMs) on hour-level, long-context audio understanding. It contains 1,500 curated question-answer triplets over real-world recordings central to real-world long-form audio understanding — broadcast, sports/esports commentary, news, and TV drama — organized into a dual-level taxonomy of single-hop perception and multi-hop reasoning. This repo is public and holds the annotations… See the full description on the dataset page: https://huggingface.co/datasets/Jashin-Yeah/VoiceGiraffe.textaudio-classification1K<n<10K0 likes47 downloads2mo agoHugging Face26audiospan /LongAudioSpan LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension Introduction LongAudioSpan is a benchmark for long-form audio comprehension, spanning diverse durations and cognitive depths. Questions come from two complementary paths: Native QA: questions drawn from the audio's natural content. Anchor QA: questions built around acoustic anchors planted into the audio. Each path is scored in its own mode: Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/audiospan/LongAudioSpan.textaudio-text-to-text1K<n<10K0 likes41 downloads20d agoHugging Face27Haopeng /PhoneticQA-L2ARCTIC PhoneticQA-L2ARCTIC v0.1 PhoneticQA-L2ARCTIC is a small, word-level multiple-choice AudioQA probe derived from L2-ARCTIC. It was created and released by Haopeng Geng to study the gap between phoneme-level error annotations and perceptually salient mispronunciations. Benchmark design 120 full-utterance audio questions: 24 dev and 96 test. Each item shows the canonical transcript and four candidate words. The task is to select the word that sounds most clearly… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-L2ARCTIC.audioaudio-classificationn<1K0 likes35 downloads1mo agoHugging Face28Rakancorle1 /hans-sft-4k Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans Supervised fine-tuning (SFT) data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-SFT-4K is the 3,834-sample SFT mix that teaches models to… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.audioaudio-classification1K<n<10K1 likes31 downloads5mo agoHugging Face29Rakancorle1 /vggsync-3k VGGSync-3K · out-of-domain audio-visual sync benchmark Out-of-domain evaluation set used in the paper When Vision Speaks for Sound. Derived from VGGSoundSync, this 3,000-clip slice tests whether a video-capable MLLM can detect audio temporal offsets on everyday sound events outside the THUD in-domain training distribution. Each item is one VGGSound clip in one of three conditions: Condition Count gt_synced gt_direction gt_offset_sec Audio aligned (no shift) 1,000 true… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/vggsync-3k.audioaudio-classification1K<n<10K0 likes29 downloads5mo agoHugging Face30hamedfrogh /StethoBench StethoBench StethoBench is a comprehensive benchmark for cardiopulmonary auscultation, comprising 77,027 instruction–response pairs synthesized from 16,125 labeled recordings across 11 public datasets. It is the training and evaluation benchmark for StethoLM, published in the Transactions on Machine Learning Research (TMLR). Dataset Description StethoBench was constructed by synthesizing instruction–response pairs from existing labeled cardiopulmonary audio datasets… See the full description on the dataset page: https://huggingface.co/datasets/hamedfrogh/StethoBench.textaudio-classification10K<n<100K0 likes24 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.