Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /svq Simple Voice Questions Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions. It serves as a core evaluation componenet for Massive Sound Embedding Benchmark (MSEB). Technical Specifications Feature Details Locales 26 Languages 17 Total Speakers ~700 (Capped at 250 recordings per speaker) Audio Conditions Clean, Background Speech, Media, Traffic Noise Gender… See the full description on the dataset page: https://huggingface.co/datasets/google/svq.audioquestion-answering1M<n<10M62 likes57k downloads11d agoHugging Face02qyang1021 /AIR-Bench-Dataset AIR-Bench Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists of 19 tasks with approximately 19k single-choice questions. The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon). Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.audioquestion-answeringn<1K8 likes11k downloads2y agoHugging Face03ZJTustc /SciTS SciTS: Scientific Time Series Understanding and Generation with LLMs This repository contains the official dataset for SciTS: Scientific Time Series Understanding and Generation with LLMs (ICLR 2026). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances. Dataset Structure The benchmark is organized… See the full description on the dataset page: https://huggingface.co/datasets/ZJTustc/SciTS.audiotime-series-forecasting10K<n<100K0 likes5.7k downloads6mo agoHugging Face04HumanBehaviorAtlas /human_behavior_atlas Human Behavior Atlas A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features. This dataset was used to train OmniSapiens, a foundation model for social behavior processing. Papers: Human Behavior Atlas: Benchmarking Unified Psychological and… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.textvideo-classification100K<n<1M3 likes5.6k downloads4mo agoHugging Face05RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes5k downloads7mo agoHugging Face06yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M53 likes4.8k downloads7mo agoHugging Face07OpenTSLab /SciTS SciTS: Scientific Time Series Understanding and Generation with LLMs This repository contains the official dataset for SciTS: Scientific Time Series Understanding and Generation with LLMs (ICLR 2026). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances. Dataset Structure The benchmark is organized… See the full description on the dataset page: https://huggingface.co/datasets/OpenTSLab/SciTS.audiotime-series-forecasting10K<n<100K5 likes3.9k downloads7mo agoHugging Face08plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes3.8k downloads10mo agoHugging Face09ddwang2000 /MMSU [ICLR 2026] MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark Overview of MMSU MMSU (Massive Multi-task Spoken Language Understanding and Reasoning Benchmark) is a comprehensive benchmark for evaluating fine-grained spoken language understanding and reasoning in multimodal models. It systematically captures the variance of real-world linguistic phenomena in daily speech through 47 sub-tasks, including phonetics, prosody, rhetoric… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/MMSU.audioquestion-answering1K<n<10K16 likes3.7k downloads5mo agoHugging Face10Joysw909 /AVQA Summary | 摘要 This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.audioquestion-answering10K<n<100K2 likes3.5k downloads11mo agoHugging Face11VocalNet /VocalBench VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models This is the official release of VocalBench Citation If you find our work helpful, please cite our paper: @article{liu2025vocalbench, title={VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models}, author={Liu, Heyang and Wang, Yuhao and Cheng, Ziyang and Wu, Ronghua and Gu, Qunshan and Wang, Yanfeng and Wang, Yu}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/VocalNet/VocalBench.audioquestion-answering1K<n<10K1 likes3.2k downloads9mo agoHugging Face12Hezep /AudioMarathon 🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audioaudio-classification1K<n<10K4 likes3.2k downloads11mo agoHugging Face13zlinao /WearVox WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables Paper: WearVox: An Egocentric Multichannel Voice Assistant Benchmark for WearablesAuthors: Zhaojiang Lin*, Yong Xu*, Kai Sun*, Jing Zheng, Yin Huang, Surya Appini, Krish Narang, Renjie Tao, Ishan Kapil Jain, Siddhant Arora, Ruizhi Li, Yiteng Huang, Kaushik Patnaik, Wenfang Xu, Suwon Shon, Yue Liu, Ahmed Aly, Anuj Kumar, Florian Metze, Luna DongAffiliations: Meta Reality Labs, Meta 📝 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zlinao/WearVox.audioquestion-answering1K<n<10K5 likes3k downloads9mo agoHugging Face14WueNLP /belebele-fleurs Belebele-Fleurs Belebele-Fleurs is a dataset suitable to evaluate two core tasks: Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form. Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.audioaudio-classification10K<n<100K9 likes2.8k downloads2y agoHugging Face15NJU-LINK /OmniVideoBenchgated OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs ✨ Overview Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction. 🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.texttext-generation1K<n<10K5 likes2.6k downloads6mo agoHugging Face16facebook /2M-Belebele 2M-Belebele Highly-Multilingual Speech and American Sign Language Comprehension Dataset We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL). The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.tabularquestion-answering10K<n<100K13 likes2.3k downloads2y agoHugging Face17MBZUAI /AudioJailbreak Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly. 📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.audioquestion-answering1K<n<10K9 likes2.2k downloads1y agoHugging Face18WueNLP /sib-fleurs SIB-Fleurs SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to. The topics are: Science/Technology Travel Politics Sports Health Entertainment Geography Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*. Dataset creation This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.audioaudio-classification10K<n<100K15 likes2.2k downloads1y agoHugging Face19Kkryptonite /CUE-Mem 🧠 CUE-Mem Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations 多模态对话隐式线索驱动的长期用户记忆评测基准 Text · Image · Audio | Explicit & Implicit Cues | Long-Term User Memory English · 中文说明 · Quick Start / 快速开始 · Citation / 引用 From visible routines to subtle clues: a cat bowl, a cat tree, and a background meow jointly suggest that the user has a cat. 从日常活动到隐式线索:猫碗、猫爬架与背景中的猫叫声,共同指向用户养猫这一信息。 English Overview What… See the full description on the dataset page: https://huggingface.co/datasets/Kkryptonite/CUE-Mem.audioquestion-answering1K<n<10K0 likes1.8k downloads6d agoHugging Face20FBK-MT /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.audioautomatic-speech-recognition1K<n<10K70 likes1.6k downloads3mo agoHugging Face21rajjanardhan00 /Seamless_Dummy_Dataset_Fixed MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes1.3k downloads1y agoHugging Face22yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.3k downloads7mo agoHugging Face23DennisDengHUst /human_behavior_atlas Human Behavior Atlas A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features. This dataset was used to train OmniSapiens, a foundation model for social behavior processing. Papers: Human Behavior Atlas: Benchmarking Unified Psychological… See the full description on the dataset page: https://huggingface.co/datasets/DennisDengHUst/human_behavior_atlas.textvideo-classification100K<n<1M0 likes1.2k downloads21d agoHugging Face24Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes1.2k downloads5mo agoHugging Face25jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face26RUC-NLPIR /OmniGAIA OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is designed to evaluate long-horizon, multi-hop, open-form problem solving in realistic settings rather than short perception-only QA. Benchmark Construction The OmniGAIA construction… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/OmniGAIA.audioquestion-answeringn<1K6 likes961 downloads7mo agoHugging Face27holvan /LongAudioSpan LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension Introduction LongAudioSpan is a benchmark for long-form audio comprehension, spanning diverse durations and cognitive depths. Questions come from two complementary paths: Native QA: questions drawn from the audio's natural content. Anchor QA: questions built around acoustic anchors planted into the audio. Each path is scored in its own mode: Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/holvan/LongAudioSpan.textaudio-text-to-text1K<n<10K9 likes957 downloads20d agoHugging Face28vector-institute /sonic-o1 &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding 🎯 What is SONIC-O1? The first open-source benchmark for evaluating omnimodal video understanding with systematic fairness analysis. SONIC-O1 requires models to jointly process audio, video, and social context from real-world interactions—not just transcripts. Key… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/sonic-o1.audiovisual-question-answering1K<n<10K5 likes954 downloads4mo agoHugging Face29CentificAIResearch /MedMosaic MedMosaic Dataset A comprehensive medical audio question-answering dataset designed for evaluating audio understanding models in clinical and healthcare contexts. Dataset Description This dataset contains audio recordings paired with clinical questions and answers across multiple QA types. It is designed to benchmark audio-language models on medical reasoning tasks. Dataset Structure The dataset is organized into 7 subfolders, each representing a… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/MedMosaic.audioaudio-classification1K<n<10K4 likes871 downloads4mo agoHugging Face30yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes837 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.