Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ShareGPT4Video /ShareGPT4Video ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.imagevisual-question-answering10K<n<100K204 likes10k downloads2y agoHugging Face02buaaplay /SVCBench SVCBench: Streaming Video Counting Benchmark This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline. Project Page: https://buaa-colalab.github.io/SVCBench/ Code: https://github.com/buaa-colalab/SVCBench Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/buaaplay/SVCBench.tabularvideo-classification1K<n<10K10 likes3.4k downloads3mo agoHugging Face03bigai-nlco /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/bigai-nlco/VideoHallucer.textquestion-answering1K<n<10K4 likes2.6k downloads1y agoHugging Face04franky-veteran /SITE-BenchThis dataset contains image and video QA test sets for SITE-Bench evaluation. imagequestion-answering1K<n<10K3 likes2.3k downloads8mo agoHugging Face05michalsr /molmo2-moments Molmo-2 Moments (M2M) Long-video QA dataset where every question is anchored to a specific [start, end] clip interval in seconds. Released alongside the ToolMerge paper, "Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval". ⚠️ Source videos & ownership The videos/*.mp4 files in this repository were collected from YouTube. We do not own these videos and claim no copyright over them. All rights to the video content remain with the original… See the full description on the dataset page: https://huggingface.co/datasets/michalsr/molmo2-moments.tabularvideo-text-to-text10K<n<100K0 likes1.3k downloads2mo agoHugging Face06a8cheng /SR-3D-Bench Spatial Region 3D (SR-3D) Aware Benchmark Paper: https://arxiv.org/abs/2509.13317Project page: https://www.anjiecheng.me/sr3dCode: https://github.com/AnjieCheng/SR-3D [!IMPORTANT] [Feb. 18, 2026] UPDATE: To improve compatibility with general-purpose VLMs, the benchmark is reformulated into multiple-choice and numerical questions following the VSI-Bench evaluation protocol. Videos are annotated with set-of-marks to explicitly indicate regions. The benchmark will be compatible… See the full description on the dataset page: https://huggingface.co/datasets/a8cheng/SR-3D-Bench.textquestion-answering1K<n<10K1 likes1.2k downloads8mo agoHugging Face07jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face08agentvidbench /agentvidbench AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI) Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench.textvideo-text-to-textn<1K0 likes513 downloads2mo agoHugging Face09LIQIIIII /ViMU ViMU: Benchmarking Video Metaphorical Understanding Qi Li, Xinchao Wang* *Corresponding author xML Lab, National University of Singapore Our GitHub repository contains the evaluation scripts for ViMU, a benchmark for video metaphorical understanding. The code evaluates multimodal models on four tasks: Open-ended interpretation (OE) Evidence grounding (EG) Rhetoric mechanism identification (RM) Social value signal identification (SV) Directory Structure Expected… See the full description on the dataset page: https://huggingface.co/datasets/LIQIIIII/ViMU.imagevisual-question-answering1K<n<10K6 likes401 downloads5mo agoHugging Face10iesc /Ava-100 Empowering Agentic Video Analytics Systems with Video Language Models [🖥️ Project Code] [📖 arXiv Paper] [📊 Dataset] Introduction AVA-100 is an ultra-long video benchmark specially designed to evaluate video analysis capabilities Avas-100 consists of 8 videos, each exceeding 10 hours in length, and includes a total of 120 manually annotated questions. The benchmark covers four typical video analytics scenarios: human daily activities, city walking, wildlife… See the full description on the dataset page: https://huggingface.co/datasets/iesc/Ava-100.textmultiple-choicen<1K2 likes358 downloads11mo agoHugging Face11Ustiniansy /SportsTimegated SportsTime SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026. It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball. Dataset This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.textvisual-question-answering10K<n<100K3 likes268 downloads1mo agoHugging Face12ov015 /VideoHallucer VideoHallucer Paper: https://huggingface.co/papers/2406.16338 This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/ov015/VideoHallucer.textquestion-answering1K<n<10K0 likes259 downloads5mo agoHugging Face13QLGalaxy /VUDG VUDG: A Dataset for Video Understanding Domain Generalization VUDG is a benchmark dataset for evaluating domain generalization (DG) in video understanding. It contains 7,899 video clips and 36,388 high-quality QA pairs, covering 11 diverse visual domains, such as cartoon, egocentric, surveillance, rainy, snowy, etc. Each video is annotated with both multiple-choice and open-ended question-answer pairs, designed via a multi-expert progressive annotation pipeline using large… See the full description on the dataset page: https://huggingface.co/datasets/QLGalaxy/VUDG.textquestion-answering10K<n<100K3 likes230 downloads8mo agoHugging Face14inesriahi /valor32k-avqa-v2 Valor32k-AVQA v2.0 Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position. Links Paper: ACM Digital Library Project page: inesriahi.github.io/valor32k-avqa-2 Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.tabularquestion-answering100K<n<1M0 likes215 downloads4mo agoHugging Face15rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes212 downloads1y agoHugging Face16KamiKrafton /agentvidbench AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI) Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.textvideo-text-to-textn<1K0 likes204 downloads2mo agoHugging Face17Inst-IT /Inst-It-Dataset Inst-IT Dataset: An Instruction Tuning Dataset with Multi-level Fine-Grained Annotations introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning 🌐 Homepage | Code | 🤗 Paper | 📖 arXiv Inst-IT Dataset Overview We create a large-scale instruction tuning dataset, the Inst-it Dataset. To the best of our knowledge, this is the first dataset that provides fine-grained annotations centric on specific… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Dataset.textquestion-answering10K<n<100K10 likes189 downloads2y agoHugging Face18n0nam4 /WereBench Anonymization For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information. WereBench WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior… See the full description on the dataset page: https://huggingface.co/datasets/n0nam4/WereBench.tabularquestion-answeringn<1K0 likes189 downloads9mo agoHugging Face19CrossVideoReasoning /SYNCR SYNCR SYNCR is a simulator-grounded framework for cross-video reasoning: questions that cannot be answered from any single video, but require aligning events, matching identities, comparing motion, or integrating partial observations across several videos. Because the videos are produced in simulation, every answer is derived from environment state rather than from human annotation. The same generators produce both an evaluation benchmark and a training set over disjoint videos… See the full description on the dataset page: https://huggingface.co/datasets/CrossVideoReasoning/SYNCR.textvideo-classification10K<n<100K0 likes182 downloads10d agoHugging Face20shuaishuaicdp /OmniCoding OmniCoding A multimodal terminal-tool-use SFT/RL dataset. Each record is a question + verifiable answer + media (video/audio/image) — the target agent is expected to operate on the media via a Linux terminal (ffmpeg, ffprobe, whisper, python, etc.) rather than a GUI. Aggregated and filtered from four upstream sources, with a single unified schema, global dedup, and category-balanced sampling. Records Source n Omnimodal-Agent-SFT-2K (RUC-NLPIR) — agentic… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/OmniCoding.textquestion-answering10K<n<100K0 likes172 downloads2mo agoHugging Face21meituan-longcat /MineExplorer MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft Tianjie Ju · Yueqing Sun · Zheng Wu · Wei Zhang · Yaqi Huo · Xi Su · Qi Gu · Xunliang Cai · Gongshen Liu · Zhuosheng Zhang Abstract Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains unclear. Existing embodied and… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/MineExplorer.textreinforcement-learningn<1K3 likes155 downloads4mo agoHugging Face22Yuan4629 /WereBench Anonymization For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information. WereBench WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior with… See the full description on the dataset page: https://huggingface.co/datasets/Yuan4629/WereBench.tabularquestion-answeringn<1K1 likes142 downloads9mo agoHugging Face23iLearn-Lab /FineBadmintonBenchmark FineBadmintonBenchmark Fine-grained badminton video question answering benchmark. Dataset Structure hf_video_clips_qa/: one clip per QA item, named by video_uid (for example video_000001.mp4). finebadmintonbenchmark/: annotation JSON files. each item contains video_uid each QA item maps to exactly one video clip through video_uid Citation @inproceedings{he2025finebadminton, title={Finebadminton: A multi-level dataset for fine-grained badminton video… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/FineBadmintonBenchmark.textquestion-answering1K<n<10K0 likes98 downloads7mo agoHugging Face24LordUky /EMCompressEMCompress A Benchmark for Endomorphic Multimodal Compression on Long Cooking Videos 📰 News 2026.05   🎉 Dataset + reproduction code released on HuggingFace & GitHub. 2026.04   📝 Paper accepted to ACL 2026 Findings. 🧠 About EMCompress is the first benchmark dedicated to evaluating the Endomorphic Multimodal Compression (EMC) task: an endomorphic transformation F_EMC : (V, Q) → (v, q) that compresses a (video, question) pair into a shorter… See the full description on the dataset page: https://huggingface.co/datasets/LordUky/EMCompress.textvideo-text-to-text1K<n<10K2 likes83 downloads5mo agoHugging Face25shuzhig /elv-halluc-videos ELV-Halluc — videos + annotations A self-contained mirror of the ELV-Halluc benchmark (CVPR 2026), bundling the raw .mp4 files together with the annotations so the benchmark can be run without sourcing videos separately. Paper: arXiv:2508.21496 Original annotations: HLSv/ELV-Halluc (no videos) Project page: https://elv-halluc.github.io/ This is an unofficial mirror. All credit for the benchmark goes to the original authors; please cite their paper (below) rather than this… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/elv-halluc-videos.tabularvideo-text-to-text1K<n<10K0 likes82 downloads2mo agoHugging Face26KamiKrafton /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.textvideo-text-to-textn<1K0 likes82 downloads2mo agoHugging Face27agentvidbench /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench-sample.textvideo-text-to-textn<1K0 likes78 downloads2mo agoHugging Face28AgentVidBench123 /agentvidbench AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) Anonymous authors — under review. Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/ │ └── video*.mp4 # 71 video files └── transcripts/… See the full description on the dataset page: https://huggingface.co/datasets/AgentVidBench123/agentvidbench.textvideo-text-to-textn<1K0 likes67 downloads13d agoHugging Face29UBC-ViL /BlackSwanSuite-MCQgated Black Swan Suite Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events Aditya Chinchure*, Sahithya Ravi*, Raymond Ng, Vered Shwartz, Boyang Li, Leonid Sigal (* equal) 🎉 Accepted at CVPR 2025 arXiv | Website Dataset Information Black Swan has three variants of questions. Please find the data in the appropriate repositories: BlackSwanSuite-Gen -- link BlackSwanSuite-MCQ (this) BlackSwanSuite-YN -- link This dataset contains questions for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/UBC-ViL/BlackSwanSuite-MCQ.tabularvisual-question-answering1K<n<10K2 likes66 downloads2y agoHugging Face30Rakancorle1 /hans-10k Hans-10K · DPO recipe for the audio-visual Clever Hans DPO training data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans 🐎 — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-10K is the 10,383-sample best-recipe preference-pair dataset that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.audioaudio-classification10K<n<100K0 likes62 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.