AgentVidBench123/agentvidbench
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) Anonymous authors — under review. Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/ │ └── video*.mp4 # 71 video files └── transcripts/… See the full description on the dataset page: https://huggingface.co/datasets/AgentVidBench123/agentvidbench.
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline)
Anonymous authors — under review.
<div align="center"> <img src="assets/example_q56.png" alt="Q56 — Bicep Curls Before "One More" (example task with human-curated reasoning trajectory)" width="75%"> </div>
Layout
.
├── README.md
├── questions.jsonl # 100 rows — one per question
├── videos.jsonl # 71 rows — one per unique video
├── videos/
│ └── video*.mp4 # 71 video files
└── transcripts/
└── video*.srt # 71 Whisper transcripts (28 are empty for silent videos)questions.jsonl schema
videos.jsonl schema
Download
hf download AgentVidBench123/agentvidbench --repo-type dataset --local-dir datasetCitation
@misc{anonymous2026agentvidbench,
title = {AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents},
author = {Anonymous},
year = {2026},
note = {Under review}
}