wangyz1999/GameplayQA
GameplayQA: A Decision-Dense POV-Synced Multi-Video Understanding Benchmark of 3D Virtual Agents Yunzhe Wang Runhui Xu Kexin Zheng Tianyi Zhang Jayavibhav N. Kogundi Soham Hans Volkan Ustun University of Southern California ACL 2026 Corresponding Author: yunzhewa@usc.edu Overview GameplayQA is the first benchmark for POV-Synced Multi-Video… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.
<!-- MARK: Title --> <h1 align="center"> GameplayQA: A Decision-Dense POV-Synced Multi-Video<br> Understanding Benchmark of 3D Virtual Agents </h1>
<!-- MARK: Authors --> <div align="center"> <a href="https://www.linkedin.com/in/yunzhe-wang/">Yunzhe Wang</a>   <a href="https://www.linkedin.com/in/runhui-xu-a9895a2b3/">Runhui Xu</a>   <a href="https://www.linkedin.com/in/kexin-zheng-8b5a7a232/">Kexin Zheng</a>   <a href="https://www.linkedin.com/in/tianyi-mavis-zhang-3b9514128/">Tianyi Zhang</a> </div> <div align="center"> <a href="https://www.linkedin.com/in/jayavibhav/">Jayavibhav N. Kogundi</a>   <a href="https://www.linkedin.com/in/soham-hans-b6295215a/"> Soham Hans</a>   <a href="https://www.linkedin.com/in/volkan-ustun-9883a74/">Volkan Ustun</a> </div>
<!-- MARK: Affiliation --> <div align="center"> University of Southern California </div> <div align="center"> <b>ACL 2026</b> </div> <div align="center"> <b>Corresponding Author</b>: yunzhewa@usc.edu </div>
<!-- MARK: Badges --> <br> <div align="center"> <a href="https://hats-ict.github.io/gameplayqa/"><img src="https://img.shields.io/static/v1?label=GameplayQA%20Project%20Homepage&message=Website&color=9a33fc&logo=githubpages" style="height: 25px;"></a> <a href="https://arxiv.org/abs/2603.24329"><img src="https://img.shields.io/static/v1?label=Paper&message=arXiv&color=FF0066&logo=arxiv" style="height: 25px;"></a> <a href="https://huggingface.co/datasets/wangyz1999/GameplayQA"><img src="https://img.shields.io/static/v1?label=Dataset&message=HuggingFace&color=FF6600&logo=huggingface" style="height: 25px;"></a> <a href="https://github.com/HATS-ICT/GameplayQA"><img src="https://img.shields.io/static/v1?label=Benchmark%20Code&message=Github&color=181717&logo=github" style="height: 25px;"></a> <br> <a href="https://github.com/wangyz1999/sync-video-label"><img src="https://img.shields.io/static/v1?label=Annotation%20Tool&message=Github&color=6699FF&logo=github" style="height: 25px;"></a> <a href="https://sync-video-label.vercel.app/"><img src="https://img.shields.io/static/v1?label=Annotation%20Tool&message=Live%20Demo&color=33CCCC&logo=vercel" style="height: 25px;"></a> <a href="https://www.youtube.com/watch?v=PKedELJ4XT0"><img src="https://img.shields.io/static/v1?label=Annotation%20Tool%20Demo&message=YouTube&color=FF0000&logo=youtube" style="height: 25px;"></a> <a href="https://huggingface.co/datasets/wangyz1999/X-EGO-CS"><img src="https://img.shields.io/static/v1?label=Related&message=X-EGO-CS&color=FFCC00&logo=huggingface" style="height: 25px;"></a> </div>
<!-- MARK: Overview -->
Overview
GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and Multi-Agent Video Understanding, built from ego-centric gameplay footage across 9 commercial 3D games. It features 2.4K multiple-choice questions spanning three levels of reasoning complexity — from single-clip action recognition to synchronized cross-video understanding.
Why synchronized multi-viewpoint reasoning matters
The ability to reason across multiple synchronized viewpoints is critical in many real-world domains: sports analytics leveraging multiple camera angles, autonomous driving requiring sensor fusion from surround cameras, law enforcement reviewing multiple dashcam feeds, and coordinated robot or drone fleets operating in shared environments. In esports and gaming, cross-POV synchronization and collective reasoning are fundamental to interpreting multi-agent collaboration — making gameplay an ideal controlled testbed for developing and evaluating these capabilities in video-language models.
<video controls> <source src="https://hats-ict.github.io/projects/gameplayqa/static/videos/synced_pov.mp4" type="video/mp4"> </video>
<details> <summary><b>Abstract</b></summary> <br> Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct entities, and reason about concurrent multi-agent behaviors from a first-person perspective — capabilities that existing benchmarks do not adequately evaluate. We introduce <b>GameplayQA</b>, a framework for evaluating agentic-centric perception and reasoning through video understanding. Specifically, we densely annotate multiplayer 3D gameplay videos at 1.22 labels/second, with time-synced, concurrent captions of states, actions, and events structured around a triadic system of <b>Self</b>, <b>Other Agents</b>, and the <b>World</b> — a natural decomposition for multi-agent environments. From these annotations, we generate 2.4K diagnostic QA pairs organized into three levels of cognitive complexity, accompanied by a structured distractor taxonomy that enables fine-grained analysis of where models hallucinate. Evaluation of frontier MLLMs reveals a substantial gap from human performance, with common failures in temporal and cross-video grounding, agent-role attribution, and handling the decision density of the game. We hope GameplayQA stimulates future research at the intersection of embodied AI, agentic perception, and world modeling. </details>
<!-- MARK: Statistics --> <div align="center"> <table> <tr> <td align="center"><b>9</b><br>Games</td> <td align="center"><b>2.4 K</b><br>QA Pairs</td> <td align="center"><b>2,709</b><br>Annotations</td> </tr> <tr> <td align="center"><b>2,219 s</b><br>Annotated Footage</td> <td align="center"><b>1.22 /s</b><br>Decision Density</td> <td align="center"><b>15</b><br>Task Categories</td> </tr> </table> </div>
<!-- MARK: Dataset Files -->
Dataset Files
Raw annotation files and video clips are located in the annotation/ folder, organized by project batch (e.g. annotation/project-batch1/, annotation/project-multi-batch1/). Processed benchmark files ready for evaluation are provided as CSV files at the root level.
Download/Loading the Dataset
The QA metadata (CSV files) can be loaded directly via the HuggingFace datasets library:
from datasets import load_dataset
# Main benchmark (default)
ds = load_dataset("wangyz1999/GameplayQA")
# Generalization benchmark
ds = load_dataset("wangyz1999/GameplayQA", "generalization")The raw video clips are stored under annotation/ and tracked with Git LFS. To download them, clone the full repository:
git lfs install
git clone https://huggingface.co/datasets/wangyz1999/GameplayQAEach QA row contains video_start / video_end timestamps in seconds. Evaluation typically requires cropping the referenced clip to the relevant segment before passing it to a model.
<!-- MARK: Column Reference --> <details> <summary><b>Column Reference</b></summary>
</details>
<!-- MARK: Taxonomy -->
Taxonomy
Entity Types: Self — Other — World
GameplayQA organizes perception in interactive 3D environments around a tripartite entity decomposition that mirrors how agents must reason in multi-agent settings.
<div align="center"> <table> <tr> <td align="center"><img src="sow.jpg" alt="Self-Other-World framework" width="400"/></td> <td align="center"><img src="taxonomy.jpg" alt="Question taxonomy" width="556"/></td> </tr> </table> </div>
Cognitive Levels
Questions are organized by the amount of temporal and cross-video context required:
Task Categories
Questions are organized into 15 task categories across 3 context scope levels (2,365 total questions).
<!-- MARK: Source Data -->
Source Data
GameplayQA is built from first-person gameplay footage sourced from 9 commercially released multiplayer games spanning diverse genres:
- Single-POV games: Minecraft, Apex Legends, No Man's Sky, Elden Ring, Cyberpunk 2077, Valheim
- Multi-POV synchronized games: Counter-Strike 2, Battlefield 6, ARC Raiders
Videos were sourced from YouTube, Twitch streams, and existing datasets (Counter-Strike 2 footage from X-EGO-CS). For multi-POV games, groups of streamers who played together in the same match were identified and their individual recordings were manually time-aligned to construct temporally synchronized multi-video sets.
<!-- MARK: Question Generation -->
Question Generation
Questions are generated through a combinatorial template-based algorithm that systematically combines verified annotation labels across five orthogonal dimensions: number of videos (single/multi), context target (summative/timestamp/entity/cross-video referring), entity type (SA/SS/OA/OS/WO/WE), distractor type, and question form. The algorithm initially produces ~400K candidate QA pairs; strategic downsampling to 4K enforces balanced category coverage before quality assurance yields the final 2,365 gold-standard pairs.
<details> <summary><b>Distractor Types</b></summary>
</details>
<!-- MARK: Quality Assurance -->
Quality Assurance
Language prior filtering: Each question is queried with text only (no video) using Gemini Flash with k=3 trials. Questions where the model consistently selects the correct answer without visual grounding are removed to prevent exploitation of statistical regularities in question phrasing.
Human evaluation: A stratified sample of 120 questions covering all question types was reviewed by annotators, who verified that (1) the video contains exactly one unambiguous correct answer, and (2) the question adheres to the semantics of its question code. Questions flagged as faulty (~8%) were corrected or removed.
<!-- MARK: Annotations -->
Annotations
Annotation Process
Videos were annotated using dense multi-track timeline captioning via a custom-built annotation tool — [sync-video-label](https://github.com/wangyz1999/sync-video-label). Each of the six entity types (SA, SS, OA, OS, WO, WE) is treated as an independent annotation track, and labels within and across tracks can overlap temporally to capture concurrent events.
The process follows a two-stage human-in-the-loop workflow:
- Stage 1 — AI-assisted generation + human verification: Gemini Pro generates candidate labels and distractors (3,632 predictions). Four graduate student annotators then verify and refine: 31.1% of predicted labels were deleted, 42.7% were edited (61.9% requiring caption changes, 42.2% requiring temporal boundary adjustments), and 7.6% of final labels were added manually.
- Stage 2 — Independent review: A separate annotator reviews all labels, making further adjustments to ~12% of labels.
A live read-only demo of the annotation interface is available at sync-video-label.vercel.app. See the demo video for a walkthrough.
Annotators
The annotation team consisted of 5 graduate students (ages 21–31). All annotators were experienced gamers: 60% play 3–5 times per week, 60% have 8+ years of gaming experience. Roles were distributed as 4 labelers and 2 evaluators, with one participant serving in both capacities for cross-stage consistency.
Label Statistics
A total of 2,709 true labels were annotated across 2,219 seconds of footage, yielding a decision density of ρ ≈ 1.22 labels/second — roughly one decision-relevant event per second.
<!-- MARK: Related Datasets -->
Related Datasets
- [wangyz1999/X-EGO-CS](https://huggingface.co/datasets/wangyz1999/X-EGO-CS) — Cross-ego synchronized gameplay data from Counter-Strike 2, used as the source for the
CounterStrike2split in this benchmark.
<!-- MARK: Citation -->
Citation
If our research is helpful to you, please cite our paper:
@article{wang2026gameplayqa,
title = {GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents},
author = {Wang, Yunzhe and Xu, Runhui and Zheng, Kexin and Zhang, Tianyi and Kogundi, Jayavibhav Niranjan and Hans, Soham and Ustun, Volkan},
year = {2026},
eprint = {2603.24329},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2603.24329}
}
