datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VideoThinkBench
[CVPR 2026] Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
🎊 News
[2026.02] 🔥🔥Our work has been accepted by CVPR 2026! 🎉🎉🎉
[2025.11] Our paper "Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm" has been released on arXiv! 📄 [Paper] On HuggingFace, it has achieved "#1 Paper of the Day"!
[2025.11] 🔥We release "minitest" of our VideoThinkBench, including 500… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/VideoThinkBench.GameQA-140K
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.LearnFromMove
Learn from Move: LatentGUIWorld Benchmark
LatentGUIWorld is the interactive GUI benchmark introduced in
Learn from Move: the Next Step for GUI Agents. It contains 900 test episodes
across six environments, with 150 episodes per environment and a
1280 × 720 viewport.
Code and environment runtime
Environments
Configuration
Task
Episodes
drag_egocentric
Egocentric Drag
150
drag_exocentric
Exocentric Drag
150
rotation_inner
Inner Rotation
150… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/LearnFromMove.GameQA-5KIn this repository, we specifically provide the 5k training samples from the complete GameQA-140K dataset used in our work for GRPO training of the models.
Refer to our paper for details. And our code for training and evaluation is at https://github.com/tongjingqi/Code2Logic.
Code2Logic: Game-Code-Driven Data Synthesis for Enhancing VLMs General Reasoning
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-5K.
