datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpatialForge
SpatialForge-10M
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
📑 Paper
Zishan Liu, Ruoxi Zang, Yanglin Zhang, Wei Liu, Yin Zhang, Jian Yao, Jiayin Zheng, Zhengzhe Liu
Lingnan University · XPENG Robotics
📦 SpatialForge-10M
A large-scale vision-language dataset designed for 3D-aware spatial perception and reasoning from open-world 2D images.
SpatialForge-10M contains over 10 million QA pairs generated from 2.8 million curated… See the full description on the dataset page: https://huggingface.co/datasets/shana643/SpatialForge.Spatial-SSRL-81k
Spatial-SSRL-81k
📖Paper| 🏠Github |🤗Spatial-SSRL-7B Model |
🤗Spatial-SSRL-3B Model | 🤗Spatial-SSRL-Qwen3VL-4B Model |
🤗Spatial-SSRL-81k Dataset | 📰Daily Paper
Spatial-SSRL-81k is a training dataset for enhancing spatial understanding in large vision-language models. It contains 81,053 samples of five pretext tasks for self-supervised learning, offering simple, intrinsic supervision that scales RLVR efficiently.
📢 News
🚀 [2026/04/05] We have released… See the full description on the dataset page: https://huggingface.co/datasets/internlm/Spatial-SSRL-81k.Q-Spatial-Bench
Dataset Card for Q-Spatial Bench
Q-Spatial Bench is a benchmark designed to measure the quantitative spatial reasoning 📏 in large vision-language models.
🔥The paper associated with Q-Spatial Bench is accepted by EMNLP 2024 main track!
Our paper: Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models [arXiv link]
Project website: [link]
Dataset Details
Q-Spatial Bench is a benchmark designed to measure the… See the full description on the dataset page: https://huggingface.co/datasets/andrewliao11/Q-Spatial-Bench.SpatialBench
SpatialBench: A Benchmark for Video Spatial Understanding
SpatialBench is a benchmark suite designed to evaluate the video spatial understanding capabilities of Multimodal Large Language Models (MLLMs). This project uses an OpenAI-compatible API interface to send video frames and related spatial reasoning questions to models, automatically evaluating their response accuracy.
Features
Multi-dimensional Evaluation: Covers 5 major categories and 15 sub-categories of… See the full description on the dataset page: https://huggingface.co/datasets/XPR2004/SpatialBench.SpatialLadder-26k
SpatialLadder-26k
This repository contains the SpatialLadder-26k, introduced in SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models.
Dataset Description
SpatialLadder-26k is a large-scale training dataset designed to develop spatial perception and reasoning capabilities in Vision-Language Models (VLMs). It contains 26,610 multimodal samples spanning four complementary task categories, forming a… See the full description on the dataset page: https://huggingface.co/datasets/hongxingli/SpatialLadder-26k.LLaVA-Spatial-Instruct-850K
LLaVA Spatial Instruct 850K Dataset Card
Dataset type:
LLaVA Spatial Instruct 850K is a combined set of LLaVA-v1.5 instruction tuning mixture dataset (llava_v1_5_mix665k.json), common benchmark training datasets, including Clevr, Textcaps, Visualmrc, VQAv2 fetched from the_cauldron, and spatial relation dataset including OpenSpaces and SpatialQA dataset created from the data pipeline of SpatialRGPT on OpenImages dataset.
Dataset proportion
LLaVA-v1.5 instruction tuning mixture… See the full description on the dataset page: https://huggingface.co/datasets/rogerxi/LLaVA-Spatial-Instruct-850K.SpatiaLQA
SpatiaLQA
Dataset Summary
SpatiaLQA is a benchmark for evaluating spatial logical reasoning in vision-language models.It focuses on reasoning abilities in real-world indoor environments, including spatial relations among objects and logical dependencies in multi-step tasks.
This dataset is introduced in the paper:
SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language ModelsAuthors: Yuechen Xie, Xiaoyan Zhang, Yicheng Shan, Hao Zhu, Rui Tang… See the full description on the dataset page: https://huggingface.co/datasets/xyc99/SpatiaLQA.SpatiaLab
SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?
Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik, Munem Shahriar, Mohsin Mahmud Topu, Sadia Tasnim Meem, Rahatun Nesa Priti, Sabrina Afroz Mitu, Md. Iqramul Hoque, Shahriyar Zaman Ridoy, Mohammed Eunus Ali, Majd Hawasly, Mohammad Raza, Md Rizwan Parvez
Computational Intelligence and Operations Laboratory (CIOL) • Shahjalal University of Science and Technology (SUST) • Monash… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/SpatiaLab.this-that-spatial-bench
spatial-decisions
7,305 multiple-choice decision questions over 6,525 distinct simulated states, in
15 families and two environments. Every answer is computed from the simulator, not
annotated by a person and not taken from a model. That is the point of the set: on a question whose
answer is derived from the rules of the environment, a disagreement is a mistake, and there is
nothing to argue about.
The set was built to replace a much narrower public artefact: a recording of 68… See the full description on the dataset page: https://huggingface.co/datasets/limberc/this-that-spatial-bench.2d_3d_seq_path_spatial_reasoning
Spatial Reasoning Dataset
A synthetic dataset of Hamiltonian path puzzles with rich chain-of-thought reasoning, designed for training and evaluating spatial reasoning in language models.
Overview
Each sample presents a grid-based puzzle where the solver must find a path visiting every cell exactly once, moving only up/down/left/right (plus above/below for 3D). Puzzles span 2D grids (3x3 to 8x8) and 3D cubes (3x3x3 to 4x4x4), covering solvable, impossible, and multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/eousphoros/2d_3d_seq_path_spatial_reasoning.SpatialReasoning
The Spatial Reasoning Dataset
The Spatial Reasoning Dataset comprises semantically meaningful question-answer pairs focused on the relative locations of geographic divisions within the United States — including states, counties, and ZIP codes.
The dataset is designed for spatial question answering and includes three types of questions:
Binary (Yes/No)
Single-choice (Radio)
Multi-choice (Checkbox)
All questions include correct answers for training and evaluation purposes.… See the full description on the dataset page: https://huggingface.co/datasets/Rammen/SpatialReasoning.Spatial-SSRL-81k
Spatial-SSRL-81k
📖Paper| 🏠Github |🤗Spatial-SSRL-7B Model |
🤗Spatial-SSRL-3B Model | 🤗Spatial-SSRL-Qwen3VL-4B Model |
🤗Spatial-SSRL-81k Dataset | 📰Daily Paper
Spatial-SSRL-81k is a training dataset for enhancing spatial understanding in large vision-language models. It contains 81,053 samples of five pretext tasks for self-supervised learning, offering simple, intrinsic supervision that scales RLVR efficiently.
📢 News
🚀 [2026/04/05] We have released… See the full description on the dataset page: https://huggingface.co/datasets/cvis-tmu/Spatial-SSRL-81k.SpatialMemoryspatial-trace-dataset
Dataset Summary
SpatialTraceGen is a dataset of multi-hop spatial reasoning traces generated by Large Language Models (LLMs) integrated with computer vision tools. The framework is designed to produce step-by-step reasoning for complex spatial queries. This dataset contains the generated reasoning traces under different levels of automated verification.
The dataset was created using the CLEVR dataset as a base. The traces were generated by providing questions from CLEVR to an LLM… See the full description on the dataset page: https://huggingface.co/datasets/dhruvmsheth/spatial-trace-dataset.vantage-spatial-pov-mix
👁️ PyIntel Vantage: Spatial & Perspective Mixture (vantage-spatial-pov-mix)
PyIntel Vantage Spatial-POV Mix is a curated, multimodal-ready instruction dataset designed to teach compact edge language and vision models (specifically the Google Gemma family, including Gemma 4) advanced 3D spatial reasoning, camera perspective translations, and multi-agent Theory of Mind (POV of others).
Curated and published by PyIntel Research.
📊 Dataset Overview
Total Samples:… See the full description on the dataset page: https://huggingface.co/datasets/pyintel/vantage-spatial-pov-mix.SpatialMed
SpatialMed: Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
Quoc-Huy Trinh, Xi Ding, Yang Liu, Zhenyue Qin, Xingjian Li, Gorkem Durak, Halil Ertugrul Aktas, Elif Keles, Ulas Bagci, Min Xu
Abstract
Current evaluations of medical multimodal large language models (MLLMs) focus on diagnostic accuracy but overlook spatial intelligence -- a critical capability for interpreting 3D medical images. We address this gap by developing… See the full description on the dataset page: https://huggingface.co/datasets/huyquoctrinh/SpatialMed.arabic-spatial-reasoning-v1
