datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen_trajectories_finalSpatialEdit-500K
SpatialEdit-500K
SpatialEdit-500K is a synthetic training dataset for fine-grained image spatial editing. It is built for learning geometry-aware edits such as object moving, object rotation, and camera viewpoint change.
The dataset was introduced in the paper SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing. It is generated with a controllable rendering pipeline to provide structured spatial transformations at scale.
Project Resources
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/EasonXiao-888/SpatialEdit-500K.SpatialCorpus-110MSpatialVID-HQSpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang1*
Yufeng Yuan1*
Rujie Zheng1*
Youtian Lin1
Jian Gao1
Lin-Zhuo Chen1
Yajie Bao1
Yi Zhang1
Chang Zeng1
Yanxi Zhou1
Xiaoxiao Long1
Hao Zhu1
Zhaoxiang Zhang2
Xun Cao1
Yao Yao1†
1Nanjing University 2Institute of Automation, Chinese Academy of Science
*Equal Contribution †Corresponding Author
CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/FelixYuan-YF/SpatialVID-HQ.SpatialVIDSpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang1*
Yufeng Yuan1*
Rujie Zheng1*
Youtian Lin1
Jian Gao1
Lin-Zhuo Chen1
Yajie Bao1
Yi Zhang1
Chang Zeng1
Yanxi Zhou1
Xiaoxiao Long1
Hao Zhu1
Zhaoxiang Zhang2
Xun Cao1
Yao Yao1†
1Nanjing University 2Institute of Automation, Chinese Academy of Science
*Equal Contribution †Corresponding Author
CVPR 2026… See the full description on the dataset page: https://huggingface.co/datasets/SpatialVID/SpatialVID.Awesome_Spatial_VQA_BenchmarksGeoSR-Bench
GeoSR-Bench
Dataset and model weights for the paper:
Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration [arXiv]
The code is available on GitHub: https://github.com/ai-spatial/GeoSR-Bench
Dataset Description
GeoSR-Bench directly connects super-resolution (SR) with downstream Earth monitoring tasks, moving beyond conventional fidelity-based evaluation. It comprises spatially co-located… See the full description on the dataset page: https://huggingface.co/datasets/ai-spatial/GeoSR-Bench.SpatialForge
SpatialForge-10M
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
📑 Paper
Zishan Liu, Ruoxi Zang, Yanglin Zhang, Wei Liu, Yin Zhang, Jian Yao, Jiayin Zheng, Zhengzhe Liu
Lingnan University · XPENG Robotics
📦 SpatialForge-10M
A large-scale vision-language dataset designed for 3D-aware spatial perception and reasoning from open-world 2D images.
SpatialForge-10M contains over 10 million QA pairs generated from 2.8 million curated… See the full description on the dataset page: https://huggingface.co/datasets/shana643/SpatialForge.SpatialRGPT-Benchamara-spatial-10k
AmaraSpatial-10K
A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing
10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines.
Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.SpatialLM-Testset
SpatialLM Testset
Project page | Paper | Code
We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchspatialvlm_qa
Synthetic Spatial Visual Language Question Answering Dataset
Automatically generated from the Blender Scene Dataset. Each example contains an image and a question-answer pair to probe metric (numeric) and relation (true/false) spatial reasoning skills.
Images: Rendered using Blender (1000 scenes, 5 random primitives each, random cameras and lighting).
Metadata: Object name, position, scale, color, material flags.
Questions: 10 per image, drawn from handcrafted templates… See the full description on the dataset page: https://huggingface.co/datasets/Litian2002/spatialvlm_qa.SpatialVID_HD_InP_encodedSpatialLM-Dataset
SpatialLM Dataset
The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.SpatialBlock-15k
SpatialBlock-15k
This dataset accompanies the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem. It contains 15,000 synthetic block-stacking problems for training large vision-language models (LVLMs) to improve spatial reasoning. The dataset includes three types of multiple-choice questions:
Q1: 3D-to-2D projection
Q2: viewpoint transformation
Q3: structural combination
The dataset is organized into a train split of 15,000… See the full description on the dataset page: https://huggingface.co/datasets/rsoohyun/SpatialBlock-15k.Spatial-DISE
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
📋 Overview
Spatial-DISE is a comprehensive benchmark dataset designed to evaluate spatial reasoning capabilities in vision-language models. The dataset focuses on various aspects of spatial intelligence including 3D perception, spatial transformation, and geometric reasoning across multiple difficulty levels.
🧪 Evaluation Support
Supported:… See the full description on the dataset page: https://huggingface.co/datasets/TACPS-liv/Spatial-DISE.Spatial457SpatialEval
🤔 About SpatialEval
SpatialEval is a comprehensive benchmark for evaluating spatial intelligence in LLMs and VLMs across four key dimensions:
Spatial relationships
Positional understanding
Object counting
Navigation
Benchmark Tasks
Spatial-Map: Understanding spatial relationships between objects in map-based scenarios
Maze-Nav: Testing navigation through complex environments
Spatial-Grid: Evaluating spatial reasoning within structured environments
Spatial-Real:… See the full description on the dataset page: https://huggingface.co/datasets/MilaWang/SpatialEval.Q-Spatial-Bench
Dataset Card for Q-Spatial Bench
Q-Spatial Bench is a benchmark designed to measure the quantitative spatial reasoning 📏 in large vision-language models.
🔥The paper associated with Q-Spatial Bench is accepted by EMNLP 2024 main track!
Our paper: Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models [arXiv link]
Project website: [link]
Dataset Details
Q-Spatial Bench is a benchmark designed to measure the… See the full description on the dataset page: https://huggingface.co/datasets/andrewliao11/Q-Spatial-Bench.embodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.map-spatial-benchmark
Map-based Spatial Reasoning Benchmark
A multi-view map-based spatial reasoning benchmark. Each row is one multiple-choice
question instance over a registered map image; models must answer with a single option
letter. Four tasks (T1–T4), four base-map views, and controlled evidence conditions
(direct / query / oracle) and world perturbations (transform / world layers) allow
fine-grained analysis of spatial reasoning robustness.
Task overview
Task
Question… See the full description on the dataset page: https://huggingface.co/datasets/mapspatial/map-spatial-benchmark.thai_spatial_reasoning
Thai Spatial Reasoning 1.0.0
Thai spatial captions and question-answer pairs for continual pretraining of a Thai
foundation VLM. Synthetic means the Thai text and the spatial annotations are generated or
processed; the primary images are real photographs.
Images are not covered by one blanket licence, so the card does not point at one file:
every image has its own licence, creator, and attribution in
rights/attribution.csv and
rights/ledger.parquet. The dataset-level other… See the full description on the dataset page: https://huggingface.co/datasets/SmartWhatt/thai_spatial_reasoning.spatial-mmcot-motif
Spatial MMCoT v1 · motif
MoTiF (OpenRaiser/MoTiF), the naive arm of its procedurally generated multi-step tasks: ball_tracking_naive, manipulation_naive, maze_naive (manipulation shows MoTiF's own CLEVR-style 3D renders of solids, not images from the CLEVR dataset). Each step has its own target image. The reflexion arm, which pairs each problem with a deliberately corrupted frame, is not included. sokoban_naive is excluded (7,997 upstream rows, S0.task_excluded): its eight… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-motif.SpatialVQASpatialMQAWelcome to explore our work titled "Can Multimodal Large Language Models Understand Spatial Relations".arXiv link: https://arxiv.org/abs/2505.19015.For more information about the paper and the SpatialMQA dataset, please visit our GitHub repository at https://github.com/ziyan-xiaoyu/SpatialMQA.
license: cc-by-4.0
grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.spatial-visual-reasoning-66kspatial-mmcot-zebra_multihop
Spatial MMCoT v1 · zebra_multihop
Zebra-CoT multi-hop object counting over rendered 3D scenes (primitive objects on textured ground under a sky). Many traces open with a viewpoint change, so the first target is often a novel view. The final thought re-examines the last target image, but on some rows it only says it counts the objects and never states the number; the count is then only in <answer>. Released rows carry 2 to 5 target images. 2,057 traces with more target images… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-zebra_multihop.SpatialBench
SpatialBench: A Benchmark for Video Spatial Understanding
SpatialBench is a benchmark suite designed to evaluate the video spatial understanding capabilities of Multimodal Large Language Models (MLLMs). This project uses an OpenAI-compatible API interface to send video frames and related spatial reasoning questions to models, automatically evaluating their response accuracy.
Features
Multi-dimensional Evaluation: Covers 5 major categories and 15 sub-categories of… See the full description on the dataset page: https://huggingface.co/datasets/XPR2004/SpatialBench.
