datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.map-spatial-benchmark
Map-based Spatial Reasoning Benchmark
A multi-view map-based spatial reasoning benchmark. Each row is one multiple-choice
question instance over a registered map image; models must answer with a single option
letter. Four tasks (T1–T4), four base-map views, and controlled evidence conditions
(direct / query / oracle) and world perturbations (transform / world layers) allow
fine-grained analysis of spatial reasoning robustness.
Task overview
Task
Question… See the full description on the dataset page: https://huggingface.co/datasets/mapspatial/map-spatial-benchmark.SpatialMQAWelcome to explore our work titled "Can Multimodal Large Language Models Understand Spatial Relations".arXiv link: https://arxiv.org/abs/2505.19015.For more information about the paper and the SpatialMQA dataset, please visit our GitHub repository at https://github.com/ziyan-xiaoyu/SpatialMQA.
license: cc-by-4.0
grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.SpatialLadder-26k
SpatialLadder-26k
This repository contains the SpatialLadder-26k, introduced in SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models.
Dataset Description
SpatialLadder-26k is a large-scale training dataset designed to develop spatial perception and reasoning capabilities in Vision-Language Models (VLMs). It contains 26,610 multimodal samples spanning four complementary task categories, forming a… See the full description on the dataset page: https://huggingface.co/datasets/hongxingli/SpatialLadder-26k.SpatialCLI-Data
SpatialCLI-Bench
Viewer data: data/eval/SpatialCLI-Bench/full.jsonl
Split: test
Number of records: 516
Shared visual assets: data/assets/
Paths embedded in the records are relative to the repository root.
Paper
This dataset accompanies the SpatialCLI paper.
Code
The code and model checkpoints are available at SpatialCLI GitHub repository.
SpatialGen-Bench
Benchmark
Each record retains its source-task metric target as integer, text, point, mask, or polyline. The frozen visual-answer contract is available at protocols/spatialgen_bench.yaml, with runtime parsers in ProVisE.
Quick Start
from datasets import load_dataset
dataset = load_dataset("wx91726/SpatialGen-Bench", split="test")
print(dataset[0])
Download the complete media and evaluation package for local evaluation:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/wx91726/SpatialGen-Bench.Spatial-Awareness-DatasetSpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialGen-Testset.Spatial-SSRL-81k
Spatial-SSRL-81k
📖Paper| 🏠Github |🤗Spatial-SSRL-7B Model |
🤗Spatial-SSRL-3B Model | 🤗Spatial-SSRL-Qwen3VL-4B Model |
🤗Spatial-SSRL-81k Dataset | 📰Daily Paper
Spatial-SSRL-81k is a training dataset for enhancing spatial understanding in large vision-language models. It contains 81,053 samples of five pretext tasks for self-supervised learning, offering simple, intrinsic supervision that scales RLVR efficiently.
📢 News
🚀 [2026/04/05] We have released… See the full description on the dataset page: https://huggingface.co/datasets/cvis-tmu/Spatial-SSRL-81k.Spatialworld-bench
SpatialWorld Benchmark
A Multi-Platform Benchmark for Spatial Reasoning and Spatial Task Execution
🎯 Overview
SpatialWorld is a comprehensive benchmark designed to evaluate spatial reasoning and spatial task execution capabilities of Multi-modal Large Language Models (MLLMs) and Vision-Language Models (VLMs). The benchmark spans multiple simulation platforms and diverse task categories.
Key Features
Multi-Platform Coverage: AI2Thor, CARLA, ProcTHOR… See the full description on the dataset page: https://huggingface.co/datasets/Spatialworld/Spatialworld-bench.SpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/BenjaminChai579/SpatialGen-Testset.
