datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3d-spatial-reasoning-2embodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.thai_spatial_reasoning
Thai Spatial Reasoning 1.0.0
Thai spatial captions and question-answer pairs for continual pretraining of a Thai
foundation VLM. Synthetic means the Thai text and the spatial annotations are generated or
processed; the primary images are real photographs.
Images are not covered by one blanket licence, so the card does not point at one file:
every image has its own licence, creator, and attribution in
rights/attribution.csv and
rights/ledger.parquet. The dataset-level other… See the full description on the dataset page: https://huggingface.co/datasets/SmartWhatt/thai_spatial_reasoning.grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.spatial-visual-reasoning-66kopen-spatial-reasoning
Open Spatial Reasoning
A multiple-choice dataset of spatial reasoning questions and answers for evaluating 3D spatial reasoning from single driving images. Each image contains numbered bounding boxes referencing objects in the scene, and each question probes a model's ability to reconstruct the real 3D scene rather than rely on flat-image shortcuts (e.g. "lower in the frame = closer", "bigger box = nearer").
Dataset Description
Frontier vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/ReasonCore/open-spatial-reasoning.geometric_spatial_compose_reasoningViRL39K-Spatial_ReasoningObject-Centric-Spatial-Relation-Reasoningscaled_synthetic_spatial_reasoningqwen-spatial-reasoning-incorrect-examples
Incorrect spatial reasoning examples for Qwen/Qwen3.5-0.8B-Base
Overview
This dataset contains incorrect non-empty predictions made by Qwen/Qwen3.5-0.8B-Base on a synthetic spatial reasoning benchmark built from 4x4 object-grid images.
I evaluated the model on 84 questions. It answered 54 of them incorrectly and achieved an overall accuracy of 35.714%. Some incorrect rows had an empty parsed pred_final, which I treat as formatting failures rather than useful supervised… See the full description on the dataset page: https://huggingface.co/datasets/safaeid48/qwen-spatial-reasoning-incorrect-examples.wm_benchmark_spatial_reasoningpara_Spatial_Reasoning
