MilaWang/SpatialEval
🤔 About SpatialEval SpatialEval is a comprehensive benchmark for evaluating spatial intelligence in LLMs and VLMs across four key dimensions: Spatial relationships Positional understanding Object counting Navigation Benchmark Tasks Spatial-Map: Understanding spatial relationships between objects in map-based scenarios Maze-Nav: Testing navigation through complex environments Spatial-Grid: Evaluating spatial reasoning within structured environments… See the full description on the dataset page: https://huggingface.co/datasets/MilaWang/SpatialEval.
🤔 About SpatialEval
SpatialEval is a comprehensive benchmark for evaluating spatial intelligence in LLMs and VLMs across four key dimensions:
- Spatial relationships
- Positional understanding
- Object counting
- Navigation
Benchmark Tasks
- Spatial-Map: Understanding spatial relationships between objects in map-based scenarios
- Maze-Nav: Testing navigation through complex environments
- Spatial-Grid: Evaluating spatial reasoning within structured environments
- Spatial-Real: Assessing real-world spatial understanding
Each task supports three input modalities:
- Text-only (TQA)
- Vision-only (VQA)
- Vision-Text (VTQA)

📌 Quick Links
Project Page: https://spatialeval.github.io/
Paper: https://arxiv.org/pdf/2406.14852
Code: https://github.com/jiayuww/SpatialEval
Talk: https://neurips.cc/virtual/2024/poster/94371
🚀 Quick Start
📍 Load Dataset
SpatialEval provides three input modalities—TQA (Text-only), VQA (Vision-only), and VTQA (Vision-text)—across four tasks: Spatial-Map, Maze-Nav, Spatial-Grid, and Spatial-Real. Each modality and task is easily accessible via Hugging Face. Ensure you have installed the packages:
from datasets import load_dataset
tqa = load_dataset("MilaWang/SpatialEval", "tqa", split="test")
vqa = load_dataset("MilaWang/SpatialEval", "vqa", split="test")
vtqa = load_dataset("MilaWang/SpatialEval", "vtqa", split="test")⭐ Citation
If you find our work helpful, please consider citing our paper 😊
@inproceedings{wang2024spatial,
title={Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models},
author={Wang, Jiayu and Ming, Yifei and Shi, Zhenmei and Vineet, Vibhav and Wang, Xin and Li, Yixuan and Joshi, Neel},
booktitle={The Thirty-Eighth Annual Conference on Neural Information Processing Systems},
year={2024}
}💬 Questions
Have questions? We're here to help!
- Open an issue in the github repository
- Contact us through the channels listed on our project page
