Pokerme/view2space-v1
VIEW2SPACE v1 VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning. Associated paper: VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 π arXiv: 2603.16506 Project Page: Project Page Related VIEW2SPACE Releases Training release: Pokerme/view2space-train 4B model checkpoint: Pokerme/view2space_4b Collection: Pokerme/view2space The public release is organized into three subsets: countβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-v1.
VIEW2SPACE v1
VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning.
Associated paper:
- VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 π
- arXiv: 2603.16506
- Project Page: Project Page
Related VIEW2SPACE Releases
- Training release: Pokerme/view2space-train
- 4B model checkpoint: Pokerme/view2space_4b
- Collection: Pokerme/view2space
The public release is organized into three subsets:
countdetectmcq
For official code, updates, and usage instructions, please see:
GitHubπ»: VIEW2SPACE GitHub repository
Directory Structure
view2space-v1-release/
README.md
count/
overall.jsonl
detect/
overall.jsonl
mcq/
overall.jsonl
images/
img_000001.png
img_000002.png
...Each subset contains one overall.jsonl file. Each line is one question.
Data Format
Each JSONL record contains the following public fields:
q_idx: question id, such asmcq_000001q_type: question family, one ofmcq,detect, orcountquestion: question textoptions: multiple-choice options for MCQ questions; empty for non-MCQ questionsquestion_prompt: extra prompt text when applicableanswer: ground-truth answerimage_paths: relative paths to the images used by the questionsupporting.draw_boxes: optional input boxes associated with one or more images
Notes:
image_pathsare relative to the dataset root, not to the subset folder.- Example image path:
images/img_000123.png q_typeuses only the public family labelsmcq,detect, andcount.
Example Record
{
"q_idx": "mcq_000001",
"q_type": "mcq",
"question": "All views are captured from the same static scene arrangement, but from different positions and angles.\nIn view 2, consider the nearest 'man with an orange hat, crouching'. Where is the 'white temple' relative to this 'man with an orange hat, crouching'?",
"options": {
"A": "left",
"B": "right",
"C": "front",
"D": "back"
},
"question_prompt": "",
"answer": "C",
"image_paths": [
"images/img_000813.png",
"images/img_000806.png",
"images/img_000815.png"
],
"supporting": {
"draw_boxes": null
}
}Basic Usage
import json
from pathlib import Path
dataset_root = Path("/path/to/view2space-v1-release")
jsonl_path = dataset_root / "mcq" / "overall.jsonl"
with jsonl_path.open("r", encoding="utf-8") as f:
first = json.loads(next(f))
image_files = [dataset_root / rel_path for rel_path in first["image_paths"]]
print(first["q_idx"])
print(first["q_type"])
print(image_files)Citation
License and Citation
This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Attribution is required under the license. For academic and research use, citation in the resulting paper, report, model card, dataset card, or public documentation is expected. Please cite this dataset whenever it is used for training, evaluation, benchmarking, data analysis, or as part of a larger dataset mixture.
@article{ke2026view2space,
title={VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations},
author={Ke, Fucai and Cai, Zhixi and Li, Boying and Chen, Long and Lin, Beibei and Wang, Weiqing and Haghighi, Pari Delir and Haffari, Gholamreza and Rezatofighi, Hamid},
journal={arXiv preprint arXiv:2603.16506},
year={2026}
}