datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3dfront_render_views3dfront-render-viewsViewSpatial-Bench
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dataset Description
We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.canopyboard-chm-viewerAwesome_Spatial_VQA_Benchmarks_ViewSpatial-BenchViewGraphBench
ViewGraphBench
Synthetic view graphs with exact ground truth for the global stage of Structure-from-Motion.
Toolkit and generator on GitHub | viewgraphforge on PyPI | Citation | MIT license
ViewGraphBench is a benchmark for rotation averaging, translation averaging, SE(3) pose-graph optimisation and the detection of outlier edges. A node of a view graph is a calibrated camera with a known pose. An edge is a measured relative pose between two cameras (rotation, unit translation… See the full description on the dataset page: https://huggingface.co/datasets/ezharjan/ViewGraphBench.partnetsim-1024-fixed-viewpointsUAV-Aerial-View-Battle-Tank-Detection-Dataset
Aerial UAV Perspective: Battle Tank Detection Dataset
Overview
This is an open-source synthetic dataset specifically engineered to train computer vision models in identifying main battle tanks (MBTs) and armored combat vehicles from tactical overhead and drone perspectives.
Obtaining real-world tactical aerial imagery for defense analytics is heavily restricted, operationally dangerous, or classified. This dataset addresses that critical data bottleneck by… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/UAV-Aerial-View-Battle-Tank-Detection-Dataset.viewpoint-aware-pig-posture-recognition
Viewpoint-Aware Pig Posture Recognition Dataset
This dataset supports multi-camera, viewpoint-aware pig posture recognition in livestock barn environments. It contains real-world pig images, bounding box annotations, posture class labels, and per-instance camera viewpoint angles (azimuth and elevation) derived from PnP-based camera calibration.
Code: Anil-Bhujel/viewpoint-aware-pig-posture-recognition on GitHub
Dataset Summary
Images were captured from 2… See the full description on the dataset page: https://huggingface.co/datasets/anilbhujel/viewpoint-aware-pig-posture-recognition.view_variation_finalViewSpatial_lmmsevalasset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.simready-usd-web-viewers
SimReady assets in browser USD viewers
Six real SimReady OpenUSD packages from the Hub, loaded in five browser USD libraries and a pre-converted GLB baseline. Each cell is what the library drew.
Asset
Reference
three.js 0.174
three.js 0.186
three.js 0.186 + crawler
Needle
tinyusdz
cinevva usdjs
GLB (pre-converted)
LG laptopusdc · 11 MB
❌ zip error
✅ renders1.3 s · 279 MB
✅ renders1.3 s · 378 MB
✅ renders1.2 s · 1.3 GB
✅ renders1.3 s · 446 MB
✅ renders2.2 s · 327 MB
✅… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/simready-usd-web-viewers.ego4d-random-views-20k
Ego4D Random Views Dataset
This dataset contains 20,000 random view frames sampled from the Ego4D dataset using a high-performance multi-process generation system.
Dataset Overview
Total Images: 20,000 high-quality frames
Image Format: PNG (1024×1024 resolution)
Source: Ego4D v2 dataset (52,665+ video files)
Sampling Method: Multi-process random sampling with maximum diversity
Generation Time: 797.57 seconds (~13 minutes)
Generation Speed: 25.08 frames/second… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/ego4d-random-views-20k.view2space-v1
VIEW2SPACE v1
VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning.
Associated paper:
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 🚀
arXiv: 2603.16506
Project Page: Project Page
Related VIEW2SPACE Releases
Training release: Pokerme/view2space-train
4B model checkpoint: Pokerme/view2space_4b
Collection: Pokerme/view2space
The public release is organized into three subsets:
count… See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-v1.view2space-train
VIEW2SPACE Training
This repository contains the public VIEW2SPACE training release for multi-view vision-language spatial reasoning.
Associated paper:
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 🚀
arXiv: 2603.16506
Project Page: Project Page
The training release contains:
604,779 training examples
22,205 public images
three public question families: count, detect, and mcq
Unlike the evaluation release, this training package keeps… See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-train.reachability-viewfinder
取景通路 Reachability Viewfinder Benchmark (v1)
同一数据集也发布于 ModelScope:https://modelscope.cn/datasets/tysgydk/reachability-viewfinder
500 题多模态空间可达性 / 连通性推理基准测试。
任务定义
每道题向模型展示 5 张图片:1 张主图 + 4 张候选局部图(A/B/C/D)。
主图:一个三维场景,绿色旗帜为起点、红色旗帜为终点,两点间的路径中段被虚线框遮挡(标记问号),遮挡区内的真实路径已隐藏;
候选图:遮挡区域的 4 种可能填充内容,图片边缘保留少量场景参照,用于对齐位置、高度与朝向。
模型需判断哪个选项能使路径从起点到终点物理连通。单选题,有且仅有一个正确答案,模型仅输出一个字母。
数据组织
questions/
q0001/
main.png 主图(含遮挡区)
A.png ~ D.png 候选局部图… See the full description on the dataset page: https://huggingface.co/datasets/tysgydke/reachability-viewfinder.Cross-view_Wireless_MIMO_Datasetpersian-ocr-gemini37-wins-bina-misses-viewer
Gemini 3.7 exact / Bina miss OCR crops
53 non-empty bbox crops from the PersianVLM submitted-10 benchmark where
google/gemini-3.7-flash was normalized-exact and Bina Koochik 0.1 was not.
This is a minimal Hugging Face ImageFolder dataset for reliable viewer support.
It contains exactly two columns: image and ocr. The ocr value is Gemini's
actual output for the corresponding crop.
ai2thor-random-views-20kgeo_viewer_tropomi_locationsshard-viewer-gridMulti_View_Industrial_part_Dataset
Dataset Card for Dataset Name
This dataset contains 7 industrial parts taken from a 360-degree all-around view with one axis of rotation.
It contains a train set of all-around images and test sets with ground truth anomaly masks.
Paper: 3D Gaussian Reference Parts for Robust Free-Viewpoint Visual Inspection
BibTeX entry and citation info
@ARTICLE{11424398,
author={Ito, Kenta and Ueda, Shiori and Mori, Shohei and Sugano, Junichi and Adachi, Hideyuki and Saito, Hideo}… See the full description on the dataset page: https://huggingface.co/datasets/kentaito321/Multi_View_Industrial_part_Dataset.modelnet40_multi_viewhuhb-viewer-new-version-2026-06-01
Huhb3D-Industrial-100: 6DoF Pose Estimation Dataset with Topology Labels
Dataset Summary
A synthetic 6DoF pose estimation dataset featuring 22 industrial parts with 15 topology categories parsed directly from STEP CAD files. All data is real OpenGL rendering — zero fabrication.
Metric
Value
Objects
22 industrial parts
Frames
2,200 (100 per object)
Image size
640×480
Mask categories
Up to 9 per object (15 total)
Depth precision
16-bit, mm… See the full description on the dataset page: https://huggingface.co/datasets/Hgodwarrior/huhb-viewer-new-version-2026-06-01.dart-libero-viewpointsviewer_gs
🔒 Licence
Proprietary – All rights reserved
L'intégralité de ce dataset est protégée par le droit d’auteur.
Tous les fichiers sont © 2025 Mika,
Aucun fichier ne peut être copié, modifié, distribué, ou utilisé sans autorisation écrite préalable,
📬 Contact
Pour toute demande de licence, collaboration ou utilisation commerciale, merci de contacter : contact.mikafilleul@gmail.com
habitat-views-20kdifferent_view_dataset
Dataset Card for "different_view_dataset_train"
More Information needed
