datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sentence_transformer_kmeans100_viewerViewSpatial-Bench
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dataset Description
We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.Awesome_Spatial_VQA_Benchmarks_ViewSpatial-BenchViewGraphBench
ViewGraphBench
Synthetic view graphs with exact ground truth for the global stage of Structure-from-Motion.
Toolkit and generator on GitHub | viewgraphforge on PyPI | Citation | MIT license
ViewGraphBench is a benchmark for rotation averaging, translation averaging, SE(3) pose-graph optimisation and the detection of outlier edges. A node of a view graph is a calibrated camera with a known pose. An edge is a measured relative pose between two cameras (rotation, unit translation… See the full description on the dataset page: https://huggingface.co/datasets/ezharjan/ViewGraphBench.UAV-Aerial-View-Battle-Tank-Detection-Dataset
Aerial UAV Perspective: Battle Tank Detection Dataset
Overview
This is an open-source synthetic dataset specifically engineered to train computer vision models in identifying main battle tanks (MBTs) and armored combat vehicles from tactical overhead and drone perspectives.
Obtaining real-world tactical aerial imagery for defense analytics is heavily restricted, operationally dangerous, or classified. This dataset addresses that critical data bottleneck by… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/UAV-Aerial-View-Battle-Tank-Detection-Dataset.wikipedia-viewerwikipedia dataset, now with viewer enabled! :D
Dataset Card for Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The datasets are built from the Wikipedia dump
(https://dumps.wikimedia.org/) with one split per language. Each example
contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
The articles are parsed using the mwparserfromhell tool, which can be… See the full description on the dataset page: https://huggingface.co/datasets/GGUFGuy/wikipedia-viewer.dreamdojo-ego-view
DreamDojo ego-view — video + instruction for Cosmos-Predict2.5 post-training
3,168 ego-view robot manipulation clips in the flat videos/ + metas/ layout that
cosmos-predict2.5's VideoDataset reads directly, with no conversion step.
Only the observation.images.ego_view camera is included. Episodes shorter than 94
frames are excluded: VideoDataset samples a random 93-frame window, and on a
93-frame video its np.random.randint(0, 0) raises rather than returning 0.… See the full description on the dataset page: https://huggingface.co/datasets/Eurong2/dreamdojo-ego-view.ViewSpatial_lmmsevalHEC3R-ckpt-fixed_view_axis_5robotsHEC3R-ckpt-fixed_view_axis_unfrozen_decasset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.simready-usd-web-viewers
SimReady assets in browser USD viewers
Six real SimReady OpenUSD packages from the Hub, loaded in five browser USD libraries and a pre-converted GLB baseline. Each cell is what the library drew.
Asset
Reference
three.js 0.174
three.js 0.186
three.js 0.186 + crawler
Needle
tinyusdz
cinevva usdjs
GLB (pre-converted)
LG laptopusdc · 11 MB
❌ zip error
✅ renders1.3 s · 279 MB
✅ renders1.3 s · 378 MB
✅ renders1.2 s · 1.3 GB
✅ renders1.3 s · 446 MB
✅ renders2.2 s · 327 MB
✅… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/simready-usd-web-viewers.twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 300,
"total_frames": 100000,
"total_tasks": 33245,
"chunks_size": 1000,
"data_files_size_in_mb": 300,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50.zillow-viewer
Housing Data Provided by Zillow
Updated: 2023-02-01
This dataset contains several configs produced based on files available at https://www.zillow.com/research/data/.
Processing Notes
This dataset contains only parquet files created from the raw Zillow data. For more information, as well as code related to processing that data and creating the parquet files see https://huggingface.co/datasets/misikoff/zillow.
Supported configs:
days_on_market: Days to pending, days to… See the full description on the dataset page: https://huggingface.co/datasets/misikoff/zillow-viewer.ego4d-random-views-20k
Ego4D Random Views Dataset
This dataset contains 20,000 random view frames sampled from the Ego4D dataset using a high-performance multi-process generation system.
Dataset Overview
Total Images: 20,000 high-quality frames
Image Format: PNG (1024×1024 resolution)
Source: Ego4D v2 dataset (52,665+ video files)
Sampling Method: Multi-process random sampling with maximum diversity
Generation Time: 797.57 seconds (~13 minutes)
Generation Speed: 25.08 frames/second… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/ego4d-random-views-20k.view2space-v1
VIEW2SPACE v1
VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning.
Associated paper:
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 🚀
arXiv: 2603.16506
Project Page: Project Page
Related VIEW2SPACE Releases
Training release: Pokerme/view2space-train
4B model checkpoint: Pokerme/view2space_4b
Collection: Pokerme/view2space
The public release is organized into three subsets:
count… See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-v1.persian-ocr-gemini37-wins-bina-misses-viewer
Gemini 3.7 exact / Bina miss OCR crops
53 non-empty bbox crops from the PersianVLM submitted-10 benchmark where
google/gemini-3.7-flash was normalized-exact and Bina Koochik 0.1 was not.
This is a minimal Hugging Face ImageFolder dataset for reliable viewer support.
It contains exactly two columns: image and ocr. The ocr value is Gemini's
actual output for the corresponding crop.
View-Spatial-Benchai2thor-random-views-20kHEC3R-ckpt-fixed_view_axis_freeze_decCS50-rawlongbench-view
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.auxiliary-views-knowledge-acquisition
Auxiliary Views Knowledge Acquisition
This repository contains the cleaned source documents and evaluation
probes used in Knowledge Acquisition During Pre-training? Large Language Models
Learn Better With Auxiliary Views (arXiv:2609.04180).
News
August 21, 2026: Our paper was accepted to Findings of EMNLP 2026.
Configurations
Configuration
Split
Rows
documents
train
30
factual_cloze
test
6,435
factual_mcqa_5shot
test
4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.pi-sessions-viewer
Coding agent session traces for aaaaliou/pi-sessions-viewer
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.Viewpoint-100Kchange-my-view-subreddit-cleaned
Opinionated LLM
modelnet40_multi_viewhuhb-viewer-new-version-2026-06-01
Huhb3D-Industrial-100: 6DoF Pose Estimation Dataset with Topology Labels
Dataset Summary
A synthetic 6DoF pose estimation dataset featuring 22 industrial parts with 15 topology categories parsed directly from STEP CAD files. All data is real OpenGL rendering — zero fabrication.
Metric
Value
Objects
22 industrial parts
Frames
2,200 (100 per object)
Image size
640×480
Mask categories
Up to 9 per object (15 total)
Depth precision
16-bit, mm… See the full description on the dataset page: https://huggingface.co/datasets/Hgodwarrior/huhb-viewer-new-version-2026-06-01.habitat-views-20k
