datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiViewBench
MultiView-Bench
Evaluation data for MultiView-Bench: A Diagnostic Benchmark for
World-Centric Multi-View Integration in VLMs
(arXiv:2607.08970).
This repository currently contains the synthetic image-and-text portion of the
benchmark. The real-world subset is not included; see the scope note below.
Contents
2,400 evaluation samples across 24 benchmark variants
7,900 PNG views (2.94 GiB logical image data)
One, three, or six images per sample
English question text… See the full description on the dataset page: https://huggingface.co/datasets/suijinru/MultiViewBench.aloha_multiviewmulti-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.PixelArt_Multiview
Multiview PixelArt
Dataset Summary
Contains sets of images representing a full 360° turnaround of characters, animals and objects in pixel art.
Each row contains 9 images from all angles.
Camera Data can be downloaded
Examples
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9… See the full description on the dataset page: https://huggingface.co/datasets/Scaryplasmon96/PixelArt_Multiview.multiview-sslsurgical-tool-recognition-full-multiview
Surgical Tool Recognition Full Multiview
Summary
This dataset contains images of individual surgical instruments for object detection.It was originally created in YOLO format and exported here to a Hugging Face-friendly structure with metadata.jsonl files for each split.
Splits
train: 2016
validation: 252
test: 252
Total: 2520 images
Classes
0 = clamp
1 = needle_holder
2 = scalpel
3 = shear
4 = tweezer
File naming convention
Each image… See the full description on the dataset page: https://huggingface.co/datasets/joonhaim/surgical-tool-recognition-full-multiview.messytable-split-multiviewcell-seg-multiview_fixedcell-seg-multiview_fixedV2Multi_View_Industrial_part_Dataset
Dataset Card for Dataset Name
This dataset contains 7 industrial parts taken from a 360-degree all-around view with one axis of rotation.
It contains a train set of all-around images and test sets with ground truth anomaly masks.
Paper: 3D Gaussian Reference Parts for Robust Free-Viewpoint Visual Inspection
BibTeX entry and citation info
@ARTICLE{11424398,
author={Ito, Kenta and Ueda, Shiori and Mori, Shohei and Sugano, Junichi and Adachi, Hideyuki and Saito, Hideo}… See the full description on the dataset page: https://huggingface.co/datasets/kentaito321/Multi_View_Industrial_part_Dataset.classroom-multiview
Classroom Multiview Renders
Four diverse 960×540 views of a Blender classroom scene, rendered with Cycles + OptiX on an NVIDIA H100 (64 samples).
File
View
views/view_00.png
rear-left → front(后墙字母墙侧)
views/view_01.png
rear-right → front(窗户侧)
views/view_02.png
front-left → rear(黑板侧)
views/view_03.png
front-right → rear(门侧)
contact_sheet_diverse.jpg — 2×2 contact sheet
private_camera_metadata.json — camera locations / targets / focal length (24mm)
Rendered… See the full description on the dataset page: https://huggingface.co/datasets/yitongl/classroom-multiview.modelnet40_multi_viewsim_transferCube_multiview_scripted_piThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/stevex0/sim_transferCube_multiview_scripted_pi.cell-seg-multiviewsim_transferCube_multiview_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/stevex0/sim_transferCube_multiview_scripted.droid-multiviewmulti-view_caption
Multi-View Caption
Per-segment captions for multi-view human (DNA-Rendering, ActorsHQ) and animal
(Artemis / DFA) video datasets. Captions are generated from masked multi-view
composites with Gemini 3 Flash and follow the ActivityNet-style dense
video-captioning layout: one row per video, with parallel captions and
timestamps lists plus a representative thumbnail.
Citation
If you find our caption useful, please cite our paper
Flex4DHuman: Flexible Multi-view Video… See the full description on the dataset page: https://huggingface.co/datasets/andaba/multi-view_caption.fiftyone-multiview-reid-attributes
📦 FiftyOne-Compatible Multiview Person ReID with Visual Attributes
A curated, attribute-rich person re-identification dataset based on Market-1501, enhanced with:
✅ Multi-view images per person
✅ Detailed physical and clothing attributes
✅ Natural language descriptions
✅ Global attribute consolidation
📊 Dataset Statistics
Subset
Samples
Train
3,181
Query
1,726
Gallery
1,548
Total
6,455
📥 Installation
Install the required… See the full description on the dataset page: https://huggingface.co/datasets/adonaivera/fiftyone-multiview-reid-attributes.ai2thor-multiview-counting-val-800-v2Multi-view-CXR
EVOKE
Patient-Specific Multimodal Learning with Multi-View Contrastive Alignment for Chest X-ray Report Generation
Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology report generation presents a promising solution, existing methods often rely on single-view radiographs, which constrain… See the full description on the dataset page: https://huggingface.co/datasets/MK-runner/Multi-view-CXR.multiview_panohead
Dataset Card for "multiview_panohead"
More Information needed
ordinary-bench-multiview
ORDINARY-BENCH Multi-View Dataset
A multi-view version of the ORDINARY-BENCH benchmark for evaluating Vision-Language Models (VLMs) on ordinal spatial reasoning in 3D scenes. Each sample includes 4 camera views of the same scene.
Single-view version: TYTSTQ/ordinary-bench
Source code & evaluation pipeline: GitHub - tasd12-ty/ordinary-bench-core
Overview
Scenes
700 synthetic 3D scenes (Blender, CLEVR-style)
Complexity
7 levels: 4 to 10 objects per… See the full description on the dataset page: https://huggingface.co/datasets/TYTSTQ/ordinary-bench-multiview.multiview_eval_filteredmultiview_eval
Dataset Card for "multiview_eval"
More Information needed
fitcheck-scraped-multiviewmultiview_panohead
Dataset Card for "multiview_panohead"
More Information needed
text2cad_multiview
Dataset Summary
This dataset is designed for instruction tuning vision-language models (VLMs) on 3D Computer-Aided Design (CAD) comprehension and interactive engineering reasoning. It scales up the Text2CAD dataset dataset by transforming static CAD assets and multi-level design prompts into a multimodal, conversational format.
To bridge the gap between static 3D CAD data and conversational AI, we extended the original dataset by:
Rendering multi-view images for each CAD asset… See the full description on the dataset page: https://huggingface.co/datasets/trislee02/text2cad_multiview.multiview-datasetmulti-view-room-datasetmulti-view-room-dataset/
README.md
Metadata UI
license
task_categories
language
tags
pretty_name
size_categories
Add num of elements
Import dataset card template
1
2
3
4
⌄
license: cc-by-nc-sa-4.0
Commit directly to the
main
branch
Open as a pull request to the
main
branch
Commit changes
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
qwen-vl-multiview-analysis-1748101249
Dataset Card for "qwen-vl-multiview-analysis-1748101249"
More Information needed
