datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiviewX_LabelsMultiViewBench
MultiView-Bench
Evaluation data for MultiView-Bench: A Diagnostic Benchmark for
World-Centric Multi-View Integration in VLMs
(arXiv:2607.08970).
This repository currently contains the synthetic image-and-text portion of the
benchmark. The real-world subset is not included; see the scope note below.
Contents
2,400 evaluation samples across 24 benchmark variants
7,900 PNG views (2.94 GiB logical image data)
One, three, or six images per sample
English question text… See the full description on the dataset page: https://huggingface.co/datasets/suijinru/MultiViewBench.multiview-pouring
MultiView Pouring Dataset, v1.0.
by Pierre Sermanet, Corey Lynch, Jasmine Hsu and Eric Jang
License
This data is licensed by Google Inc. under a Creative Commons Attribution 4.0 International License.
Downloading
Because of some downloading issues for a specific file, the file was split in two, call https://huggingface.co/datasets/sermanet/multiview-pouring/blob/main/tfrecords/test/whiteorange_to_clear1_real_combining.sh to recombine the parts.… See the full description on the dataset page: https://huggingface.co/datasets/sermanet/multiview-pouring.test_multiview_3d_reconstruction_3camsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 375,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/test_multiview_3d_reconstruction_3cams.aloha_multiviewEgoDex-PickPlace-YAM-14dof-multiview
EgoDex → YAM 14-DOF, Multiview (LeRobot v2.1)
Egocentric human hand-manipulation demonstrations from EgoDex retargeted to a
YAM bimanual robot (14-DOF), packaged as a LeRobot v2.1 dataset with three
synthesized camera views. The observation schema is drop-in compatible with
angkul07/abc-teleop for
cotraining (identical Hz, camera keys, and action convention).
At a glance
Episodes
8,842
Frames
1,074,893
Control rate
30 Hz (30 fps video)
Robot
YAM… See the full description on the dataset page: https://huggingface.co/datasets/angkul07/EgoDex-PickPlace-YAM-14dof-multiview.test_multiview_3d_reconstructionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 757,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/test_multiview_3d_reconstruction.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.PixelArt_Multiview
Multiview PixelArt
Dataset Summary
Contains sets of images representing a full 360° turnaround of characters, animals and objects in pixel art.
Each row contains 9 images from all angles.
Camera Data can be downloaded
Examples
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9… See the full description on the dataset page: https://huggingface.co/datasets/Scaryplasmon96/PixelArt_Multiview.surgical-tool-recognition-full-multiview
Surgical Tool Recognition Full Multiview
Summary
This dataset contains images of individual surgical instruments for object detection.It was originally created in YOLO format and exported here to a Hugging Face-friendly structure with metadata.jsonl files for each split.
Splits
train: 2016
validation: 252
test: 252
Total: 2520 images
Classes
0 = clamp
1 = needle_holder
2 = scalpel
3 = shear
4 = tweezer
File naming convention
Each image… See the full description on the dataset page: https://huggingface.co/datasets/joonhaim/surgical-tool-recognition-full-multiview.twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 300,
"total_frames": 100000,
"total_tasks": 33245,
"chunks_size": 1000,
"data_files_size_in_mb": 300,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50.fold_onesie_multiview_20251025This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 75749,
"total_tasks": 1,
"total_videos": 250,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_onesie_multiview_20251025.robofactory-camera-alignment-multiviewrobofactory-long-pipeline-delivery-multiviewmultiview_incabin_dataset
Multi-View In-Cabin Dataset
We introduce a multi-view in-cabin monitoring dataset for public transportation with synchronized RGB and depth images from four inward-facing cameras and a rotating LiDAR covering the vehicle interior of a digitalized and partly automated German city bus. The dataset contains 9.136 synchronized samples with annotations and is accompanied by a calibration and pseudo-labeling pipeline that generates 3D human pose estimates and oriented 3D bounding… See the full description on the dataset page: https://huggingface.co/datasets/evgenygorelik96/multiview_incabin_dataset.robofactory-place-food-multiviewmessytable-split-multiviewwire_pick_place_multi_view_expandedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "xarm_end_effector",
"total_episodes": 1,
"total_frames": 745,
"total_tasks": 0,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/spesrobotics/wire_pick_place_multi_view_expanded.robofactory-lift-barrier-multiviewmultiview-datasets
DX.GL Multi-View Datasets for NeRF & 3D Gaussian Splatting
Multi-view training datasets rendered from CC0 3D models via DX.GL. Each dataset includes calibrated camera poses, depth maps, normal maps, binary masks, and point clouds — ready for nerfstudio out of the box.
10 objects × 196 views × 1024×1024 resolution × full sphere coverage.
Quick Start
# Download a dataset (Apple, 196 views, 1024x1024)
wget https://dx.gl/api/v/EJbs8npt2RVM/vCHDLxjWG65d/dataset -O… See the full description on the dataset page: https://huggingface.co/datasets/dxgl/multiview-datasets.fold_onesie_reverse_human_multiview_20251210This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 25,
"total_frames": 11889,
"total_tasks": 1,
"total_videos": 125,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_onesie_reverse_human_multiview_20251210.fold_bottoms_multiview_20251031This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 25175,
"total_tasks": 1,
"total_videos": 250,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_bottoms_multiview_20251031.fold_towel_human_multiview_20251122This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 17026,
"total_tasks": 1,
"total_videos": 250,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_towel_human_multiview_20251122.fold_towel_multiview_20251030This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 43140,
"total_tasks": 1,
"total_videos": 250,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_towel_multiview_20251030.cell-seg-multiview_fixedrobofactory-take-photo-multiviewmultiview-sslfold_towel_human_multiview_20251210This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 15267,
"total_tasks": 1,
"total_videos": 250,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_towel_human_multiview_20251210.fold_shirt_multiview_20251030This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 72754,
"total_tasks": 1,
"total_videos": 250,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_shirt_multiview_20251030.fold_onesie_human_multiview_20251210This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 25,
"total_frames": 11798,
"total_tasks": 1,
"total_videos": 125,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/fold_onesie_human_multiview_20251210.
