multi-view
qwen3_4b_multiview-i1-GGUFintent-aware-lfqa-qwen3-8b-multiview-GGUFintent-aware-lfqa-llama3-8b-multiview-GGUFintent-aware-lfqa-qwen3-4b-multiview-GGUFil-xl-charturn-merged-multi-view-turnaround-model-sheet-character-designqwen3_4b_multiview-GGUFintent-aware-lfqa-llama3-8b-multiviewintent-aware-lfqa-qwen3-4b-multiview
MultiviewX_LabelsMultiViewBench
MultiView-Bench
Evaluation data for MultiView-Bench: A Diagnostic Benchmark for
World-Centric Multi-View Integration in VLMs
(arXiv:2607.08970).
This repository currently contains the synthetic image-and-text portion of the
benchmark. The real-world subset is not included; see the scope note below.
Contents
2,400 evaluation samples across 24 benchmark variants
7,900 PNG views (2.94 GiB logical image data)
One, three, or six images per sample
English question text… See the full description on the dataset page: https://huggingface.co/datasets/suijinru/MultiViewBench.multiview-pouring
MultiView Pouring Dataset, v1.0.
by Pierre Sermanet, Corey Lynch, Jasmine Hsu and Eric Jang
License
This data is licensed by Google Inc. under a Creative Commons Attribution 4.0 International License.
Downloading
Because of some downloading issues for a specific file, the file was split in two, call https://huggingface.co/datasets/sermanet/multiview-pouring/blob/main/tfrecords/test/whiteorange_to_clear1_real_combining.sh to recombine the parts.… See the full description on the dataset page: https://huggingface.co/datasets/sermanet/multiview-pouring.test_multiview_3d_reconstruction_3camsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 375,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/test_multiview_3d_reconstruction_3cams.aloha_multiviewEgoDex-PickPlace-YAM-14dof-multiview
EgoDex → YAM 14-DOF, Multiview (LeRobot v2.1)
Egocentric human hand-manipulation demonstrations from EgoDex retargeted to a
YAM bimanual robot (14-DOF), packaged as a LeRobot v2.1 dataset with three
synthesized camera views. The observation schema is drop-in compatible with
angkul07/abc-teleop for
cotraining (identical Hz, camera keys, and action convention).
At a glance
Episodes
8,842
Frames
1,074,893
Control rate
30 Hz (30 fps video)
Robot
YAM… See the full description on the dataset page: https://huggingface.co/datasets/angkul07/EgoDex-PickPlace-YAM-14dof-multiview.
