datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reef-guidance-system
Dataset Card for Reef Guidance System
This dataset provides imagery used for training and evaluation of models in the Reef Guidance System. All imagery was collected by the Australian Institute of Marine Science using the ReefScan™ Transom Marine Monitoring System.
If you use this dataset in your work, please cite the associated paper: AI-driven dispensing of coral reseeding devices for broad-scale restoration of the Great Barrier Reef (citations provided at bottom of this… See the full description on the dataset page: https://huggingface.co/datasets/QCR-Underwater-Perception/reef-guidance-system.synthetic-bathroom-dataset-for-robotic-perception
Synthetic Bathroom Dataset for Robotic Perception
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-bathroom-dataset-for-robotic-perception.synthetic-living-room-dataset-for-robotic-perception
Synthetic Living Room Dataset for Robotic Perception
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-living-room-dataset-for-robotic-perception.PerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.artificial-foveated-perception
Artificial Foveated Perception (AFP) Dataset
Training data for Artificial Foveated Perception (AFP), a task-conditioned mask predictor for robotic
foundation models (paper, code,
labeling tool).
The dataset contains 786 robot-manipulation episodes from real-world and simulated manipulation data. Every frame is paired with a continuous task-relevance matte: an alpha map in [0, 1]
that is 1 on the task-relevant objects and the robot end-effector, 0 on the background, and graded in… See the full description on the dataset page: https://huggingface.co/datasets/yalesunxiatao/artificial-foveated-perception.lighting-invariant-bedroom-perception-robustness-benchmark
Lighting-Invariant Bedroom Perception & Robustness Benchmark
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.AirV2X-Perceptionimaginative-perception-token-pet-ipt
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pet-ipt.gently-perception-benchmark
Gently Perception Agent Benchmark
Light-sheet microscopy volumes of C. elegans embryo development, intended
for evaluating vision-based perception agents on embryo stage classification.
The dataset has two tiers:
Annotated benchmark set (embryo_1–embryo_8) — human ground-truth
stage transitions. Use this for evaluation.
Unannotated corpus (embryo_9–embryo_105) — 97 additional real embryo
timelapses with no human labels, provided for developing and stress-
testing perception… See the full description on the dataset page: https://huggingface.co/datasets/gently-project/gently-perception-benchmark.PerceptionComp
PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning
PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.imaginative-perception-token-mvc-ipt
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-mvc-ipt.PerceptionTest_Valrss-robot-perception-evidence
Robot perception research artifacts
This is an exploratory experiment archive for an RSS research project. It is not a claim of acceptance, a finalized paper, or an official VER reproduction.
Code and protocol: https://github.com/YananZHOU5555/rss-robot-perception (private).
Weights: https://huggingface.co/B111ue/rss-robot-perception-checkpoints
Evidence: https://huggingface.co/datasets/B111ue/rss-robot-perception-evidence
This snapshot contains 228 complete evaluation cohorts… See the full description on the dataset page: https://huggingface.co/datasets/B111ue/rss-robot-perception-evidence.perception-testPerceptionTestperception_test_mcq
Perception Test MCQ Dataset
Dataset Description
This dataset contains 1000 video question-answering entries from the Perception Test dataset. Each entry includes a video and a multiple-choice question about the video content, testing various aspects of video understanding including object tracking, action recognition, and temporal reasoning.
Dataset Structure
This dataset follows the VideoFolder format with the following structure:
dataset/
├── data/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/advaitgupta/perception_test_mcq.visual-perception-benchmark-anon
Anonymous Visual Perception Benchmark
This evaluation dataset contains 2,876 image-question pairs with Chinese and
English questions, reference answers, and hierarchical visual-perception labels.
Images are embedded in Parquet files; no external image service is required.
Data Format
The dataset has one test split. Each row contains:
Field
Description
id
Release-local sample identifier, unrelated to internal identifiers.
image
Embedded PNG image… See the full description on the dataset page: https://huggingface.co/datasets/Susan0803/visual-perception-benchmark-anon.perception_lm_test_videosperception_lm_test_imagesskylink_perceptionperception-benchmarkimaginative-perception-token-pet-eval-habitat
Dataset Card for "habitat_perspective_eval"
More Information needed
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pet-eval-habitat.imaginative-perception-token-mvc-textcot
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-mvc-textcot.imaginative-perception-token-pet-eval-ai2thor
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pet-eval-ai2thor.so101_newton_perception_return_dr500_20260924
SO-101 Newton perception and return DR500
500 successful scripted simulation demonstrations, selected from 1,646 attempts.
267,345 frames at 30 FPS (148.525 minutes); LeRobot v3.0 format, external and
wrist RGB at 640 x 480. Only the final 500-episode dataset is included.
Task: "Pick up the vial and place it in the rack".
What this dataset tests
Combined perception-gap reduction, measured vial geometry, variable vial counts,
and complete post-placement return. The… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/so101_newton_perception_return_dr500_20260924.CWRU-perception
Dataset Card
Dataset Description
[Placeholder]
Dataset Structure
[Placeholder]
Uses
[Placeholder]
Limitations
[Placeholder]
License
[Placeholder]
Citation
[Placeholder]
V-Perception-40KLabHorizon-3D-Asset-Perception
LabHorizon 3D Asset Perception
Pushing the Limits of Laboratory 3D Perception and Long-Horizon Planning via Protocol-Aligned Action Prediction
Overview | News | Highlights | Dataset | Evaluation | Leaderboard | Training | Citation
🔎 Overview
This dataset is the Level 1 split of LabHorizon. Each example pairs three rendered views of the same laboratory asset with historical experimental actions and a set of candidate… See the full description on the dataset page: https://huggingface.co/datasets/black-yt/LabHorizon-3D-Asset-Perception.PADERBORN-perception
Dataset Card
Dataset Description
[Placeholder]
Dataset Structure
[Placeholder]
Uses
[Placeholder]
Limitations
[Placeholder]
License
[Placeholder]
Citation
[Placeholder]
OTTAWA-VARSPEED-perception
Dataset Card
Dataset Description
[Placeholder]
Dataset Structure
[Placeholder]
Uses
[Placeholder]
Limitations
[Placeholder]
License
[Placeholder]
Citation
[Placeholder]
