datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CameraBench
📷 CameraBench: Towards Understanding Camera Motions in Any Video
SfMs and VLMs performance on CameraBench: Generative VLMs (evaluated with VQAScore) trail classical SfM/SLAM in pure geometry, yet they outperform discriminative VLMs that rely on CLIPScore/ITMScore and—even better—capture scene‑aware semantic cues missed by SfM
After simple supervised fine‑tuning (SFT) on ≈1,400 extra annotated clips, our 7B Qwen2.5‑VL doubles its AP, outperforming the current best… See the full description on the dataset page: https://huggingface.co/datasets/syCen/CameraBench.IDLE-OO-Camera-Traps
Dataset Card for IDLE-OO Camera Traps
IDLE-OO Camera Traps is a 5-dataset benchmark of camera trap images from the Labeled Information Library of Alexandria: Biology and Conservation (LILA BC) with a total of 2,586 images for species classification. Each of the 5 benchmarks is balanced to have the same number of images for each species within it (between 310 and 1120 images), representing between 16 and 39 species.
Supported Tasks and Leaderboards
Image… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/IDLE-OO-Camera-Traps.CameraBenchProcamera_pizza_additionalcamera-movement-human-preference-324k
Rapidata Camera Movement Benchmark
Built by Rapidata.
This dataset contains 324,044 human responses, collected with the
Rapidata Python SDK, comparing how well 15 image-to-video models and world models
execute a described camera movement from a single still image. Each row is a head-to-head comparison between
two models' clips generated from the same still and the same instruction, judged by human annotators who
watched a reference animation of the requested movement.
The task… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/camera-movement-human-preference-324k.Robocasa_Camera_Space_Eef
RoboCasa Camera-Space EEF (ACE-Ego-0)
LeRobot v2.0 demonstrations for the 24 GR1 RoboCasa TableTop tasks used by
ACE-Ego-0 SFT. Actions and states
are stored in the observation camera frame (end-effector position + rot6d, plus
gripper and waist).
This dataset does not include unused conversion caches (experiments/) or
legacy per-dataset index files. Training reads only the files listed below and
the shipped merged q99 statistics. The training code never recomputes norms.… See the full description on the dataset page: https://huggingface.co/datasets/ACERobotics/Robocasa_Camera_Space_Eef.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.G1_Dex3_CameraPackaging_DatasetThis dataset was created using LeRobot.
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in Part 5 of AVP Teleoperation Documentation.
Data collection is not completed in a single session, and variations between data entries exist. Ensure these variations are accounted for during model training.
Dataset Structure
meta/info.json:
{
"codebase_version":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex3_CameraPackaging_Dataset.flir-camera-objects
Flir Camera Objects
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
9,306
Validation
2,854
Test
1,452
Total
13,612
Classes (4)
bicycle
car
dog
person
Usage
With LibreYOLO
from libreyolo import LIBREYOLO
# Load a model
model = LIBREYOLO(model_path="libreyoloXnano.pt")
# Train on this dataset… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/flir-camera-objects.thewilds_cameratraps
Dataset Card for The Wilds Camera Trap Data
This dataset contains images and video captured from camera traps deployed at The Wilds safari park in Ohio during Summer 2025. It supports ecological monitoring, animal behavior analysis, and biodiversity studies.
Dataset Details
This dataset was created to support wildlife monitoring research using camera traps. Data were collected using various camera trap models, with each camera recording photos and videos in the… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/thewilds_cameratraps.InteriorVerse-Camera
InteriorVerse-Camera
Per-image camera parameter annotations for the InteriorVerse dataset
(a large-scale photorealistic synthetic indoor dataset rendered from professionally designed scenes, shipped with per-frame albedo / depth / normal / material maps; 61,231 images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows:… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/InteriorVerse-Camera.PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.camera
Dataset Card for CAMERA📷:
Table of Contents:
Dataset Card for Camera
Table of Contents
Dataset Details
Dataset Description
Dataset Sources
Uses
Direct Use
Dataset Information
Data Example
Dataset Structure
Citation
Dataset Details
Dataset Description
CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) is the Japanese ad text generation dataset, which comprises actual data sourced from Japanese search ads and incorporates… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/camera.Megalith-10M-Camera
Megalith-10M-Camera
Per-image camera parameter annotations for the Megalith-10M dataset
(the image-bearing build drawthingsai/megalith-10m; ~9.58M Flickr photos across
959 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Megalith-10M-Camera.camerabench_cutcamera-motion-dataset-and-benchmark
Camera Motion Dataset and Benchmark
This repository packages the dataset and benchmark release for our paper "Geometry-Guided Camera Motion Understanding in VideoLLMs", accepted to the CVPR 2026 Workshop on Pixel-level Video Understanding in the Wild (PVUW2026), as a Hugging Face dataset repo. Derived from the KlingTeam/MultiCamVideo-Dataset.
It contains:
WebDataset tar shards for train, val, and test
Split manifests (train.json, val.json, test.json) that map source clips to sample… See the full description on the dataset page: https://huggingface.co/datasets/fengyee/camera-motion-dataset-and-benchmark.pseudo-camera-10k
pseudo-camera-10k dataset
Contents
This dataset contains 10k free images from world class photographers. The images have been resized using Lanczos antialiasing, with their smaller edge shifted to 1024px.
The aim of this dataset is a highly variable but high quality and high resolution set of images containing difficult concepts, with about half of the images being numbered group shots and family portraits with the number of subjects labeled.
No images were upsampled in… See the full description on the dataset page: https://huggingface.co/datasets/bghira/pseudo-camera-10k.jetson1-top-camera-083126
jetson1-top-camera-083126
Recorded dataset — captured on jetson1 — 5 episodes · 222 frames @ 20 fps (~0 min of demonstration).
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting board, aligned vertically in the image, with the back end of the cucumber close to the upper edge of the cutting board
2
Recording
Rig
jetson1 (calibration sidecar)
Recorded
2026-08-31
Operator
dorischen
Episode… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-top-camera-083126.RealEstate10K-Absolute-Camera
RealEstate10K-Absolute-Camera
Per-frame camera parameter annotations for the RealEstate10K (RE10K) dataset
(5,134 scenes; 595,382 valid per-frame annotations),
captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/RealEstate10K-Absolute-Camera.CameraClone-Dataset
CamCloneMaster: Enabling reference-based camera control for video generation
Paper:https://arxiv.org/abs/2506.03140
Project Page:https://camclonemaster.github.io/
Dataset:https://huggingface.co/datasets/KwaiVGI/CameraClone-Dataset
Training & Inference Code:https://github.com/KwaiVGI/CamCloneMaster
Camera Clone Dataset
1. Dataset Introduction
TL;DR: The Camera Clone Dataset, introduced in CamCloneMaster, is a large-scale synthetic dataset designed… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/CameraClone-Dataset.jetson1-top-camera-test082526
jetson1-top-camera-test082526
Recorded dataset — captured on jetson1 — 1 episodes · 157 frames @ 20 fps (~0 min of demonstration).
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting board, aligned vertically in the image, with the back end of the cucumber close to the upper edge of the cutting board
1
Recording
Rig
jetson1 (calibration sidecar)
Recorded
2026-08-25
Operator
dorischen
Episode… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-top-camera-test082526.ur5_camera_ready_routingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5",
"total_episodes": 30,
"total_frames": 18010,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DORLR/ur5_camera_ready_routing.Open-Images-Camera
Open-Images-Camera
Per-image camera parameter annotations for the Open Images dataset
(the image-bearing WebDataset build dalle-mini/open-images; 8,338,780 images across 1,838 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Open-Images-Camera.relative-camera-movement-human-preference-900k
Rapidata Relative Camera Movement Benchmark
Built by Rapidata.
This dataset contains 900,318 human responses, collected with the
Rapidata Python SDK, comparing how well 15 image-to-video models and world models
move the camera relative to what is in the scene. Each row is a head-to-head comparison between two models' clips
generated from the same still and the same instruction, judged by human annotators.
Our Camera Movement Benchmark asks for scene-agnostic
moves — "tilt thirty… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/relative-camera-movement-human-preference-900k.G1_Dex1_Store_Cameralila_camera_trapsLILA Camera Traps is an aggregate data set of images taken by camera traps, which are devices that automatically (e.g. via motion detection) capture images of wild animals to help ecological research.
This data set is the first time when disparate camera trap data sets have been aggregated into a single training environment with a single taxonomy.
This data set consists of only camera trap image data sets, whereas the broader LILA website also has other data sets related to biology and conservation, intended as a resource for both machine learning (ML) researchers and those that want to harness ML for this topic.omr-primus-cameraImageNet-1K-Camera
ImageNet-1K-Camera
Per-image camera parameter annotations for the full ImageNet-1K dataset
(1,000 training classes + the 50,000-image validation split, ~1.35M images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/ImageNet-1K-Camera.record-test-camera
LeRobot aims to provide models, datasets, and tools for real-world robotics in PyTorch. The goal is to lower the barrier to entry so that everyone can contribute to and benefit from shared datasets and pretrained models.
🤗 A hardware-agnostic, Python-native interface that standardizes control across diverse platforms, from low-cost arms (SO-100) to humanoids.
🤗 A standardized, scalable LeRobotDataset format (Parquet + MP4 or images) hosted on the Hugging Face Hub, enabling… See the full description on the dataset page: https://huggingface.co/datasets/phuongpm72/record-test-camera.openve_fixed_bernini1p3_opd_vmse_allin_rlv2_r256_eb48_s200_step200_camera
