datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CameraBench
📷 CameraBench: Towards Understanding Camera Motions in Any Video
SfMs and VLMs performance on CameraBench: Generative VLMs (evaluated with VQAScore) trail classical SfM/SLAM in pure geometry, yet they outperform discriminative VLMs that rely on CLIPScore/ITMScore and—even better—capture scene‑aware semantic cues missed by SfM
After simple supervised fine‑tuning (SFT) on ≈1,400 extra annotated clips, our 7B Qwen2.5‑VL doubles its AP, outperforming the current best… See the full description on the dataset page: https://huggingface.co/datasets/syCen/CameraBench.IDLE-OO-Camera-Traps
Dataset Card for IDLE-OO Camera Traps
IDLE-OO Camera Traps is a 5-dataset benchmark of camera trap images from the Labeled Information Library of Alexandria: Biology and Conservation (LILA BC) with a total of 2,586 images for species classification. Each of the 5 benchmarks is balanced to have the same number of images for each species within it (between 310 and 1120 images), representing between 16 and 39 species.
Supported Tasks and Leaderboards
Image… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/IDLE-OO-Camera-Traps.camera-movement-human-preference-324k
Rapidata Camera Movement Benchmark
Built by Rapidata.
This dataset contains 324,044 human responses, collected with the
Rapidata Python SDK, comparing how well 15 image-to-video models and world models
execute a described camera movement from a single still image. Each row is a head-to-head comparison between
two models' clips generated from the same still and the same instruction, judged by human annotators who
watched a reference animation of the requested movement.
The task… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/camera-movement-human-preference-324k.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.camera
Dataset Card for CAMERA📷:
Table of Contents:
Dataset Card for Camera
Table of Contents
Dataset Details
Dataset Description
Dataset Sources
Uses
Direct Use
Dataset Information
Data Example
Dataset Structure
Citation
Dataset Details
Dataset Description
CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) is the Japanese ad text generation dataset, which comprises actual data sourced from Japanese search ads and incorporates… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/camera.pseudo-camera-10k
pseudo-camera-10k dataset
Contents
This dataset contains 10k free images from world class photographers. The images have been resized using Lanczos antialiasing, with their smaller edge shifted to 1024px.
The aim of this dataset is a highly variable but high quality and high resolution set of images containing difficult concepts, with about half of the images being numbered group shots and family portraits with the number of subjects labeled.
No images were upsampled in… See the full description on the dataset page: https://huggingface.co/datasets/bghira/pseudo-camera-10k.CameraClone-Dataset
CamCloneMaster: Enabling reference-based camera control for video generation
Paper:https://arxiv.org/abs/2506.03140
Project Page:https://camclonemaster.github.io/
Dataset:https://huggingface.co/datasets/KwaiVGI/CameraClone-Dataset
Training & Inference Code:https://github.com/KwaiVGI/CamCloneMaster
Camera Clone Dataset
1. Dataset Introduction
TL;DR: The Camera Clone Dataset, introduced in CamCloneMaster, is a large-scale synthetic dataset designed… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/CameraClone-Dataset.relative-camera-movement-human-preference-900k
Rapidata Relative Camera Movement Benchmark
Built by Rapidata.
This dataset contains 900,318 human responses, collected with the
Rapidata Python SDK, comparing how well 15 image-to-video models and world models
move the camera relative to what is in the scene. Each row is a head-to-head comparison between two models' clips
generated from the same still and the same instruction, judged by human annotators.
Our Camera Movement Benchmark asks for scene-agnostic
moves — "tilt thirty… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/relative-camera-movement-human-preference-900k.CameraBench-Pro
CameraBench-Pro
This dataset contains the testing split for the CameraBench-Pro evaluation.
pseudo-camera-10k-structured-json
pseudo-camera-10k, structured JSON captions
The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled.
The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.CameraOperator-BlockCam
CameraOperator-BlockCam (Synthetic)
CameraOperator-BlockCam (Synthetic) contains 37,499 repaired-and-audited synthetic annotation-label trajectory records pairing English camera-motion descriptions, 150-frame camera trajectories, and time-varying target-object 3D oriented bounding boxes (OBBs). They are grouped into 5,180 reconstructed source events and 13,121 augmentation families; the 37,499 records should not be interpreted as 37,499 independent scenes or events. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/huuuuuuuuu/CameraOperator-BlockCam.camera-grandstaffEgocentric-3-Camera-Array
Egocentric 3-Camera Array (Head + Both Wrists)
Three long-form household and warehouse tasks recorded simultaneously from three body-mounted cameras — head, left wrist and right wrist — each with per-frame timestamps and its own high-rate gyroscope and accelerometer.
This is a bimanual manipulation dataset: the wrist cameras see what each hand is doing at close range while the head camera carries the scene context.
Preview: 45 s of the Cooking task, all three views at the same… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Egocentric-3-Camera-Array.vggt-camera-moveruncam-feed-camera-egocentric-rgb-imu
RunCam Feed Camera - Egocentric RGB + IMU Sample Dataset
A small sample dataset captured with the RunCam Feed Camera for egocentric video and synchronized motion-sensor workflows.
Capture Device
Video: H.265 MP4, 1920x1080, 60 fps for V01-V06
Nominal video bitrate: 18 Mbps
Horizontal field of view: 126 degrees
Device weight: approximately 26 g
IMU: ICM-42607, 6-axis
IMU sampling rate: 800 Hz for the recordings in this sample
Firmware reported in the GCSV files:… See the full description on the dataset page: https://huggingface.co/datasets/RunCam/runcam-feed-camera-egocentric-rgb-imu.Arena-DROID-Camera-Sensitivity-Workflow-Sample
Arena DROID Camera Sensitivity Workflow Sample
Dataset Description
Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep.
The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.drone-navigation-event-camera
Edged-USLAM Dataset: Drone Navigation with Event Camera
🔗 Citation
@inproceedings{sariozkan2026edged,
title={Edged USLAM: Edge-Aware Event-Based SLAM with Learning-Based Depth Priors},
author={Sarıözkan, Şebnem and Şahin, Hürkan and Álvarez-Tuñón, Olaya and Kayacan, Erdal},
booktitle={IEEE International Conference on Robotics and Automation (ICRA)},
year={2026}
}
ℹ️ Extra Info
Project Page
This dataset contains synchronized Event Camera (DAVIS346)… See the full description on the dataset page: https://huggingface.co/datasets/sebnem-byte/drone-navigation-event-camera.parliamentrag-camera-leg19
ParliamentRAG — Italian Chamber of Deputies, 19th legislature
Full proceedings of the Italian Chamber of Deputies (Camera dei deputati) for the 19th legislature, from the first sitting on 13 October 2022 through 6 August 2026: verbatim speech transcripts, roll-call votes with every individual ballot, parliamentary acts with EuroVoc subjects, and the deputies' group, committee and government memberships over time. The dataset is refreshed as new sittings are ingested.
The tables… See the full description on the dataset page: https://huggingface.co/datasets/emeierkeio/parliamentrag-camera-leg19.camera-calib-and-scene-alignment-dataferma-camera-yolo
Ferma Camera YOLO
Dataset Summary
Цель датасета — обучить YOLO распознавать людей, собак, коров и хищников на камерах в ферме.
Data Sources
Данные собраны из нескольких источников:
часть изображений взята из OIDv6
часть изображений взята из датасета 8 calves
часть изображений взята с YouTube и размечена вручную
Data Structure
images/ — изображения .jpg
labels/ — разметка YOLO .txt (same stem)
Annotation Format
Формат YOLO: class_id… See the full description on the dataset page: https://huggingface.co/datasets/I77/ferma-camera-yolo.WDC-cameraschannel-islands-camera-traps
Channel Islands Camera Traps — Detection
Camera-trap images with bounding-box annotations from the Channel Islands, California, provided by
The Nature Conservancy and distributed via
LILA BC.
Unofficial community mirror — not affiliated with or endorsed by LILA BC. This is an
HF-native, viewer-ready object-detection build (parquet, dataset viewer renders boxes,
trains in one command). For the full multi-dataset collection use the streaming loader
society-ethics/lila_camera_traps.… See the full description on the dataset page: https://huggingface.co/datasets/lila-bc-community/channel-islands-camera-traps.camerabench_vqa_lmms_eval2026-24679-camera-lens-dataset
Camera Lens Prime/Zoom Dataset
Dataset summary
This dataset contains physical and design specifications for 35 Sony camera-lens models. The binary classification task predicts whether a lens is prime or zoom.
Focal length is intentionally excluded because it would directly reveal the target.
Collection
The specifications were compiled from official Sony E-mount lens support pages. They are published manufacturer specifications, not physical… See the full description on the dataset page: https://huggingface.co/datasets/cannj/2026-24679-camera-lens-dataset.bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.act-mobile-speed-camera-visits
Mobile speed camera visits, ACT
Every visit of an ACT mobile speed camera since July 2016, one row per stay, with the site, hours on site, vehicles checked and posted speed limit.
This is a copy of the 2026-10-04 version of this dataset on publicdata.au, where it has 94,329 rows and 9 fields. The same version is kept at https://publicdata.au/d/act-mobile-speed-camera-visits/v/2026-10-04/ in 9 formats, with every earlier version and a query API.
Attribution
The… See the full description on the dataset page: https://huggingface.co/datasets/National-Digital/act-mobile-speed-camera-visits.wildlife_conservation_camera_trap_datasetA camera trap dataset created as a subset of the Wildlife Conservation Society (WCS) camera trap dataset, af featured
on: https://lila.science/datasets/wcscameratraps.
The subset contains 14,000 images split into a test set of 2,000 images and a pool of
12,000 images of 6 species of animal, namely pecari tajacu, loxodonta africana, aepyceros melampus, equus quagga, madoqua guentheri
and crax rubra that can be used to generate various different splits.
The dataset also includes 3 pre-done… See the full description on the dataset page: https://huggingface.co/datasets/jnle/wildlife_conservation_camera_trap_dataset.ai-video-camera-movements
AI Video Camera Movements: 43 prompt recipes with example clips
43 camera moves for AI video (pans, tilts, zooms, dollies, tracking shots, orbits, cranes, drone moves, FPV, and specials like infinite zoom and time-lapse). Each record has a ready-to-paste prompt recipe, practical tips, negatives, a recommended clip length, and a 1280×720 example clip.
The recipes are plain cinematography language, so they work with any text-to-video or image-to-video model.
Watch every move in… See the full description on the dataset page: https://huggingface.co/datasets/FueledByImagination/ai-video-camera-movements.act-traffic-camera-offences
Traffic and camera offences and fines, ACT
Traffic and camera infringements issued in the ACT since July 2021, one row per offence, day, camera location, registration state and client type, with the count and the penalty total.
This is a copy of the 2026-10-06 version of this dataset on publicdata.au, where it has 199,446 rows and 8 fields. The same version is kept at https://publicdata.au/d/act-traffic-camera-offences/v/2026-10-06/ in 9 formats, with every earlier version and a… See the full description on the dataset page: https://huggingface.co/datasets/National-Digital/act-traffic-camera-offences.
