datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eurosat-rgb
EuroSat (RGB)
Description
A dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples. This is the RGB version of the dataset with visible bands encoded as JPEG images.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/eurosat-rgb.table_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.EuroSAT_RGB
EuroSAT RGB
EUROSAT RGB is the RGB version of the EUROSAT dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples.
Paper: https://arxiv.org/abs/1709.00029
Homepage: https://github.com/phelber/EuroSAT
Description
The EuroSAT dataset is a comprehensive land cover classification dataset that focuses on images taken by the ESA Sentinel-2 satellite. It contains a total of 27… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/EuroSAT_RGB.M3FD_RGBTfMoW_rgbdroid-3d-rgb-68epThis is a FiftyOne dataset with 68 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/droid-3d-rgb-68ep")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for droid_3d (68-episode FiftyOne RGB subset)
A… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/droid-3d-rgb-68ep.SPADES-RGBrgb_articubotdataset for rgb articubot
bigearthnet-v2-rgb
BigEarthNet v2.0 (reBEN) - Sentinel-2 RGB
Description
BigEarthNet v2.0 (reBEN, "refined BigEarthNet"), a multi-label land-cover classification dataset of Sentinel-2 image patches from 10 European countries. Each 1200 m x 1200 m patch (120x120 pixels at 10 m) is labelled with one or more of 19 land-cover classes. The labels come from the CORINE Land Cover 2018 map, using the 19-class BigEarthNet nomenclature.
This variant is RGB-only, stored as JPEG (quality 98… See the full description on the dataset page: https://huggingface.co/datasets/timm/bigearthnet-v2-rgb.citrus-fruit-rgbdThis dataset contains
train: 1500 images test: 500 images val: 200 images. Each RGB image also has its corresponding depth file (.npy), and the B (Laplacian-based convexity cues), N (surface bumpness), generated from the depth image.
so101_pick_place_blue_dual_rgb_20260929
SO-101 pick place — dual RGB
Task: Pick up the black 25 mm cube, place it in the blue region, and release the gripper.
LeRobot v3.0, 20 FPS, 37 recorded episodes, 30 human-labeled successful demonstrations. Cameras: observation.images.wrist and observation.images.third_person, 640×480 RGB. State/action: six joints; body joints in degrees, gripper 0–100. Actions are sent joint targets, not measured end-effector commands.
All recorded episodes, including failures and excluded… See the full description on the dataset page: https://huggingface.co/datasets/LUOSYrrrrr/so101_pick_place_blue_dual_rgb_20260929.egocentric-stereo-rgbd
Hub Egocentric: Stereo RGB-D
9 egocentric stereo RGB-D clips with dense metric depth, IMU, 6-DoF VIO pose, and calibration. Each clip carries a self-contained LeRobot v3.0 dataset (loads on lerobot >= 0.6.0) plus side-by-side stereo, mono, a colorized depth preview, and a Foxglove MCAP recording.
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on StereoLabs ZED X Mini. Egocentric, human-demonstration data (passive; no robot action stream). July… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-stereo-rgbd.egocentric-gopro-rgb-imu
Hub Egocentric: GoPro RGB+IMU
19 egocentric human-manipulation clips captured on GoPro HERO13, with high-rate IMU (~200 Hz GPMF) delivered as CSV/Parquet/JSON sidecars plus a Foxglove MCAP recording.
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on GoPro HERO13. Egocentric, human-demonstration data (passive; no robot action stream). July 2026.
Dataset structure
Each clip is a top-level folder named #NN_... holding its media, a… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-gopro-rgb-imu.piperx-old-ab-rgbd-5903
PiperX Old AB RGB-D — 5,903 accepted points
This public dataset export contains the exact 5,903 accepted training Episodes recorded by the old AB SE(3) perturbation collector. The export is bound to the aggregate state.json truth and contains 5,903 unique schedule_index values across 124 closed shards (shard-00000 through shard-00123). Audit rejections remain audit records and are not training Episodes.
Data
Schema: piperx_lerobot_se3_rgbd_v2
Accepted Episodes /… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/piperx-old-ab-rgbd-5903.Pixel-aligned_RGB-NIR_stereo_dataset
Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision
CVPR 2025Jinnyeong Kim, Seung-Hwan BaekPOSTECH[arXiv] • [Code] • [Video] • [Dataset on HuggingFace]
Overview
This repository provides the code and dataset accompanying our CVPR 2025 paper:
"Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision"
We propose a novel robotic vision system equipped with two pixel-aligned RGB-NIR stereo cameras and a LiDAR sensor mounted on a mobile robot. Our… See the full description on the dataset page: https://huggingface.co/datasets/DivisonOfficer/Pixel-aligned_RGB-NIR_stereo_dataset.eurosat-rgb
EuroSat (RGB)
Description
A dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples. This is the RGB version of the dataset with visible bands encoded as JPEG images.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/mteb/eurosat-rgb.so101_pick_lift_dual_rgb_20260929
SO-101 pick lift — dual RGB
Task: Pick up the black 25 mm cube from the blue spawn region, lift it approximately 5 cm above the table, and hold it for 1 second.
LeRobot v3.0, 20 FPS, 32 recorded episodes, 30 human-labeled successful demonstrations. Cameras: observation.images.wrist and observation.images.third_person, 640×480 RGB. State/action: six joints; body joints in degrees, gripper 0–100. Actions are sent joint targets, not measured end-effector commands.
All recorded… See the full description on the dataset page: https://huggingface.co/datasets/LUOSYrrrrr/so101_pick_lift_dual_rgb_20260929.game-scenes-posed-rgbd
Origin Lab Game Scenes: Posed RGB-D Flythroughs of Game Worlds
Every frame carries the camera that rendered it and the depth the engine computed for it. Ten game worlds, with the camera released from the player for 60% of the footage: metric depth, world-space normals, 4x4 pose, and per-frame intrinsics on one frame index, plus hundreds of full in-place turns and long stretches in which the world is frozen and only the camera moves. Two trajectories per world and one whole… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-scenes-posed-rgbd.rgbt234robotwin-blocks-ranking-rgb-rollouts
RoboTwin blocks_ranking_rgb — Wan2.2 TI2V Rollouts
160 text+image-to-video rollouts (10 initial conditions × 16 random seeds) for the
blocks_ranking_rgb task from RoboTwin, generated with the Wan2.2 TI2V (5B)
diffusion model fine-tuned with a merged Vidar LoRA adapter, and scored with the
blocks_ranking_v2 reward (SAM3 object tracking + IDM inverse-dynamics + FK
gripper ↔ block position matching).
Companion to the EmbodiedVideoRL / DanceGRPO
reward-model work.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/VincentNi/robotwin-blocks-ranking-rgb-rollouts.egocentric-iphone-rgb-imu
Hub Egocentric: iPhone RGB+IMU
12 egocentric human-manipulation clips captured on iPhone, with nominal 30 Hz CoreMotion + ARKit IMU/attitude on the shared media timeline, delivered as CSV/Parquet/JSON sidecars plus a Foxglove MCAP recording and per-clip camera intrinsics.
Across this 12-clip sample, IMU row count is 0–4 boundary rows lower than decoded video frame count (≤0.035%).
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on iPhone 13… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-iphone-rgb-imu.EgoManipulation-RGBD-Beta
EgoManipulation RGB-D Beta
EgoManipulation RGB-D Beta is a collection of egocentric manipulation recordings containing RGB and depth streams time-synchronized using timestamps from a shared device clock.
A Parquet catalog provides recording-level metadata and previews for browsing the dataset without downloading the full MCAP collection.
Data recordings are stored as MCAP files, with one selected recording chunk per file.
Repository layout
README.md
data/… See the full description on the dataset page: https://huggingface.co/datasets/PI-Dojo/EgoManipulation-RGBD-Beta.fashion-second-hand-front-only-rgb
Clothing Dataset for Second-Hand Fashion
This dataset contains only the front image and labels from version 3 of the following dataset released on zenodo:
Clothing Dataset for Second-Hand Fashion
Three changes were made:
Front image: Only front image is uploaded here. Back and brand image are not.
Background removal: The background from the front image was removed using BiRefNet, which only supports up to 1024x1024 images - larger images were resized. The background removal is… See the full description on the dataset page: https://huggingface.co/datasets/fnauman/fashion-second-hand-front-only-rgb.bigearthnet-v2-rgb-nir-swir
BigEarthNet v2.0 (reBEN) - Sentinel-2 RGB + NIR + SWIR
Description
BigEarthNet v2.0 (reBEN, "refined BigEarthNet"), a multi-label land-cover classification dataset of Sentinel-2 image patches from 10 European countries. Each 1200 m x 1200 m patch (120x120 pixels at 10 m) is labelled with one or more of 19 land-cover classes. The labels come from the CORINE Land Cover 2018 map, using the 19-class BigEarthNet nomenclature.
This variant stores the tone-mapped RGB as… See the full description on the dataset page: https://huggingface.co/datasets/timm/bigearthnet-v2-rgb-nir-swir.EuroSAT_RGB
EuroSAT RGB
Dataset Description
EuroSAT is a dataset for land use and land cover (LULC) classification using Sentinel-2 satellite imagery. This version contains the RGB (visible spectrum) bands encoded as JPEG images at 64x64 pixel resolution.
The dataset covers 10 land use/land cover classes across 27,000 geo-referenced images from 34 European countries.
Source: https://zenodo.org/records/7711810
DOI: 10.5281/zenodo.7711810
License: MIT
Paper: EuroSAT: A Novel Dataset… See the full description on the dataset page: https://huggingface.co/datasets/giswqs/EuroSAT_RGB.xd-violence-rgb-videomae-chunked-testsentinel-2-rgb-captionedGitHub: https://github.com/sshh12/terrain-diffusion
Source
COPERNICUS_S2_SR
Contains modified Copernicus Service information 2023
Captions
Generated using openai/clip-vit-large-patch14 on a fixed set of captions.
Data Processing
Pick region, filter for low clouds, get median over fixed timespan
Normalize RGB by min/max channel values
Generate captions for each tile using CLIP
Filter out tiles with black parts or overlap lines (automatically)
Manually pass to… See the full description on the dataset page: https://huggingface.co/datasets/sshh12/sentinel-2-rgb-captioned.overhead-people-rgb
Overhead People RGB
Unified overhead RGB people detection dataset converted from local Roboflow YOLOv8 exports.
Images are stored as original encoded bytes without resizing, recompression, or preprocessing. All retained boxes use a single category:
category_id: 0
category: person
bbox: COCO-style [x, y, width, height] in pixel coordinates
Source labels such as Man, Woman, Person, ero, and 0 are preserved in objects.source_category. Source label object is omitted.
## Loading The… See the full description on the dataset page: https://huggingface.co/datasets/bdanko/overhead-people-rgb.runcam-feed-camera-egocentric-rgb-imu
RunCam Feed Camera - Egocentric RGB + IMU Sample Dataset
A small sample dataset captured with the RunCam Feed Camera for egocentric video and synchronized motion-sensor workflows.
Capture Device
Video: H.265 MP4, 1920x1080, 60 fps for V01-V06
Nominal video bitrate: 18 Mbps
Horizontal field of view: 126 degrees
Device weight: approximately 26 g
IMU: ICM-42607, 6-axis
IMU sampling rate: 800 Hz for the recordings in this sample
Firmware reported in the GCSV files:… See the full description on the dataset page: https://huggingface.co/datasets/RunCam/runcam-feed-camera-egocentric-rgb-imu.five-cam-xyz-rgb-1024
FiveCam Dataset
XYZ-RGB Pairs (1024*1024) simultaneously captured by five different smartphones.
Data Format
Keys of each data item in the self-camera XYZ-RGB datasets:
raw_device
The name of the device that captures the current raw (XYZ) image.
rgb_device
The name of the device that captures the current RGB image. For the self-camera dataset, the value of rgb_device is the same as the value of raw_device.
raw_image
Resized & cropped XYZ image in PIL.Image format.… See the full description on the dataset page: https://huggingface.co/datasets/l-li/five-cam-xyz-rgb-1024.
