datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3d-front-rgbbehavior_224_rgbThis is the compressed version of the original BEHAVIOR dataset
It contains only RGB videos compressed to 224x224 as well as actions, annotations, and metadata files.
Depth and segmentation data are removed.
The dataset is just ~260GB, which makes it easier to use than the original one if you don't need all the data.
We used this dataset for our 1st place solution in the NeurIPS 2025 BEHAVIOR Challenge. Code, tech report.
Citation
@article{li2024behavior,
title={Behavior-1k:… See the full description on the dataset page: https://huggingface.co/datasets/IliaLarchenko/behavior_224_rgb.RGB-Event-ISP-Datasetagibotworld-beta-rgb-lance
AgiBotWorld-Beta (LeRobot lance format)
agibot-world/AgiBotWorld-Beta
converted to the LeRobot lance storage format, uploaded in coordination with the AgiBot team.
160,454 episodes, 286,556,463 frames, 8 RGB cameras (AV1, copied from the source without re-encoding),
state[20] and action[22] at 30 fps.
Read it in place, no download needed (lerobot with the lancedb extra):
from lerobot.datasets import LeRobotDataset
ds = LeRobotDataset("lance-format/agibotworld-beta-rgb-lance")… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/agibotworld-beta-rgb-lance.RAW-RAIN-rgb
Dataset Card for R³ RGB Dataset
Dataset Summary
The R³ (Reconstruction, Raw, and Rain) RGB Dataset is a large-scale, real-world stereo dataset designed for deraining tasks directly in the raw sRGB domain. It was collected using a dual-camera synchronized setup capturing raw Bayer images, which under went a software ISP. Unlike popular synthetic datasets, R³ uses a real-world rain simulation system to generate realistic rain artifacts, including scene depth effects and… See the full description on the dataset page: https://huggingface.co/datasets/realrainmaker/RAW-RAIN-rgb.TUM_RGBD-SLAMRGBench-Cloth-Sim2Real-v1
RGBench Cloth Sim-to-Real (v1)
🌐 Project page: https://rgbench.github.io/ · 📦 Code: https://github.com/hwk0809/RGBench
Nine carefully captured garments — three bimanual manipulation actions
each (fling / fold / grasp) — with real-world ground truth point
clouds for evaluating any cloth simulator's sim-to-real gap. Released
as the evaluation half of the AAAI 2026 paper Real Garment Benchmark
(RGBench).
The larger 6 000+ garment-mesh asset library and the GarmentDynamics… See the full description on the dataset page: https://huggingface.co/datasets/RGBench/RGBench-Cloth-Sim2Real-v1.eurosat-rgb
EuroSat (RGB)
Description
A dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples. This is the RGB version of the dataset with visible bands encoded as JPEG images.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/eurosat-rgb.gso-orbit-rgbatable_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.demogen-rgbd-fr3-cube-stacking
demogen-rgbd — FR3 cube stacking, 10,000 augmented episodes
Augmented teleoperation data for a Franka FR3 cube-stacking task, produced by
demogen-rgbd: a DemoGen-style pipeline that turns one teleoperated episode
into many by translating the manipulated cube in 3-D and re-rendering the scene
from the original RGB-D observations.
No generative model is involved. Every pixel is a real observation, warped
by a dense inverse RGB-D transform. The only synthesised regions are the… See the full description on the dataset page: https://huggingface.co/datasets/Jordano/demogen-rgbd-fr3-cube-stacking.SPAR-7M-RGBD
📦 Spatial Perception And Reasoning Dataset – RGBD (SPAR-7M-RGBD)
A large-scale multimodal dataset for 3D-aware spatial perception and reasoning in vision-language models.
SPAR-7M-RGBD extends the original SPAR-7M with additional depths, camera intrinsics, and pose information. It contains over 7 million QA pairs across 33 spatial tasks, built from 4,500+ richly annotated indoor 3D scenes.
This version supports single-view, multi-view, and… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-7M-RGBD.RGB-Event-ISP-DatasetEuroSAT_RGB
EuroSAT RGB
EUROSAT RGB is the RGB version of the EUROSAT dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples.
Paper: https://arxiv.org/abs/1709.00029
Homepage: https://github.com/phelber/EuroSAT
Description
The EuroSAT dataset is a comprehensive land cover classification dataset that focuses on images taken by the ESA Sentinel-2 satellite. It contains a total of 27… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/EuroSAT_RGB.OVIS_RGBD
OVIS_RGBD
The dataset is for paper "Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation".
You can find the usages in GitHub.
The original frame images and annotations are from OVIS. We use DepthAnythingV2 to perform monocular depth estimation on all images. We concatenate depth map on the channel demension and each image is in RGBD format.
Citations
@InProceedings{niu2025,
author = {Niu, Quanzhu and Zhou, Yikang and Chen, Shihao and… See the full description on the dataset page: https://huggingface.co/datasets/QuanzhuNiu/OVIS_RGBD.M3FD_RGBTbehavior1k-only-rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/behavior1k-only-rgb.Syn4D_RGBD
Dataset Card for Syn4D RGBD
Syn4D is a large-scale, fully-synthetic multiview dataset of dynamic scenes designed to advance research in 4D reconstruction, depth estimation, 3D point tracking, novel-view synthesis, and human pose estimation. It provides dense, complete, and accurate geometric annotations — including per-pixel depth maps, multi-view camera trajectories, dense long-range 3D point tracks, and parametric SMPL-X human body annotations — across a diverse collection of… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Syn4D_RGBD.ouro-1.4b-thinking-evals
Ouro looped-LM experiments
Experiments on ByteDance's Ouro-1.4B-Thinking looped language model, run on a Colab T4 on 2026-09-20:
small GSM8K / MBPP banks, activation-size accounting, linear probes for the loop index, a per-loop logit lens,
and Contrastive Activation Addition steering with a loop sweep.
ouro_eval.ipynb is the full Colab notebook with outputs (includes the transformers 4.54 cache patch Ouro needs).
The recorded residual-stream activations (1.9 GB, probe/after{0,6… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/ouro-1.4b-thinking-evals.TUM_RGBD-SLAMfMoW_rgbAnti-UAV-RGBTRGB-D-SegmentEgocentricBodiesannotations_creators:
- other
language:
- en
language_creators:
- other
license:
- odc-by
multilinguality:
- monolingual
pretty_name: 'RGB-D-SegmentEgocentricBodies '
size_categories:
- 1K<n<10K
source_datasets:
- original
tags:
- egocentric segmentation
- extended reality
- xr
- human-body
- mixed-reality
- avatar
task_categories:
- image-segmentation
- depth-estimation
task_ids:
- semantic-segmentation
- features:
- name: image
dtype: image
- name: depth
dtype: image… See the full description on the dataset page: https://huggingface.co/datasets/ExtendedRealityLab/RGB-D-SegmentEgocentricBodies.eurosat-rgbwb_pointed_chair_pull_push_rgb
wb_pointed_chair_pull_push_rgb
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_pointed_chair_pull_push_rgb
Task — pull out the chair indicated by the human gesture, then push it… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_pointed_chair_pull_push_rgb.RGBT-Ground-Datasetbert_cot_em
Can you tell a model is about to misbehave by reading its reasoning?
Short answer: no — but you can change what it does by writing its reasoning for it.
This repo studies a large language model that has been deliberately made
misaligned, and asks whether its chain-of-thought (the "thinking out loud" it
does before answering) gives away that a harmful answer is coming.
The setup in plain terms
Researchers found that fine-tuning a model on bad medical advice makes… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/bert_cot_em.droid-3d-rgb-68epThis is a FiftyOne dataset with 68 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/droid-3d-rgb-68ep")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for droid_3d (68-episode FiftyOne RGB subset)
A… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/droid-3d-rgb-68ep.Huntington_and_Caltech_Koi_RGB_and_Sonar
Huntington and Caltech Koi RGB and Sonar
Huntington paired and supplementary RGB/sonar exports for 15ft, 40ft, 100ft and vertical_plane.
The calibrated runs contain full uncropped 989×512 grayscale sonar with complete outer bounds and the original top margin. One corrected-paper matrix per run incorporates RGB rotation and maps directly onto native RGB. See homography instructions.
Oversized folders retain part_001/part_002 subdivisions, and JSON paths match those folders. RGB… See the full description on the dataset page: https://huggingface.co/datasets/perona-lab/Huntington_and_Caltech_Koi_RGB_and_Sonar.SPADES-RGB
