datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/droid_1.0.1.3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.thinking_droid_lerobot_output_qwen3vlMolmoAct2-DROID-Dataset
MolmoAct2-DROID Dataset
This dataset was created using LeRobot.
Language Annotations
This dataset includes annotated language instructions in meta/tasks_annotated.parquet. The file is indexed by episode_index and has a task column containing our per-episode annotated instruction.
The standard LeRobot loader resolves a frame's language instruction through task_index: each data row stores a task_index, which is looked up in meta/tasks.parquet. When you use these… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-DROID-Dataset.droid_1.0.1_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.droid_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DAVIAN-Robotics/droid_v3.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1.droid-3d-rgb-68epThis is a FiftyOne dataset with 68 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/droid-3d-rgb-68ep")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for droid_3d (68-episode FiftyOne RGB subset)
A… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/droid-3d-rgb-68ep.DroidCall
DroidCall: A Dataset for LLM-powered Android Intent Invocation
paper|github
DroidCall is the first open-sourced, high-quality dataset designed for fine-tuning LLMs for accurate intent invocation on Android devices.
This repo contains data generated by DroidCall. The process of data generation is shown in the figure below
Details can be found in our paper and github repository.
What is Android Intent Invocation?
Android Intent is a key machanism in Android that allows… See the full description on the dataset page: https://huggingface.co/datasets/mllmTeam/DroidCall.droid_120_stsg_bboxesdroid_1.0.1droiddroid_low_resolutiondroid-failure-sampled
DROID Robot Manipulation Dataset (Sampled)
数据集概述
这是从 DROID 1.0.1 数据集中采样的机器人操作失败案例子集。
总样本数: 2064
数据类型: failure
采样策略: balanced
视频格式: MP4, 60fps, 1280x720
数据集结构
hg_data/
├── videos/ # 视频文件
│ ├── 0000.mp4
│ ├── 0001.mp4
│ └── ...
├── metadata/ # 元数据文件
│ ├── 0000.json
│ ├── 0001.json
│ └── ...
├── dataset_info.json # 数据集总体信息
└── README.md # 本文件
任务类别分布
任务类别
数量
占比
Open a drawer and take some items out
4… See the full description on the dataset page: https://huggingface.co/datasets/JiaaqiLiu/droid-failure-sampled.droid_1.0.1
DROID 1.0.1 for pi05 training: successful, non-idle segments
lerobot/droid_1.0.1 rewritten twice, videos untouched. First
augment_droid_delta_ee_gripper_events.py (lerobot_policy_framepick) added action.delta_ee,
observation.extrinsics.static1/static2/wrist1 and the observation.gripper.time_* columns;
the flat action is the delta-EE command and the flat observation.state the EE state. Then
benchmarks/robolab/augment_droid_dataset.py (this revision) mirrored the data pipeline of… See the full description on the dataset page: https://huggingface.co/datasets/dgrachev/droid_1.0.1.droid_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 18327,
"total_frames": 3974430,
"total_tasks": 12747,
"total_videos": 128289,
"total_chunks": 19,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:18327"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/droid_lerobot.droid_dataset_segmentation_mask
DROID SAM 3.1 Segmentation Masks
This dataset is a mask-only sidecar generated from the original
droid_101/0.0.1 RLDS release. It does not redistribute DROID images or
actions. Its episode_index follows the RLDS episode order.
The same episodes appear in lerobot/droid_1.0.1, but LeRobot stores them in a
different episode order. Therefore, mask episode_index and LeRobot
episode_index must not be joined directly. Use the mapping file described
below to associate these masks with… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/droid_dataset_segmentation_mask.DroidCollection
Dataset Description
The dataset is structured into four primary classes:
Human-Written Code: Samples written entirely by humans.
AI-Generated Code: Samples generated by Large Language Models (LMs).
Machine-Refined Code: Samples representing a collaboration between humans and LMs, where human-written code is modified or extended by an AI.
AI-Generated-Adversarial Code: Samples generated by LMs with the specific intent to evade detection by mimicking human-like patterns and styles.… See the full description on the dataset page: https://huggingface.co/datasets/project-droid/DroidCollection.droid_1.0.1_v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/muacha/droid_1.0.1_v30.SemanticVLA-TraceX-240K-DROID
SemanticVLA TraceX 240K · DROID
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The DROID component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of DROID · Franka · Open-X-Embodiment DROID… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-DROID.droid_120_stsg_bboxes_v4droid_1.0.1_v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95584,
"total_frames": 27607757,
"total_tasks": 49596,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95584"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30.droid_1.0.1_v30_compact_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 10240,
"total_frames": 2988169,
"total_tasks": 6798,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:10240"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_3.droid_1.0.1_v30_compact_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_5.DROID_bench_alldroid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95617,
"total_frames": 27618651,
"total_tasks": 49611,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95617"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ygtxr1997/droid_1.0.1.droid_subsets_scooping
DROID Skill Subset: scooping
This is a skill-filtered subset of the local DROID 1.0.1 LeRobot-format dataset.
It was generated with droid_filter.py from /scratch/jellyho/prsl_skills.
Repository
Hugging Face repo: jellyho/droid_subsets_scooping
Source dataset path used locally: /scratch/jellyho/droid_1.0.1
Output subset path used locally: /scratch/jellyho/droid_subsets/scooping
Filtering
{
"skills": [
"scooping"
],
"skill_keywords": {… See the full description on the dataset page: https://huggingface.co/datasets/jellyho/droid_subsets_scooping.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ChanoYoung/droid_1.0.1.droid_subsets_picking
DROID Skill Subset: picking
This is a skill-filtered subset of the local DROID 1.0.1 LeRobot-format dataset.
It was generated with droid_filter.py from /scratch/jellyho/prsl_skills.
Repository
Hugging Face repo: jellyho/droid_subsets_picking
Source dataset path used locally: /scratch/jellyho/droid_1.0.1
Output subset path used locally: /scratch/jellyho/droid_subsets/picking
Filtering
{
"skills": [
"picking"
],
"skill_keywords": {… See the full description on the dataset page: https://huggingface.co/datasets/jellyho/droid_subsets_picking.droid_1.0.1_v30_compactThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95584,
"total_frames": 27607757,
"total_tasks": 49596,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95584"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact.
