datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoboTwin2.0robotwin-clean-and-aug-lerobot
Robotwin Dataset
Robotwin dataset in Lerobot format, with video latents already extracted in WAN 2.2 format, ready for use in Lingbot-VA post-training.
License Agreement
This project is licensed under the CC BY-NC-SA 4.0.
robotwin-random-47-500
RoboTwin Randomized 47 Tasks, 500 Episodes Each
This dataset contains a 500-episode subset for each of 47 randomized RoboTwin
tasks. Episodes 0 through 499 were selected from each task.
Structure
<task>-demo_randomized-1000/
episode_<n>/
episode_<n>.hdf5
instructions.json
Each HDF5 episode contains robot actions, joint positions, arm metadata, and
three encoded camera streams:
action
observations/qpos
observations/left_arm_dim… See the full description on the dataset page: https://huggingface.co/datasets/sunLry/robotwin-random-47-500.robotwin_3d
RoboTwin 2.0 — 3D (RGB + Depth)
Bimanual manipulation data from the RoboTwin 2.0 simulator, in LeRobot v2.1 format, with per-camera ground-truth depth alongside RGB.
Tasks
50
Episodes
27,500 (550 per task, contiguous)
Frames
6,183,813
Robot
ALOHA-style bimanual, 14-DoF
Control rate
50 Hz
Cameras
3 (cam_high, cam_left_wrist, cam_right_wrist)
Resolution
240 × 320
Language instructions
1,039,891 unique corpus-wide; 100 entries per episode
Total size… See the full description on the dataset page: https://huggingface.co/datasets/flex-pi/robotwin_3d.RoboTwin2.0-Depthrobotwin
RoboTwin Bimanual Robotic Manipulation Platform
Lastest Version: RoboTwin 2.0🤲 Webpage | Document | Paper | Community | Leaderboard
https://private-user-images.githubusercontent.com/88101805/463126988-e3ba1575-4411-4a36-ad65-f0b2f49890c3.mp4
[2.0 Version (lastest)] RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationUnder Review 2025: Webpage | Document | PDF | arXiv | Talk (in Chinese) | 机器之心 | Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/sdvkasc/robotwin.robotwin-prefvlarobotwin_unifiedrobotwin_aloha_clean_hdf5without depth data. processed data. generated from python scripts/process_data.py $task_name $setting $expert_data_num
X-WAM-RoboTwin
X-WAM
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
Dataset Summary
This is the RoboTwin 2.0 fine-tuning dataset used to train the X-WAM unified 4D World Action Model. It packages dual-arm bimanual manipulation demonstrations into a unified multi-view RGB-D video + low-dimensional state/action format, where each episode provides synchronized RGB videos, depth videos, dual-arm end-effector proprioception, actions, and a… See the full description on the dataset page: https://huggingface.co/datasets/sharinka0715/X-WAM-RoboTwin.RoboTwin-Randomizedrobotwin2.0-fastwam
RobotWin 2.0 (Preprocessed LeRobot v2.1 Release)
This repository releases our preprocessed RoboTwin / RobotWin 2.0 dataset in LeRobot v2.1 format for the open-source release of Fast-WAM: Do World Action Models Need Test-time Future Imagination?
This is not the official upstream RoboTwin release. It is our paper-specific processed version prepared to support training, evaluation, and reproducibility for our project.
This Hugging Face repository distributes the dataset as split… See the full description on the dataset page: https://huggingface.co/datasets/yuanty/robotwin2.0-fastwam.RoboTwin-Cleanrobotwin_3d_text_embeds_cache
RoboTwin 2.0 3D — T5 Text Embedding Cache
Precomputed UMT5-XXL text embeddings for the
1,039,891 unique task prompts of the RoboTwin 2.0 3D dataset
(flex-pi/robotwin_3d), as consumed by
Wan2.2-TI2V-5B. Precomputing these costs substantial GPU time; this cache skips it.
Cache key
Task strings from meta/tasks.jsonl are wrapped in a fixed template before encoding, and the cache
key is the sha256 of that templated prompt:
DEFAULT_PROMPT = "A video recorded from a… See the full description on the dataset page: https://huggingface.co/datasets/flex-pi/robotwin_3d_text_embeds_cache.RoboTwin-MeMrobotwin-icl-paired-v3
RoboTwin ICL-paired v3
Cross-embodiment paired dataset on RoboTwin: 25 manipulation tasks rendered for 3 robots (arx-x5, franka, ur5) from a shared seed pool, so episode index i corresponds to the same scene/trajectory across all three robots.
Status (snapshot 2026-05-03)
13 tasks: 150 train + 50 val per robot.
scan_object click_bell place_empty_cup press_stapler click_alarmclock
place_container_plate stamp_seal move_pillbottle_pad place_fan
place_bread_skillet lift_pot… See the full description on the dataset page: https://huggingface.co/datasets/yeeeiii111/robotwin-icl-paired-v3.robotwin-failure-recoveryrobotwin2_rldsRobotwinrobotwin_3d_missingRoboTwin-LeRobot-unseen-tasks-cross-embrobotwin_merged
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
This repository contains the preprocessed dataset used in the paper LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies.
Project Page: https://rlinf.github.io/LaWAM/
Repository: https://github.com/RLinf/LaWAM
The dataset is formatted in LeRobot format and is designed for training and evaluating dynamics-aware robot policies.
Citation
@misc{chen2026lawam… See the full description on the dataset page: https://huggingface.co/datasets/jialei02/robotwin_merged.robotwin_scan_object_place_dual_shoes_serialized_200robotwin2.0_v30
RoboTwin 2.0 v3.0 + FastWAM VAE latent cache
Author: Rui Heng Yang
RoboTwin 2.0 dataset in LeRobot v3.0 format (data/, meta/, videos/; 27,500
episodes, robot type aloha), together with the FastWAM Wan2.2 bf16 VAE latent
cache in fastwam_vae_cache_wan22_bf16/.
Rejoining the latent cache
latents.bf16 (831,040,128,000 bytes) exceeds the Hub's 500GB per-file limit,
so it is stored as 100GiB parts. After downloading, rejoin with:
cat… See the full description on the dataset page: https://huggingface.co/datasets/Zhanguang/robotwin2.0_v30.RoboTwin-LeRobotrobotwin2.0-lerobot-v3.0
RoboTwin 2.0 LeRobot v3.0
This dataset packages RoboTwin 2.0 bimanual-manipulation demonstrations in LeRobot v3.0 format for EasyWAM training. Clean and randomized demonstrations are stored as two independent datasets so they can be trained separately or combined. Both contain synchronized high-view and dual-wrist RGB video, bimanual joint state, actions, and natural-language task metadata.
Dataset Summary
Split
Episodes
Frames
Videos
Recording rate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/robotwin2.0-lerobot-v3.0.WorldArena_Robotwin2.0
RoboTwin Embodied Video Generation Dataset for WorldArena
This dataset is designed for embodied video generation and evaluation across two main leaderboards and an interactive arena of WorldArena.
0) Dataset Overview
Leaderboard (test_dataset): Evaluation set for Leaderboard.Extract the directory from test_dataset.tar.gz
Arena (val_dataset): Used for the Arena (interactive comparison). This set allows users to upload their own generated videos for a specific… See the full description on the dataset page: https://huggingface.co/datasets/WorldArena/WorldArena_Robotwin2.0.RoboTwin-LeRobot-v3.0robotwin2.0-egowam-flow-cam_high-subset
RoboTwin 2.0 cam_high EgoWAM-style 3D Flow — snapshot subset (8,522 episodes)
Query-based 3D motion-flow sidecar labels for the RoboTwin 2.0 LeRobot dataset
(yuanty/robotwin2.0-fastwam), generated with a pretrained 3D point tracker
(Track4World, DA3 backbone, metric-scale mode) from RGB only — EgoWAM-style
(arXiv 2607.08436 §4.2 conventions). This is a frozen snapshot of an in-progress
full-dataset run (27,500 episodes); see snapshot_manifest.json for the exact
episode list and… See the full description on the dataset page: https://huggingface.co/datasets/kimtaey/robotwin2.0-egowam-flow-cam_high-subset.robotwin-2.0-480p
RoboTwin 2.0 Clean-50 at 480p
This dataset contains 2,500 rerendered RoboTwin ALOHA-AgileX demonstrations:
50 tasks with 50 episodes per task. The head and left and right wrist cameras
use native 640×480 rendering. The auxiliary front camera retains its native
320×240 resolution. Images are JPEG-encoded; they are not upscaled.
The source demonstrations come from
TianxingChen/RoboTwin2.0.
Rerendering uses RoboTwin revision 6dde57155eafa3e4ebf6ad1f93a7cf7d5d41a755
and the… See the full description on the dataset page: https://huggingface.co/datasets/sai-prasanna/robotwin-2.0-480p.
