SidneyXie/pi05_robotwin
π₀.₅ for RoboTwin 2.0
This is a LeRobot π₀.₅ Vision-Language-Action policy fine-tuned from `lerobot/pi05_base` on the `lerobot/robotwin_unified` dataset. It predicts joint-space action chunks for the 14-DoF Aloha-AgileX bimanual robot used by RoboTwin 2.0.
The checkpoint in this repository is the final checkpoint from training step 50,000. The model consumes language instructions, one 14-dimensional robot state, and three RGB camera views.
Model details
Inputs and outputs
The feature names below are part of the checkpoint configuration. Camera names must match these keys, or be mapped to them with --rename_map.
Inputs
Output
The 14 state/action dimensions are ordered as follows:
left_waist, left_shoulder, left_elbow, left_forearm_roll,
left_wrist_angle, left_wrist_rotate, left_gripper,
right_waist, right_shoulder, right_elbow, right_forearm_roll,
right_wrist_angle, right_wrist_rotate, right_gripperTraining data
The policy was trained on lerobot/robotwin_unified in LeRobot v3.0 format. The dataset metadata available for this training run reports:
- 27,500 episodes
- 6,075,103 frames
- 30 FPS
- three RGB views: high, left wrist, and right wrist
- 14-dimensional Aloha joint state and action
RoboTwin 2.0 covers 50 bimanual manipulation tasks with varied objects, layouts, lighting, backgrounds, and language instructions.
Training configuration
The saved preprocessor and postprocessor files contain the normalization state needed for inference; upload them together with model.safetensors and config.json.
Installation
Install LeRobot with the π policy dependencies:
pip install "lerobot[pi]"RoboTwin evaluation additionally requires Linux, an NVIDIA GPU, and the RoboTwin SAPIEN/CuRobo environment. See the LeRobot RoboTwin guide for simulator setup.
Loading the policy
The checkpoint includes serialized pre- and postprocessing pipelines. Load all three components from the same Hub repository:
import torch
from lerobot.policies import make_pre_post_processors
from lerobot.policies.pi05 import PI05Policy
model_id = "SidneyXie/pi05_robotwin"
device = "cuda"
policy = PI05Policy.from_pretrained(model_id)
policy.eval()
preprocessor, postprocessor = make_pre_post_processors(
policy.config,
pretrained_path=model_id,
preprocessor_overrides={"device_processor": {"device": device}},
)At inference time, pass a batch containing the three image features, observation.state, and a natural-language task. Use policy.select_action(preprocessor(batch)), then apply postprocessor to the result before sending it to the robot or simulator.
RoboTwin evaluation
RoboTwin's environment camera keys differ from the names stored in this checkpoint, so the rename map is required. A quick five-episode evaluation on one task is:
lerobot-eval \
--policy.path=SidneyXie/pi05_robotwin \
--env.type=robotwin \
--env.task=beat_block_hammer \
--eval.batch_size=1 \
--eval.n_episodes=5 \
--rename_map='{"observation.images.head_camera":"observation.images.cam_high","observation.images.left_camera":"observation.images.cam_left_wrist","observation.images.right_camera":"observation.images.cam_right_wrist"}' \
--output_dir=outputs/eval/pi05_robotwin/beat_block_hammerFor an official-style result, evaluate 100 episodes per task and report Easy (demo_clean) and Hard (demo_randomized) settings separately. Consult the RoboTwin leaderboard for the current submission protocol.
Evaluation results
Intended use and limitations
This checkpoint is intended for research on RoboTwin 2.0 and compatible 14-DoF Aloha-style bimanual setups. It expects the same joint ordering, camera semantics, observation preprocessing, and action convention used during training.
- It has not been validated for direct deployment on a physical robot.
- Distribution shifts in camera placement, calibration, control frequency, joint scaling, objects, or scene appearance can substantially reduce performance.
- The policy can produce unsafe or infeasible actions. Use workspace limits, collision checking, emergency stops, and human supervision on real hardware.
- This is a learned policy and does not provide correctness or safety guarantees.
References
@article{intelligence2025pi05,
title = {Pi 0.5: a Vision-Language-Action Model with Open-World Generalization},
author = {Physical Intelligence and Kevin Black and Noah Brown and others},
journal = {arXiv preprint arXiv:2504.16054},
year = {2025}
}
@misc{cadene2024lerobot,
title = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in PyTorch},
author = {Cadene, Remi and Alibert, Simon and Soare, Alexander and others},
year = {2024},
howpublished = {\url{https://github.com/huggingface/lerobot}}
}