Team Ai
Modelpublic

SidneyXie/pi05_robotwin

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
4likes652downloads
Model Card

π₀.₅ for RoboTwin 2.0

This is a LeRobot π₀.₅ Vision-Language-Action policy fine-tuned from `lerobot/pi05_base` on the `lerobot/robotwin_unified` dataset. It predicts joint-space action chunks for the 14-DoF Aloha-AgileX bimanual robot used by RoboTwin 2.0.

The checkpoint in this repository is the final checkpoint from training step 50,000. The model consumes language instructions, one 14-dimensional robot state, and three RGB camera views.

Model details

PropertyValue
Policyπ₀.₅ (pi05)
Fine-tuned fromlerobot/pi05_base
FrameworkLeRobot / PyTorch
RobotAloha-AgileX bimanual, 14 DoF
Action representationAbsolute joint-space actions
Action chunk size50
Actions executed per prediction50
Inference steps10
Internal image resolution224 × 224
Model dtypebfloat16
LicenseApache-2.0

Inputs and outputs

The feature names below are part of the checkpoint configuration. Camera names must match these keys, or be mapped to them with --rename_map.

Inputs

FeatureTypeShape
observation.stateState(14,)
observation.images.cam_highRGB image(480, 640, 3)
observation.images.cam_left_wristRGB image(480, 640, 3)
observation.images.cam_right_wristRGB image(480, 640, 3)
taskNatural-language instructionstring

Output

FeatureTypeShape
actionAbsolute joint-space action(14,) per step, 50-step chunks

The 14 state/action dimensions are ordered as follows:

text
left_waist, left_shoulder, left_elbow, left_forearm_roll,
left_wrist_angle, left_wrist_rotate, left_gripper,
right_waist, right_shoulder, right_elbow, right_forearm_roll,
right_wrist_angle, right_wrist_rotate, right_gripper

Training data

The policy was trained on lerobot/robotwin_unified in LeRobot v3.0 format. The dataset metadata available for this training run reports:

  • —27,500 episodes
  • —6,075,103 frames
  • —30 FPS
  • —three RGB views: high, left wrist, and right wrist
  • —14-dimensional Aloha joint state and action

RoboTwin 2.0 covers 50 bimanual manipulation tasks with varied objects, layouts, lighting, backgrounds, and language instructions.

Training configuration

SettingValue
Training steps50,000
Batch size16
OptimizerAdamW
Peak learning rate2.5e-5
Weight decay0.01
LR scheduleCosine decay with 1,000 warmup steps
Decay steps / final LR30,000 / 2.5e-6
Seed1000
Gradient checkpointingEnabled
torch.compileEnabled (max-autotune)
Vision encoder frozenNo
Expert-only trainingNo
Image augmentationDisabled
State/action normalizationMean and standard deviation

The saved preprocessor and postprocessor files contain the normalization state needed for inference; upload them together with model.safetensors and config.json.

Installation

Install LeRobot with the π policy dependencies:

bash
pip install "lerobot[pi]"

RoboTwin evaluation additionally requires Linux, an NVIDIA GPU, and the RoboTwin SAPIEN/CuRobo environment. See the LeRobot RoboTwin guide for simulator setup.

Loading the policy

The checkpoint includes serialized pre- and postprocessing pipelines. Load all three components from the same Hub repository:

python
import torch

from lerobot.policies import make_pre_post_processors
from lerobot.policies.pi05 import PI05Policy

model_id = "SidneyXie/pi05_robotwin"
device = "cuda"

policy = PI05Policy.from_pretrained(model_id)
policy.eval()

preprocessor, postprocessor = make_pre_post_processors(
    policy.config,
    pretrained_path=model_id,
    preprocessor_overrides={"device_processor": {"device": device}},
)

At inference time, pass a batch containing the three image features, observation.state, and a natural-language task. Use policy.select_action(preprocessor(batch)), then apply postprocessor to the result before sending it to the robot or simulator.

RoboTwin evaluation

RoboTwin's environment camera keys differ from the names stored in this checkpoint, so the rename map is required. A quick five-episode evaluation on one task is:

bash
lerobot-eval \
  --policy.path=SidneyXie/pi05_robotwin \
  --env.type=robotwin \
  --env.task=beat_block_hammer \
  --eval.batch_size=1 \
  --eval.n_episodes=5 \
  --rename_map='{"observation.images.head_camera":"observation.images.cam_high","observation.images.left_camera":"observation.images.cam_left_wrist","observation.images.right_camera":"observation.images.cam_right_wrist"}' \
  --output_dir=outputs/eval/pi05_robotwin/beat_block_hammer

For an official-style result, evaluate 100 episodes per task and report Easy (demo_clean) and Hard (demo_randomized) settings separately. Consult the RoboTwin leaderboard for the current submission protocol.

Evaluation results

SettingTaskEpisodesSuccess rate
Easyadjust_bottle100100%
Easybeat_block_hammer10093%
Easyclick_alarmclock10090%
Easyclick_bell10086%
Easydump_bin_bigbin10096%
Easygrab_roller10099%
Easyhandover_mic10023%
Easylift_pot10057%
Easymove_can_pot10048%
Easymove_pillbottle_pad10074%
Easymove_playingcard_away10096%
Easymove_stapler_pad10014%
Easypick_diverse_bottles10059%
Easypick_dual_bottles10069%
Easyplace_a2b_left10049%
Easyplace_a2b_right10041%
Easyplace_bread_skillet10060%
Easyplace_burger_fries10096%
Easyplace_container_plate10086%
Easyplace_dual_shoes10072%
Easyplace_empty_cup10090%
Easyplace_fan10050%
Easyplace_mouse_pad10045%
Easyplace_object_scale10046%
Easyplace_object_stand10090%
Easyplace_phone_stand10066%
Easyplace_shoe10092%
Easypress_stapler10097%
Easyrotate_qrcode10071%
Easyscan_object10018%
Easystamp_seal10045%
Easyturn_switch10055%
EasyOverall (32 tasks)320067.91%

Intended use and limitations

This checkpoint is intended for research on RoboTwin 2.0 and compatible 14-DoF Aloha-style bimanual setups. It expects the same joint ordering, camera semantics, observation preprocessing, and action convention used during training.

  • —It has not been validated for direct deployment on a physical robot.
  • —Distribution shifts in camera placement, calibration, control frequency, joint scaling, objects, or scene appearance can substantially reduce performance.
  • —The policy can produce unsafe or infeasible actions. Use workspace limits, collision checking, emergency stops, and human supervision on real hardware.
  • —This is a learned policy and does not provide correctness or safety guarantees.

References

bibtex
@article{intelligence2025pi05,
  title   = {Pi 0.5: a Vision-Language-Action Model with Open-World Generalization},
  author  = {Physical Intelligence and Kevin Black and Noah Brown and others},
  journal = {arXiv preprint arXiv:2504.16054},
  year    = {2025}
}

@misc{cadene2024lerobot,
  title        = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in PyTorch},
  author       = {Cadene, Remi and Alibert, Simon and Soare, Alexander and others},
  year         = {2024},
  howpublished = {\url{https://github.com/huggingface/lerobot}}
}