vo2yager/aloha_incontext
aloha_incontext A Mobile ALOHA robot manipulation dataset for in-context imitation learning. It contains human-teleoperated demonstrations of pick-and-place, pen uncapping, placing eggs in a box and closing it, and additional bimanual tasks. 1,328 episodes / 31 task configurations / 587,000 frames, recorded at 50 Hz. The task configurations are divided into 25 seen configurations (1,318 episodes) and 6 unseen configurations (10 episodes). Observations and actions… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/aloha_incontext.
aloha_incontext
A Mobile ALOHA robot manipulation dataset for in-context imitation learning. It contains human-teleoperated demonstrations of pick-and-place, pen uncapping, placing eggs in a box and closing it, and additional bimanual tasks.
1,328 episodes / 31 task configurations / 587,000 frames, recorded at 50 Hz. The task configurations are divided into 25 seen configurations (1,318 episodes) and 6 unseen configurations (10 episodes).
Observations and actions
- Robot: Mobile ALOHA with two arms and 14 joint/gripper dimensions.
- RGB cameras:
cam_high,cam_left_wrist, andcam_right_wrist. - State:
observation.state, a 14-dimensional vector of joint and gripper positions. - Additional observations:
observation.velocityandobservation.effort. - Action: a 16-dimensional vector containing 14 joint/gripper targets and 2 base-velocity dimensions.
- Format: LeRobot v2.0, with one Parquet file per episode and camera images embedded in the Parquet data.
Tasks
The additional bimanual tasks are handover (104 episodes), cup_stack (50), stir (50), and water_wipe (50).
Pen uncapping
In every pen instruction, left/right identifies the hand that picks up the pen. The other hand grasps and removes the cap.
For example: "Pick up the gray pen with the right hand, grasp the cap with the other hand and uncap it."
For pick-and-place tasks, the hand named in the instruction picks up the object and places it in the basket.
Seen and unseen configurations
All episodes are packaged in a single Hugging Face train split. The seen/unseen division is defined by task configuration, not by separate Hugging Face splits.
The unseen configurations are:
For evaluation on unseen configurations, exclude these six configurations from training. This leaves 1,318 training episodes. Unseen demonstrations can be used as in-context examples at evaluation time.
File organization
data/chunk-*/episode_*.parquet: observations, actions, timestamps, and frame, episode, and task indices.meta/info.json: dataset schema and summary statistics.meta/tasks.jsonl: task IDs and task instructions.meta/episodes.jsonl: episode IDs, task assignments, and episode lengths.meta/stats.json: normalization statistics. Numeric-feature statistics cover all 1,328 episodes, including unseen configurations. Image normalization uses fixed per-channel statistics.
License
Apache-2.0.
