ajaysri/pi07_cable_three_vector_v1
Three-holder cable routing with vector goals Real-robot demonstrations for a vector-conditioned low-level policy: place three holders and route a cable through each holder. This is the validated LeRobot v3 dataset prepared for the first Pi0.7 four-camera-goal cable policy. Property Value Source recordings 140 Subtask episodes 840 Frames 335,897 Sampling rate 100 Hz Robot ARX bimanual Recorded state / action 14 / 14 dimensions Video resolution 448 × 448… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/pi07_cable_three_vector_v1.
Three-holder cable routing with vector goals
Real-robot demonstrations for a vector-conditioned low-level policy: place three holders and route a cable through each holder. This is the validated LeRobot v3 dataset prepared for the first Pi0.7 four-camera-goal cable policy.
Episodes and subtasks
Each source recording becomes six episodes, in order:
- Place holder 1.
- Route the cable through holder 1.
- Place holder 2.
- Route the cable through holder 2.
- Place holder 3.
- Route the cable through holder 3.
137 recordings have six original phases. Three recordings have nine original phases; their separate pick phases are merged into the corresponding holder-placement episode. Original phase IDs and labels remain in per-frame metadata. Episode boundaries, rather than video-file boundaries, define training sequences.
There are two task types (placement and routing), distinguished by task_index and the stored language instruction. Holder identity and placement vectors are also retained. All 140 recordings belong to the training corpus; this release makes no held-out or novelty claim.
Observations and vector goals
Native 1280 × 800 images are resized to 448 × 280 and padded with 84 black pixels above and below. The goal overlay uses a yellow center dot (radius 5) and red direction arrow (length 28, width 4, head size 8) at the output resolution. Vector metadata includes native and resized center coordinates and a unit direction in image x/y coordinates.
The stored fourth video contains the annotated image at each timestamp. In the training loader, the goal image is sampled uniformly from the current subtask's start through the current frame, inclusive (goal seed 7). This past-goal sampling is a loader operation; simply reading the fourth video at the current timestamp does not reproduce it. The three physical camera inputs use the current frame.
observation.state, observation.velocity, and action preserve the recorded 14-dimensional arrays. Joint ordering and action units require a verified robot-specific deployment binding; they are not inferred here.
Action chunks and normalization
Stored actions are individual 14-dimensional commands. Training constructs 100-step action chunks and repeats the final action when a chunk reaches a subtask boundary. Chunks must not draw actions from the following subtask. metadata.valid_future_action_steps records the remaining valid portion of the chunk.
norm_stats.json contains the training normalization statistics, including q01/q99: state statistics use uniform frames, and action statistics account for uniform-frame 100-step chunks with repeated tails. The model consumer pads actions and state to 32 dimensions.
The associated training recipe uses train-time RTC with a uniformly sampled delay of 0–20 steps and checkpoint intervals of 4,000 steps. These settings describe the consuming training pipeline; this repository contains the dataset, not model weights.
Files and loading
data/chunk-000/*.parquet: frame records, including state, velocity, action, language, vector annotations, source indices and phase labels.videos/<video_key>/chunk-000/*.mp4: four synchronized video streams per source recording, with no audio.meta/info.json: LeRobot features, counts and file templates.meta/episodes/chunk-000/file-000.parquet: subtask episode boundaries and video offsets.meta/tasks.parquet: the two task types.contract.json: representation, derivation, training contract and SHA-256 artifact manifest.norm_stats.json: normalization used for training.
Six subtask episodes share each source recording's data file and videos. A LeRobot v3-compatible reader must use the episode metadata's video file references and from_timestamp / to_timestamp offsets. Per-episode frame_index and timestamp restart at zero; these alone are not absolute offsets into a shared video.
Download the native layout:
from huggingface_hub import snapshot_download
root = snapshot_download(
repo_id="ajaysri/pi07_cable_three_vector_v1",
repo_type="dataset",
)Use a fixed Hub commit via revision= when reproducing a run. The Hugging Face table viewer displays frame records; videos are separate files and are not embedded in those rows. The preview above illustrates the visual inputs.
Integrity
This public upload preserves the validated training artifacts byte-for-byte. The card and preview are additional documentation.
- Dataset content revision:
821d0d27d4e1b5eeb41a65dc72adf3ec72b74dee14d0d31c1f90973a4efead62 - Source content revision:
ea6d9dd01f4acfd72369a149c828ec3b96e5e226028221bfec6c9a9e73061cf9 - Contract SHA-256:
9b9f047e79a3021bca6e70a26312c3582c69eb540f36b4c6b68fa154a2d2d7ba - Normalization SHA-256:
c7571a4e9b0316bc64d0a13f900f03e31c068ccdbfaac1c104c6f2a39c30dfa1
The dataset content revision identifies the manifest in contract.json; the Hub commit separately identifies this published snapshot.
