hello3x3/zte_embodied_2026
ZTE Embodied 2026 Dataset (LeRobot v3.0) Overview This dataset is converted from the raw data of the 2026 16th ZTE Cup Global Elite Challenge - Algorithm Elite Challenge - Embodied Intelligence (Preliminary). Competition Link: ZTE Challenge - Embodied Intelligence Format: LeRobot v3.0 Conversion Script: scripts/datasets/convert_zte_to_lerobot.py Dataset Info Property Value Codebase Version v3.0 Robot Type Dual-arm Dexterous (dual… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/zte_embodied_2026.
ZTE Embodied 2026 Dataset (LeRobot v3.0)
Overview
This dataset is converted from the raw data of the 2026 16th ZTE Cup Global Elite Challenge - Algorithm Elite Challenge - Embodied Intelligence (Preliminary).
- Competition Link: ZTE Challenge - Embodied Intelligence
- Format: LeRobot v3.0
- Conversion Script:
scripts/datasets/convert_zte_to_lerobot.py
Dataset Info
Splits
Features
observation.images.cam
- dtype: video
- shape: [720, 1280, 3]
- names: height, width, rgb
- info: codec=av1, pix_fmt=yuv420p, fps=30, channels=3
observation.state
Joint positions (26-dim float32) from joint.txt:
action
Action commands (26-dim float32) from action.txt, same structure as observation.state.
Metadata Features
Tasks
Directory Structure
zte_embodied_2026/
├── data/
│ └── chunk-000/
│ └── file-000.parquet # Frame-level data (action, state, metadata)
├── meta/
│ ├── info.json # Dataset metadata and feature definitions
│ ├── stats.json # Global statistics (mean, std, min, max)
│ ├── tasks.parquet # Task description index
│ └── episodes/
│ └── chunk-000/
│ └── file-000.parquet # Per-episode metadata and statistics
├── videos/
│ └── observation.images.cam/
│ └── chunk-000/
│ ├── file-000.mp4 # Concatenated video segments
│ ├── ...
│ └── file-007.mp4
└── README.mdUsage
Load with LeRobotDataset
from lerobot.datasets.lerobot_dataset import LeRobotDataset
ds = LeRobotDataset("zte_embodied_2026", root="./datasets/zte_embodied_2026")
print(f"Episodes: {ds.num_episodes}, Frames: {ds.num_frames}")
# Access a sample
sample = ds[0]
# sample["observation.images.cam"] -> torch.Tensor [3, 720, 1280]
# sample["observation.state"] -> torch.Tensor [26]
# sample["action"] -> torch.Tensor [26]
# sample["task"] -> strLoad with UniVAM DataLoader
In the JSONL config file, add:
{"repo_id": "./datasets/zte_embodied_2026", "dataset": "zte"}The CAMERA_KEYS and ACTION_KEYS mappings for "zte" are already configured in src/univam/utils/dataloaders/lerobot_.py.
Raw Data Format (Before Conversion)
The original data was organized as follows:
release/
├── train/ # 250 episodes (5 groups x 50)
│ ├── 1_1/
│ │ ├── action.txt # CSV: 26 columns (14 arm + 12 finger joints)
│ │ ├── joint.txt # CSV: same structure as action.txt
│ │ ├── instruction.txt # Natural language task description
│ │ ├── instruction.pt # Encoded instruction embedding [1, 50, 4096]
│ │ └── video.mp4 # 30 FPS, 1280x720
│ └── ...
├── test/ # 100 episodes (5 groups x 20), 16 frames each
└── sample_result/ # Example submission formatKey differences from the converted format:
instruction.pt(embedding tensor) is not included in the lerobot datasetjoint.txtmaps toobservation.stateaction.txtmaps toactionvideo.mp4frames are re-encoded as av1 and concatenated per chunk
Citation
If you use this dataset, please cite the original competition:
2026年第十六届中兴捧月全球精英挑战赛 - 神算师算法精英挑战赛 - 具身智能(初赛)
https://zte.uchallenge.cn/challenge/69638677c8440b6a6e14563c