Team Ai
Datasetpublic

hello3x3/zte_embodied_2026

ZTE Embodied 2026 Dataset (LeRobot v3.0) Overview This dataset is converted from the raw data of the 2026 16th ZTE Cup Global Elite Challenge - Algorithm Elite Challenge - Embodied Intelligence (Preliminary). Competition Link: ZTE Challenge - Embodied Intelligence Format: LeRobot v3.0 Conversion Script: scripts/datasets/convert_zte_to_lerobot.py Dataset Info Property Value Codebase Version v3.0 Robot Type Dual-arm Dexterous (dual… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/zte_embodied_2026.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes19downloads
Dataset Card

ZTE Embodied 2026 Dataset (LeRobot v3.0)

Overview

This dataset is converted from the raw data of the 2026 16th ZTE Cup Global Elite Challenge - Algorithm Elite Challenge - Embodied Intelligence (Preliminary).

Dataset Info

PropertyValue
Codebase Versionv3.0
Robot TypeDual-arm Dexterous (dual 7-DOF arms + dual 6-DOF hands)
Total Episodes350
Total Frames51,467
Total Tasks15
FPS30
Video Resolution1280 x 720
Video Codecav1 (libsvtav1)
Action Dimension26
State Dimension26

Splits

SplitEpisodesIndex RangeDescription
train2500:2505 groups x 50 episodes
test100250:3505 groups x 20 episodes

Features

observation.images.cam

  • —dtype: video
  • —shape: [720, 1280, 3]
  • —names: height, width, rgb
  • —info: codec=av1, pix_fmt=yuv420p, fps=30, channels=3

observation.state

Joint positions (26-dim float32) from joint.txt:

IndexNameDescription
0-6leftarmjoint1~7Left arm 7-DOF joint positions
7-13rightarmjoint1~7Right arm 7-DOF joint positions
14-19leftthumb0, leftthumb1, leftindex, leftmiddle, leftring, leftpinkyLeft hand 6-DOF finger positions
20-25rightthumb0, rightthumb1, rightindex, rightmiddle, rightring, rightpinkyRight hand 6-DOF finger positions

action

Action commands (26-dim float32) from action.txt, same structure as observation.state.

Metadata Features

FeaturedtypeDescription
timestampfloat32Frame timestamp (frame_index / fps)
frame_indexint64Index within episode
episode_indexint64Episode index
indexint64Global frame index
task_indexint64Task description index

Tasks

IndexDescription
0Pick up the Cocacola and place it into the box.
1Pick up the green tea and place it into the box.
2Pick up the Sprite and place it into the box.
3Pick up the mineral water from the box and place it on the table.
4Pick up the Fanta from the box and place it on the table.
5Pick up the Sprite from the box and place it on the table.
6Pick up the apple from the basket and pass it to me.
7Pick up the orange from the basket and place it on the table.
8Pick up the banana from the basket and place it on the table.
9Pick up the banana and place it into the box.
10Pick up the apple and place it into the box.
11Pick up the orange and place it into the box.
12Pick up the doll and place it into the box.
13Pick up the ball and place it into the box.
14Pick up the Rubik's Cube and place it into the box.

Directory Structure

zte_embodied_2026/
├── data/
│   └── chunk-000/
│       └── file-000.parquet          # Frame-level data (action, state, metadata)
├── meta/
│   ├── info.json                     # Dataset metadata and feature definitions
│   ├── stats.json                    # Global statistics (mean, std, min, max)
│   ├── tasks.parquet                 # Task description index
│   └── episodes/
│       └── chunk-000/
│           └── file-000.parquet      # Per-episode metadata and statistics
├── videos/
│   └── observation.images.cam/
│       └── chunk-000/
│           ├── file-000.mp4          # Concatenated video segments
│           ├── ...
│           └── file-007.mp4
└── README.md

Usage

Load with LeRobotDataset

python
from lerobot.datasets.lerobot_dataset import LeRobotDataset

ds = LeRobotDataset("zte_embodied_2026", root="./datasets/zte_embodied_2026")
print(f"Episodes: {ds.num_episodes}, Frames: {ds.num_frames}")

# Access a sample
sample = ds[0]
# sample["observation.images.cam"]  -> torch.Tensor [3, 720, 1280]
# sample["observation.state"]       -> torch.Tensor [26]
# sample["action"]                  -> torch.Tensor [26]
# sample["task"]                    -> str

Load with UniVAM DataLoader

In the JSONL config file, add:

json
{"repo_id": "./datasets/zte_embodied_2026", "dataset": "zte"}

The CAMERA_KEYS and ACTION_KEYS mappings for "zte" are already configured in src/univam/utils/dataloaders/lerobot_.py.

Raw Data Format (Before Conversion)

The original data was organized as follows:

release/
├── train/          # 250 episodes (5 groups x 50)
│   ├── 1_1/
│   │   ├── action.txt        # CSV: 26 columns (14 arm + 12 finger joints)
│   │   ├── joint.txt         # CSV: same structure as action.txt
│   │   ├── instruction.txt   # Natural language task description
│   │   ├── instruction.pt    # Encoded instruction embedding [1, 50, 4096]
│   │   └── video.mp4         # 30 FPS, 1280x720
│   └── ...
├── test/           # 100 episodes (5 groups x 20), 16 frames each
└── sample_result/  # Example submission format

Key differences from the converted format:

  • —instruction.pt (embedding tensor) is not included in the lerobot dataset
  • —joint.txt maps to observation.state
  • —action.txt maps to action
  • —video.mp4 frames are re-encoded as av1 and concatenated per chunk

Citation

If you use this dataset, please cite the original competition:

2026年第十六届中兴捧月全球精英挑战赛 - 神算师算法精英挑战赛 - 具身智能(初赛)
https://zte.uchallenge.cn/challenge/69638677c8440b6a6e14563c