Sharpa-Robotics/SharpaDex-v1.0
SharpaDex v1.0 A multimodal dataset of real-world teleoperated demonstrations for bimanual dexterous manipulation. Overview · Get Started · Tasks · Examples · Features · Language · Citation An animated overview of 56 tasks in a uniform 8 × 7 grid, with head-camera and wrist-camera observations distributed throughout the montage. It spans assembly, tool use, deformable-object manipulation, cleaning, material transfer, and long-horizon activities. Open MP4.… See the full description on the dataset page: https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0.
<h1 align="center">SharpaDex v1.0</h1>
<p align="center"> A multimodal dataset of real-world teleoperated demonstrations for bimanual dexterous manipulation. </p>
<p align="center"> <a href="https://www.sharpa.com/"><img src="https://img.shields.io/badge/Website-sharpa.com-555555?style=flat-square" alt="Website"></a> <a href="https://github.com/sharpa-robotics"><img src="https://img.shields.io/badge/GitHub-Sharpa-24292f?style=flat-square&logo=github&logoColor=white" alt="GitHub"></a> <a href="https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0"><img src="https://img.shields.io/badge/Dataset-Hugging%20Face-f3b000?style=flat-square" alt="Dataset"></a> <a href="https://creativecommons.org/licenses/by/4.0/"><img src="https://img.shields.io/badge/License-CC%20BY%204.0-2b7bb9?style=flat-square" alt="License"></a> </p>
<p align="center"> <a href="#overview">Overview</a> · <a href="#get-started">Get Started</a> · <a href="#task-collection">Tasks</a> · <a href="#example-observations">Examples</a> · <a href="#feature-schema">Features</a> · <a href="#language-annotations">Language</a> · <a href="#citation">Citation</a> </p>

An animated overview of 56 tasks in a uniform 8 × 7 grid, with head-camera and wrist-camera observations distributed throughout the montage. It spans assembly, tool use, deformable-object manipulation, cleaning, material transfer, and long-horizon activities. Open MP4.
Overview
SharpaDex v1.0 is a real-world teleoperation dataset for bimanual dexterous manipulation. It contains 28,993 demonstration episodes across 59 manipulation tasks, comprising 32,396,172 frames (approximately 300.0 hours at 30 FPS).
Data were collected through human teleoperation of a real bimanual robotic system with two 7-DoF arms, two 22-DoF dexterous hands, four RGB cameras, and fingertip tactile sensors. Each episode provides synchronized joint state, joint torque, tool-center-point state, observed and commanded TCP poses, action, numeric tactile measurements, RGB video, tactile video, and temporally aligned language annotations.
The task set covers object rearrangement, articulated-object interaction, assembly, tool use, deformable-object manipulation, packaging, cleaning, and long-horizon sequential manipulation. The data are distributed in LeRobot v3.0 and v2.1 formats for research on imitation learning, visuomotor control, vision-language-action models, visual-tactile learning, and hierarchical policy learning.
Dataset at a Glance
Version note: v3.0 and v2.1 are two exports of the same demonstrations. Dataset scale must be reported as 28,993 episodes and 32,396,172 frames, not the sum of both exports.
<details> <summary>Detailed dataset statistics</summary>
</details>
For new projects, we recommend starting with lerobot_v3.0. The lerobot_v2.1 export is provided for compatibility with pipelines that depend on the older LeRobot layout.
Example Observations
The previews below show synchronized observations from a real-robot teleoperation episode of deal_playing_cards (season POC22007_2026_03_26_15_55_47_train, episode 12, source interval 5–25 seconds). Each clip is displayed at 2x playback speed. The MP4 links remain available when animated GIF playback is disabled by the Markdown viewer.
Task Collection
Task directories use lowercase snake_case names with an action-object-target structure. Counts refer to one format version (v3.0); v2.1 contains the corresponding demonstrations.
<details> <summary>Complete task inventory — 59 tasks across 495 collection seasons</summary>
</details>
Get Started
Download the Dataset
Install Git LFS before cloning from Hugging Face.
git lfs install
git clone https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0To clone metadata first and fetch large files later:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0For a single task, use sparse checkout. This example downloads clean_plate_with_eraser:
git init SharpaDex-v1.0
cd SharpaDex-v1.0
git remote add origin https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0
git sparse-checkout init
git sparse-checkout set clean_plate_with_eraser README.md
git pull origin mainQuick Inspection
Inspect meta/info.json to discover the exact schema and path templates for an export.
import json
from pathlib import Path
dataset_root = Path("SharpaDex-v1.0")
episode_root = (
dataset_root
/ "clean_plate_with_eraser"
/ "season_POC22027_2026_04_17_11_01_59_train"
/ "lerobot_v3.0"
)
with open(episode_root / "meta" / "info.json", "r") as f:
info = json.load(f)
print("episodes:", info["total_episodes"])
print("frames:", info["total_frames"])
print("fps:", info["fps"])
print("features:", list(info["features"]))Dataset Structure
The repository uses a uniform task / season / format hierarchy. Every task is stored directly under the repository root.
<details> <summary>Repository layout</summary>
SharpaDex-v1.0/
├── README.md
├── clean_plate_with_eraser/
│ ├── season_POC22027_2026_04_17_11_01_59_train/
│ │ ├── lerobot_v3.0/
│ │ │ ├── meta/
│ │ │ │ ├── info.json
│ │ │ │ ├── modality.json
│ │ │ │ ├── episodes/
│ │ │ │ ├── tasks.parquet
│ │ │ │ └── subtasks.parquet
│ │ │ ├── data/
│ │ │ │ └── chunk-000/
│ │ │ └── videos/
│ │ │ ├── observation.images.head_left/
│ │ │ ├── observation.images.head_right/
│ │ │ ├── observation.images.wrist_left/
│ │ │ ├── observation.images.wrist_right/
│ │ │ ├── observation.images.tactile_deform/
│ │ │ └── observation.images.tactile_raw/
│ │ └── lerobot_v2.1/
│ │ ├── meta/
│ │ ├── data/
│ │ └── videos/
│ └── season_.../
├── cover_ball_with_cup/
├── reposition_and_stack_blocks/
├── play_tic_tac_toe/
└── .../</details>
Storage Layout
<details> <summary>LeRobot path templates</summary>
LeRobot v3.0 uses paths similar to:
data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4LeRobot v2.1 uses paths similar to:
data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet
videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4</details>
Feature Schema
All inspected seasons expose the same main feature set.
Joint-Space State and Action
The 65D state, joint-torque, and action vectors use the following order:
TCP Pose State and Action
observation.state.tcp_pose and action.tcp_pose are separate 12D float32 vectors. Components 0:6 contain the left-arm TCP pose and 6:12 contain the right-arm TCP pose.
The observation is read from the source state/left_arm/tcp_pose and state/right_arm/tcp_pose; the action is read from the corresponding action/.../tcp_pose streams. Each stream uses its own aligned_index. Source values and pose representation are preserved; commanded poses are not copied from observed poses or reconstructed using forward kinematics.
The existing 24D observation.state.tcp remains unchanged: left pose 0:6, left force/torque 6:12, right pose 12:18, right force/torque 18:24. Its two pose slices correspond to the new observed TCP-pose vector.
The export does not fully declare TCP units, reference frames or rotation convention (rotation_type is null). Use the acquisition/controller convention before geometric transformations; do not infer Euler angles or quaternions from the vector width alone.
Video Streams
Video files are MP4 without audio. Codec may differ between exports; inspect the local files if your training stack has codec restrictions.
Language Annotations
Every episode includes language annotations at two complementary levels: a global description of the complete task and temporally grounded descriptions of the subtasks performed during the trajectory. These annotations support language-conditioned policy learning, vision-language-action training, temporal grounding, skill discovery, hierarchical policy learning, and subtask-aware evaluation.
Global Task Description
The episode-level global_task stores the available structured natural-language fields below:
The same information is also exposed as separate fields under tags. All episodes have scene descriptions; 1,061 scene descriptions were supplemented from existing episode text and visual spot checks, with provenance recorded in the annotation metadata. Some of these episodes do not have separate instruction, success-criteria or SOP fields:
task_name
task_instruction
scene_description
success_criteria
score
SOPTemporally Grounded Subtasks
Each episode is decomposed into language segments. A segment contains:
Across the release, all 28,993 episodes contain language annotations. The dataset provides 208,264 temporal subtask segments spanning 28,897,608 frames, or approximately 89.2% of all frames. The segments use 45 nonempty skill labels; recovered descriptions may not have a separate skill label. Frames outside a fine-grained segment remain associated with the episode-level task description.
For example, a plate-cleaning trajectory is decomposed into phases such as:
Pick up the plate from the table with the left hand. | Skill: pick
Pick up the blackboard eraser with the right hand. | Skill: pick
Wipe away the dirty marks while the left hand holds the plate. | Skill: wipe
Place the plate and the eraser back on the table. | Skill: placeAnnotation Files and Frame Alignment
The v2.1 export stores task and subtask lookup tables as JSONL; v3.0 stores the corresponding tables as Parquet. annotations.jsonl is included with both exports for direct episode-level inspection.
The following example reads the global description and temporally grounded language from the first annotated episode:
import json
from pathlib import Path
meta_root = Path(
"SharpaDex-v1.0/clean_plate_with_eraser/"
"season_POC22027_2026_04_17_11_01_59_train/"
"lerobot_v3.0/meta"
)
with open(meta_root / "annotations.jsonl", "r") as f:
annotation = json.loads(next(f))
print(annotation["global_task"])
for segment in annotation["language_segments"]:
print(
segment["start_step"],
segment["end_step"],
segment["task_skill"],
segment["text"],
)When constructing training samples, use task_index for episode-level language conditioning and subtask_index or the explicit start_step / end_step intervals for phase-level conditioning.
Tactile Modality
The dataset provides tactile information in structured numeric and image-based forms:
observation.tactilecontains 60 values: left/rightthumb,index,middle,ring, andlittlefingertips, each withfx,fy,fz,tx,ty, andtz.observation.images.tactile_deformvisualizes contact-induced deformation patterns.observation.images.tactile_rawpreserves the raw tactile camera layout for custom preprocessing and representation learning.
All tactile modalities are synchronized with visual observations, proprioception, and actions at 30 FPS. The 60D signal provides a compact tactile representation; the tactile video streams provide spatially resolved observations for models designed to process the additional input resolution.
Usage Recommendations
For policy learning, a typical configuration is:
- Visual observations: one or more
observation.images.*streams - Proprioception:
observation.state - Optional force and contact inputs: joint torque, TCP state, and
observation.tactile - Optional tactile vision:
observation.images.tactile_deformand/orobservation.images.tactile_raw - Supervision target:
action, oraction.tcp_posefor a TCP-based action representation - Language conditioning: structured task instruction, scene description, and success criteria
- Phase conditioning: frame-aligned subtask text and skill label
For evaluation, split by collection season rather than randomly splitting frames. This reduces temporal leakage and avoids placing closely related demonstrations from the same collection session in both training and evaluation sets. For multi-task experiments, additionally report held-out tasks or task families when measuring cross-task generalization.
Dataset Notes
- The release is organized by task and collection season, not as one flattened LeRobot root.
- Every released season includes both
lerobot_v3.0andlerobot_v2.1. - Task descriptions, scene details, success criteria, and subtask decomposition may vary between seasons; use the metadata shipped with the selected season as the source of truth.
- Some long-horizon tasks contain multiple interaction phases and substantial hand-object occlusion.
- The two tactile video streams are high resolution and may dominate input bandwidth.
- The dataset contains demonstrations rather than a fixed benchmark split; users should document their task and season splits for reproducibility.
License and Terms
This dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You may share and adapt the dataset, including for commercial purposes, provided that you give appropriate attribution and indicate whether changes were made.
Citation
If this dataset contributes to your research, please cite or acknowledge the dataset:
@misc{sharpadex_v1_2026,
title = {SharpaDex v1.0: Large-Scale Bimanual Dexterous Manipulation Demonstrations},
author = {{Sharpa}},
howpublished = {\url{https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0}},
year = {2026}
}