Team Ai
Datasetpublic

Sharpa-Robotics/SharpaDex-v1.0

SharpaDex v1.0 A multimodal dataset of real-world teleoperated demonstrations for bimanual dexterous manipulation. Overview · Get Started · Tasks · Examples · Features · Language · Citation An animated overview of 56 tasks in a uniform 8 × 7 grid, with head-camera and wrist-camera observations distributed throughout the montage. It spans assembly, tool use, deformable-object manipulation, cleaning, material transfer, and long-horizon activities. Open MP4.… See the full description on the dataset page: https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0.

sourceHugging Facecc-by-4.0updated 9d agoView on Hugging Face
2likes3.9kdownloads
Dataset Card

<h1 align="center">SharpaDex v1.0</h1>

<p align="center"> A multimodal dataset of real-world teleoperated demonstrations for bimanual dexterous manipulation. </p>

<p align="center"> <a href="https://www.sharpa.com/"><img src="https://img.shields.io/badge/Website-sharpa.com-555555?style=flat-square" alt="Website"></a> <a href="https://github.com/sharpa-robotics"><img src="https://img.shields.io/badge/GitHub-Sharpa-24292f?style=flat-square&logo=github&logoColor=white" alt="GitHub"></a> <a href="https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0"><img src="https://img.shields.io/badge/Dataset-Hugging%20Face-f3b000?style=flat-square" alt="Dataset"></a> <a href="https://creativecommons.org/licenses/by/4.0/"><img src="https://img.shields.io/badge/License-CC%20BY%204.0-2b7bb9?style=flat-square" alt="License"></a> </p>

<p align="center"> <a href="#overview">Overview</a> · <a href="#get-started">Get Started</a> · <a href="#task-collection">Tasks</a> · <a href="#example-observations">Examples</a> · <a href="#feature-schema">Features</a> · <a href="#language-annotations">Language</a> · <a href="#citation">Citation</a> </p>

![Overview of diverse manipulation tasks from head and wrist cameras](assets/taskoverview56tasksgrid.mp4)

An animated overview of 56 tasks in a uniform 8 × 7 grid, with head-camera and wrist-camera observations distributed throughout the montage. It spans assembly, tool use, deformable-object manipulation, cleaning, material transfer, and long-horizon activities. Open MP4.

Overview

SharpaDex v1.0 is a real-world teleoperation dataset for bimanual dexterous manipulation. It contains 28,993 demonstration episodes across 59 manipulation tasks, comprising 32,396,172 frames (approximately 300.0 hours at 30 FPS).

Data were collected through human teleoperation of a real bimanual robotic system with two 7-DoF arms, two 22-DoF dexterous hands, four RGB cameras, and fingertip tactile sensors. Each episode provides synchronized joint state, joint torque, tool-center-point state, observed and commanded TCP poses, action, numeric tactile measurements, RGB video, tactile video, and temporally aligned language annotations.

The task set covers object rearrangement, articulated-object interaction, assembly, tool use, deformable-object manipulation, packaging, cleaning, and long-horizon sequential manipulation. The data are distributed in LeRobot v3.0 and v2.1 formats for research on imitation learning, visuomotor control, vision-language-action models, visual-tactile learning, and hierarchical policy learning.

Dataset at a Glance

TasksEpisodesDurationFrequencyVideo streamsLanguage segments
5928,993300.0 hours30 FPS6208,264
ModalityContent
Robot65D joint state, 65D action, 65D joint torque, 24D TCP state, and separate 12D observed / commanded TCP poses
VisionTwo head cameras and two wrist cameras
Tactile60D force/torque signal, deformation video, and raw tactile video
LanguageStructured task descriptions, frame-aligned subtasks, and 45 skill labels
FormatLeRobot v3.0 and v2.1 for every released season
Version note: v3.0 and v2.1 are two exports of the same demonstrations. Dataset scale must be reported as 28,993 episodes and 32,396,172 frames, not the sum of both exports.

<details> <summary>Detailed dataset statistics</summary>

ItemValue
Manipulation tasks59
Collection seasons495
Episodes (one export)28,993
Frames (one export)32,396,172
Approximate duration at 30 FPS300.0 hours
LeRobot v3.0 seasons / episodes495 / 28,993
LeRobot v2.1 seasons / episodes495 / 28,993
FPS30
Synchronized video streams6
State / action dimension65
Joint-torque dimension65
TCP-state dimension24
Observed / commanded TCP-pose dimension12 / 12
Tactile-signal dimension60
Episodes with global language annotations28,993 (100%)
Temporally grounded language segments208,264
Frames covered by temporal language segments28,897,608 (89.2%)
Distinct skill labels45

</details>

For new projects, we recommend starting with lerobot_v3.0. The lerobot_v2.1 export is provided for compatibility with pipelines that depend on the older LeRobot layout.

Example Observations

The previews below show synchronized observations from a real-robot teleoperation episode of deal_playing_cards (season POC22007_2026_03_26_15_55_47_train, episode 12, source interval 5–25 seconds). Each clip is displayed at 2x playback speed. The MP4 links remain available when animated GIF playback is disabled by the Markdown viewer.

Head camera ([MP4](assets/example_views/head_left.mp4))Wrist camera ([MP4](assets/example_views/wrist_right.mp4))
![Head-left camera observation](assets/exampleviews/headleft.mp4)![Right-wrist camera observation](assets/exampleviews/wristright.mp4)
Tactile deformation ([MP4](assets/example_views/tactile_deform.mp4))Raw tactile observation ([MP4](assets/example_views/tactile_raw.mp4))
![Tactile deformation observation](assets/exampleviews/tactiledeform.mp4)![Raw tactile observation](assets/exampleviews/tactileraw.mp4)

Task Collection

Task directories use lowercase snake_case names with an action-object-target structure. Counts refer to one format version (v3.0); v2.1 contains the corresponding demonstrations.

<details> <summary>Complete task inventory — 59 tasks across 495 collection seasons</summary>

TaskSeasonsEpisodesFrames
insert_batteries_into_charger3458841,645
scrub_cup_interior2481314,401
insert_plug_into_socket6416183,726
assemble_gears_on_base111,488983,123
clean_plate_with_eraser5343190,467
tighten_bottle_cap7131139,693
secure_coiled_cable_with_velcro_tie10473632,425
collect_waste_into_bin8508621,140
deal_playing_cards231,0882,074,644
empty_dustpan_into_bin10488379,585
cover_ball_with_cup5488188,956
insert_knife_into_cutlery_bin4495232,113
place_ball_in_box311540,672
place_ball_in_cup_then_cup_in_box4456252,933
place_cup_on_plate4429149,882
place_knife_and_fork_on_plate8461314,324
stack_two_plates4491185,481
remove_block_from_box5388142,321
reorient_cup_upright5420151,573
fit_trash_bag_into_bin47656852,443
turn_book_pages76587,056
fold_and_close_box75371,008,242
hang_garment_on_hanger15521979,299
form_kraft_paper_box5455775,252
close_kraft_paper_box5645693,683
fold_towel3459547,201
grind_medicine_with_mortar_and_pestle4131176,595
hammer_nails41,071861,668
insert_batteries_into_device7516456,513
iron_garment335172,277,751
clean_garment_with_lint_roller3810737,849
stack_napkins_in_holder5497545,816
nest_cups10464763,601
organize_objects_in_drawer595115,637
pack_blueberries_in_bag174851,533,417
pack_object_in_storage_bag9311644,006
staple_paper9695605,253
relocate_tennis_ball1450157,112
place_fruits_in_basket8416286,988
place_pens_in_holder11399427,274
place_coffee_filter_in_dripper217841,786,030
insert_socket_module_into_board2957512,619
transfer_pills_between_cups9479239,561
pry_nails_from_board32421,174
reorient_and_relocate_carton2347181,309
transfer_liquid_with_dropper12360655,829
transfer_salt_between_cups_with_spoon41,483831,390
scoop_salt_from_jar_into_cup13424485,424
sort_utensils5419460,851
reposition_and_stack_blocks151,0071,069,334
sweep_waste_into_dustpan10609516,260
seal_box_with_tape1115208,648
place_books_on_shelf9151133,230
collect_toys_into_basket58486,296
rotate_bottle_cap17583497,695
unscrew_bottle_cap3445351,637
clean_knife_with_cloth2292246,287
play_tic_tac_toe165211,317,513
tomato_boxed_lunch_packaging497243,325

</details>

Get Started

Download the Dataset

Install Git LFS before cloning from Hugging Face.

bash
git lfs install
git clone https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0

To clone metadata first and fetch large files later:

bash
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0

For a single task, use sparse checkout. This example downloads clean_plate_with_eraser:

bash
git init SharpaDex-v1.0
cd SharpaDex-v1.0
git remote add origin https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0
git sparse-checkout init
git sparse-checkout set clean_plate_with_eraser README.md
git pull origin main

Quick Inspection

Inspect meta/info.json to discover the exact schema and path templates for an export.

python
import json
from pathlib import Path

dataset_root = Path("SharpaDex-v1.0")
episode_root = (
    dataset_root
    / "clean_plate_with_eraser"
    / "season_POC22027_2026_04_17_11_01_59_train"
    / "lerobot_v3.0"
)

with open(episode_root / "meta" / "info.json", "r") as f:
    info = json.load(f)

print("episodes:", info["total_episodes"])
print("frames:", info["total_frames"])
print("fps:", info["fps"])
print("features:", list(info["features"]))

Dataset Structure

The repository uses a uniform task / season / format hierarchy. Every task is stored directly under the repository root.

<details> <summary>Repository layout</summary>

text
SharpaDex-v1.0/
├── README.md
├── clean_plate_with_eraser/
│   ├── season_POC22027_2026_04_17_11_01_59_train/
│   │   ├── lerobot_v3.0/
│   │   │   ├── meta/
│   │   │   │   ├── info.json
│   │   │   │   ├── modality.json
│   │   │   │   ├── episodes/
│   │   │   │   ├── tasks.parquet
│   │   │   │   └── subtasks.parquet
│   │   │   ├── data/
│   │   │   │   └── chunk-000/
│   │   │   └── videos/
│   │   │       ├── observation.images.head_left/
│   │   │       ├── observation.images.head_right/
│   │   │       ├── observation.images.wrist_left/
│   │   │       ├── observation.images.wrist_right/
│   │   │       ├── observation.images.tactile_deform/
│   │   │       └── observation.images.tactile_raw/
│   │   └── lerobot_v2.1/
│   │       ├── meta/
│   │       ├── data/
│   │       └── videos/
│   └── season_.../
├── cover_ball_with_cup/
├── reposition_and_stack_blocks/
├── play_tic_tac_toe/
└── .../

</details>

Storage Layout

PartDescription
meta/Dataset schema, episode metadata, statistics, tasks, subtasks, and annotations
data/Frame-aligned robot data stored as Apache Parquet files
videos/Per-camera MP4 video streams

<details> <summary>LeRobot path templates</summary>

LeRobot v3.0 uses paths similar to:

text
data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4

LeRobot v2.1 uses paths similar to:

text
data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet
videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4

</details>

Feature Schema

All inspected seasons expose the same main feature set.

FeatureTypeShapeDescription
observation.statefloat3265Joint-space robot state
actionfloat3265Joint-space action target
observation.state.joint_torquefloat3265Joint-torque signal
observation.state.tcpfloat3224Left/right tool-center-point pose and force state
observation.state.tcp_posefloat3212Left/right observed TCP poses, 6 components per arm
action.tcp_posefloat3212Left/right commanded TCP poses, 6 components per arm
observation.tactilefloat326010 fingertips x 6-axis force/torque signal
observation.images.*videovariesSix synchronized visual and tactile streams
timestampfloat321Frame timestamp
frame_indexint641Frame index within an episode
episode_indexint641Episode index
task_indexint641Task-description index
subtask_indexint641Subtask-annotation index

Joint-Space State and Action

The 65D state, joint-torque, and action vectors use the following order:

RangeNamesMeaning
0-6left_arm_j0 to left_arm_j6Left arm joints
7-28left_hand_j0 to left_hand_j21Left dexterous-hand joints
29-35right_arm_j0 to right_arm_j6Right arm joints
36-57right_hand_j0 to right_hand_j21Right dexterous-hand joints
58-64motor_j0 to motor_j6Torso / motor-related joints

TCP Pose State and Action

observation.state.tcp_pose and action.tcp_pose are separate 12D float32 vectors. Components 0:6 contain the left-arm TCP pose and 6:12 contain the right-arm TCP pose.

The observation is read from the source state/left_arm/tcp_pose and state/right_arm/tcp_pose; the action is read from the corresponding action/.../tcp_pose streams. Each stream uses its own aligned_index. Source values and pose representation are preserved; commanded poses are not copied from observed poses or reconstructed using forward kinematics.

The existing 24D observation.state.tcp remains unchanged: left pose 0:6, left force/torque 6:12, right pose 12:18, right force/torque 18:24. Its two pose slices correspond to the new observed TCP-pose vector.

The export does not fully declare TCP units, reference frames or rotation convention (rotation_type is null). Use the acquisition/controller convention before geometric transformations; do not infer Euler angles or quaternions from the vector width alone.

Video Streams

Feature keyDescriptionShape
observation.images.head_leftLeft head camera480 x 480 x 3
observation.images.head_rightRight head camera480 x 480 x 3
observation.images.wrist_leftLeft wrist camera480 x 480 x 3
observation.images.wrist_rightRight wrist camera480 x 480 x 3
observation.images.tactile_deformTactile deformation video480 x 1200 x 3
observation.images.tactile_rawRaw tactile video480 x 1600 x 3

Video files are MP4 without audio. Codec may differ between exports; inspect the local files if your training stack has codec restrictions.

Language Annotations

Every episode includes language annotations at two complementary levels: a global description of the complete task and temporally grounded descriptions of the subtasks performed during the trajectory. These annotations support language-conditioned policy learning, vision-language-action training, temporal grounding, skill discovery, hierarchical policy learning, and subtask-aware evaluation.

Global Task Description

The episode-level global_task stores the available structured natural-language fields below:

ComponentMeaningExample
TaskShort task nameClearPlate
InstructionOrdered description of the intended behaviorPick up the plate and eraser, wipe the plate, then place both objects down
SceneObjects, layout, and manipulation contextA marked plate and a blackboard eraser are on the table
SuccessCompletion criteriaThe marks are removed and both objects are placed stably

The same information is also exposed as separate fields under tags. All episodes have scene descriptions; 1,061 scene descriptions were supplemented from existing episode text and visual spot checks, with provenance recorded in the annotation metadata. Some of these episodes do not have separate instruction, success-criteria or SOP fields:

text
task_name
task_instruction
scene_description
success_criteria
score
SOP

Temporally Grounded Subtasks

Each episode is decomposed into language segments. A segment contains:

FieldDescription
start_step, end_stepFrame-index interval for the subtask
textNatural-language description of the action phase
task_skillCompact semantic skill label such as pick, place, move, insert, fold, or wipe
labelAnnotation type, including subtasks and recovered action descriptions
idEpisode-local subtask identifier

Across the release, all 28,993 episodes contain language annotations. The dataset provides 208,264 temporal subtask segments spanning 28,897,608 frames, or approximately 89.2% of all frames. The segments use 45 nonempty skill labels; recovered descriptions may not have a separate skill label. Frames outside a fine-grained segment remain associated with the episode-level task description.

For example, a plate-cleaning trajectory is decomposed into phases such as:

text
Pick up the plate from the table with the left hand. | Skill: pick
Pick up the blackboard eraser with the right hand. | Skill: pick
Wipe away the dirty marks while the left hand holds the plate. | Skill: wipe
Place the plate and the eraser back on the table. | Skill: place

Annotation Files and Frame Alignment

MetadataPurpose
meta/annotations.jsonlPer-episode global task, structured tags, temporal segments, frame ranges, and source identifiers
meta/tasks.jsonl / tasks.parquetMapping from task_index to the global task text
meta/subtasks.jsonl / subtasks.parquetMapping from subtask_index to subtask text and skill
Frame feature task_indexConnects each frame to its global task description
Frame feature subtask_indexConnects each frame to its subtask annotation

The v2.1 export stores task and subtask lookup tables as JSONL; v3.0 stores the corresponding tables as Parquet. annotations.jsonl is included with both exports for direct episode-level inspection.

The following example reads the global description and temporally grounded language from the first annotated episode:

python
import json
from pathlib import Path

meta_root = Path(
    "SharpaDex-v1.0/clean_plate_with_eraser/"
    "season_POC22027_2026_04_17_11_01_59_train/"
    "lerobot_v3.0/meta"
)

with open(meta_root / "annotations.jsonl", "r") as f:
    annotation = json.loads(next(f))

print(annotation["global_task"])

for segment in annotation["language_segments"]:
    print(
        segment["start_step"],
        segment["end_step"],
        segment["task_skill"],
        segment["text"],
    )

When constructing training samples, use task_index for episode-level language conditioning and subtask_index or the explicit start_step / end_step intervals for phase-level conditioning.

Tactile Modality

The dataset provides tactile information in structured numeric and image-based forms:

  • —observation.tactile contains 60 values: left/right thumb, index, middle, ring, and little fingertips, each with fx, fy, fz, tx, ty, and tz.
  • —observation.images.tactile_deform visualizes contact-induced deformation patterns.
  • —observation.images.tactile_raw preserves the raw tactile camera layout for custom preprocessing and representation learning.

All tactile modalities are synchronized with visual observations, proprioception, and actions at 30 FPS. The 60D signal provides a compact tactile representation; the tactile video streams provide spatially resolved observations for models designed to process the additional input resolution.

Usage Recommendations

For policy learning, a typical configuration is:

  • —Visual observations: one or more observation.images.* streams
  • —Proprioception: observation.state
  • —Optional force and contact inputs: joint torque, TCP state, and observation.tactile
  • —Optional tactile vision: observation.images.tactile_deform and/or observation.images.tactile_raw
  • —Supervision target: action, or action.tcp_pose for a TCP-based action representation
  • —Language conditioning: structured task instruction, scene description, and success criteria
  • —Phase conditioning: frame-aligned subtask text and skill label

For evaluation, split by collection season rather than randomly splitting frames. This reduces temporal leakage and avoids placing closely related demonstrations from the same collection session in both training and evaluation sets. For multi-task experiments, additionally report held-out tasks or task families when measuring cross-task generalization.

Dataset Notes

  • —The release is organized by task and collection season, not as one flattened LeRobot root.
  • —Every released season includes both lerobot_v3.0 and lerobot_v2.1.
  • —Task descriptions, scene details, success criteria, and subtask decomposition may vary between seasons; use the metadata shipped with the selected season as the source of truth.
  • —Some long-horizon tasks contain multiple interaction phases and substantial hand-object occlusion.
  • —The two tactile video streams are high resolution and may dominate input bandwidth.
  • —The dataset contains demonstrations rather than a fixed benchmark split; users should document their task and season splits for reproducibility.

License and Terms

This dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You may share and adapt the dataset, including for commercial purposes, provided that you give appropriate attribution and indicate whether changes were made.

Citation

If this dataset contributes to your research, please cite or acknowledge the dataset:

bibtex
@misc{sharpadex_v1_2026,
  title        = {SharpaDex v1.0: Large-Scale Bimanual Dexterous Manipulation Demonstrations},
  author       = {{Sharpa}},
  howpublished = {\url{https://huggingface.co/datasets/Sharpa-Robotics/SharpaDex-v1.0}},
  year         = {2026}
}
Sharpa-Robotics/SharpaDex-v1.0 · Team Ai