datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reachy-mini-emotions-library
Reachy Mini Emotions Library
Curated emotion recordings for the Reachy Mini robot, maintained by
Pollen Robotics. Each move is a JSON trajectory (head pose, antennas,
body yaw, sampled over time) paired with an Opus audio track.
Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by
the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves
non-.wav audio sidecars).
File layout
Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.reachy-mini-dances-library
Reachy Mini Dances Library
Curated dance moves for the Reachy Mini robot, maintained by Pollen
Robotics. Each move is a JSON trajectory (head pose, antennas, body
yaw, sampled over time). Motion-only — no audio tracks in this set.
File layout
Files live at the root of the dataset, named <dance>.json.
How to use
Python — via the reachy_mini package:
from reachy_mini import ReachyMini
from reachy_mini.motion.recorded_move import RecordedMoves
library… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-dances-library.DatasetDemo
Motus Training Dataset Demo
Introduction
This repository serves as a demonstration dataset illustrating the required data format for training the Motus model. It provides a reference for structuring your data to ensure compatibility with the training pipeline. The demo data come from Robotwin-clean benchmark.
Directory Structure
Data is generally organized following a hierarchy of Dataset Name, Task Name (optional), and Data Type.
Standard Format:… See the full description on the dataset page: https://huggingface.co/datasets/motus-robotics/DatasetDemo.droid_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DAVIAN-Robotics/droid_v3.Liberobioburden_labelled_dataPhysicalAI-Robotics-GraspGen
GraspGen: Scaling Sim2Real Grasping
GraspGen is a large-scale simulated grasp dataset for multiple robot embodiments and grippers.
We release over 57 million grasps, computed for a subset of 8515 objects from the Objaverse XL (LVIS) dataset. These grasps are specific to three grippers: Franka Panda, the Robotiq-2f-140 industrial gripper, and a single-contact suction gripper (30mm radius).
Dataset Format
The dataset is released in the WebDataset format. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GraspGen.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.speech-commands-v0.02
Speech Commands Dataset v0.02
This is a re-hosted copy of the Google Speech Commands v0.02 dataset in Parquet format for compatibility with the Hugging Face Dataset Viewer.
⚠️ Credits
This dataset was created by Pete Warden / Google. All credit goes to the original authors and the crowdsourcing contributors.
Original source: http://download.tensorflow.org/data/speech_commands_v0.02.tar.gz
Paper: Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/speech-commands-v0.02.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.PhysicalAI-Robotics-GR00T-Eval
EVAL-175
Dataset Description:
123 initial frame pictures from the robot's perspective before performing various tasks in the lab.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA Corporation (GEAR Lab)
Dataset Creation Date:
May 1, 2025
License/Terms of Use:
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
This dataset was created using a… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Eval.PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval
VoMP: Predicting Volumetric Mechanical Properties
Dataset Description:
The Pre-Processed 3D Dataset is a dataset that is composed of 4 individual 3D asset datasets which are processed to render them from multiple views, voxelize the assets, and propagate VLM annotations for material properties.
We release pre-processed data derived from the 3D assets, specifically: voxels, rendered images, and LLM-annotated material descriptions.
This dataset is for research and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval.PhysicalAI-Robotics-GR00T-GR1
GR1-100
92 Videos of a Fourier GR1-T2 robot performing various tasks in the lab from the third person perspective.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA Corporation (GEAR Lab)
Dataset Creation Date:
May 1, 2025
License/Terms of Use:
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
This dataset was created using a Fourier GR1-T2 robot.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-GR1.factory-manipulation-videos
Factory manipulation videos
Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks.
Contents
Seven continuous takes, 109 minutes in total.
Task
Station
Worker
Duration
File
cardboard manipulation
01
041
23.6 min
cardboard_manipulation_station01_worker041.mp4
cardboard manipulation
04
026
16.5 min
cardboard_manipulation_station04_worker026.mp4
defect… See the full description on the dataset page: https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos.qwen_robotics_open_dataset_robosense
Robot Navigation Open Scenarios: RoboSense
Image-in, trajectory-out navigation scenarios with 3D ground truth, for evaluating a navigating
robot (the intended use is humanoid navigation). Converted from
RoboSense (Su et al., CVPR 2025).
Each scenario has two halves:
Input, the past. Front-camera images of the 10 past frames and the current frame, plus the
robot's own past positions. No obstacle information.
Ground truth, the future. The robot's recorded trajectory over the 10… See the full description on the dataset page: https://huggingface.co/datasets/Jinyan0924/qwen_robotics_open_dataset_robosense.pi-embodied-rd-interactions
pi-embodied Robotics R&D Multimodal Interactions
This public dataset contains 2,396 project-scoped agent turns, including 184 turns with recoverable, published images. It preserves successful, unsuccessful, interrupted and unverified interactions without inventing outcome labels.
Storage
data/train-*.parquet: normalized interactions with an explicit Arrow schema and Zstandard compression.
media/<hash-prefix>/<sha256>.<extension>: deduplicated original image bytes… See the full description on the dataset page: https://huggingface.co/datasets/LV-Robotics-Lab/pi-embodied-rd-interactions.TaskGrasp-Pro
TaskGrasp-Pro dataset
The TaskGrasp-Pro dataset extends the original TaskGrasp dataset by providing fine-grained part decompositions and part-level physical property annotations for 190 household objects. For each object instance, we design three types of tasks: category-related tasks, part-related tasks, and part-irrelevant tasks, resulting in a total of 2,850 tasks.
Files
scans/ contains the point clouds of objects, multi-view RGB-D images, different types of task… See the full description on the dataset page: https://huggingface.co/datasets/WCL-Robotics/TaskGrasp-Pro.qwen_robotics_nav_pretrain_dedup
Navigation pretraining frames, deduplicated
A 49,391-sample training set for the final-frame pretraining task of
qwen_robotics_open_dataset: past camera views + the view
5 s ahead → the recorded motion in between. It is cut from the same four open sources as the full per-frame sets
(1.11 M frames), keeping motion diversity first and scene diversity second, so that a pass over it costs a
fraction of the time and the rare behaviours (turns, stops, starts) are not drowned by… See the full description on the dataset page: https://huggingface.co/datasets/Jinyan0924/qwen_robotics_nav_pretrain_dedup.qwen_robotics_open_dataset_egowalk
EgoWalk navigation frames
EgoWalk (MIT) converted into a per-frame training
format for image-in, trajectory-out navigation policies: 57 hours and 232 km of people walking
in Moscow, indoors (malls, stations, offices) and outdoors, filmed by a chest-mounted ZED camera, with
visual odometry and about 79,000 language goals ("Walk to the glass door on the left").
It is one of the training sources of qwen_robotics_open_dataset,
which builds training data usable by both… See the full description on the dataset page: https://huggingface.co/datasets/Jinyan0924/qwen_robotics_open_dataset_egowalk.CausalVerse_Video_Robotics_Kitchen
CausalVerse Video Dataset
Available splits: robotics_kitchen
Each record contains the following columns:
videos, metavalue, npz_data
robotics_kitchen
Examples: 7500
Columns: videos, metavalue, npz_data
PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Cube Stacking task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking.qwen_robotics_open_dataset
Robot Navigation Open Scenarios: CODa
Image-in, trajectory-out navigation scenarios with 3D ground truth, for evaluating a
navigating robot (the intended use is humanoid navigation). Converted from the
UT Campus Object Dataset (CODa).
Each scenario has two halves:
Input, the past. Front-camera images of the 10 past frames and the current frame, plus the
robot's own past positions. No obstacle information.
Ground truth, the future. The robot's recorded trajectory over the 10… See the full description on the dataset page: https://huggingface.co/datasets/Jinyan0924/qwen_robotics_open_dataset.CausalVerse_Video_Robotics_Mobile
CausalVerse Video Dataset
Available splits: robotics_mobile
Each record contains the following columns:
videos, metavalue, npz_data
robotics_mobile
Examples: 2001
Columns: videos, metavalue, npz_data
SynthRender_Robotics
SynthRender_Robotics: Synthetic Training Sets for the Robotics Sim-to-Real Benchmark
This repository hosts the synthetic training sets generated with SynthRender and used to benchmark sim-to-real transfer on the public Robotics dataset (Horváth et al., 2022), as reported in "SynthRender and IRIS: Open-Source Framework and Dataset for Bidirectional Sim-Real Transfer in Industrial Object Perception" (arXiv:2602.21141).
Sample Images (720x720)… See the full description on the dataset page: https://huggingface.co/datasets/moiaraya/SynthRender_Robotics.breakfast_robocup_homeCausalVerse_Video_Robotics_Living
CausalVerse Video Dataset
Available splits: robotics_living
Each record contains the following columns:
videos, metavalue, npz_data
robotics_living
Examples: 4800
Columns: videos, metavalue, npz_data
robotic-seminars-cacherobotics-quality-leaderboard
Robotics Dataset Quality Leaderboard
What This Is
This repository hosts an automatically-updated quality leaderboard for robotics
imitation-learning datasets on HuggingFace. Each dataset is scored by the
HaptalAI quality scorer, an
open-source tool that streams a sample of episodes from a dataset, detects the
available sensor schema, runs a suite of failure-detection checks, and computes
a single 0–100 quality score. The leaderboard is intended as a first-pass… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/robotics-quality-leaderboard.Veridis
VERIDIS Dataset
Overview
This repository contains the VERIDIS dataset, a collection of annotated agricultural images for crop detection and identification. The dataset comprises field images of beet and corn crops captured by a ground-level robotic platform, organized in YOLO format for object detection tasks.
The dataset primarily captures crops at early growth stages, which is particularly relevant for applications such as plant detection, early monitoring, and… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/Veridis.qwen_robotics_nav_eval
Robot navigation evaluation suite (v2)
150 hand-audited scenarios for image-in, path-out navigation policies of an indoor robot (a walking
humanoid, about 0.5 m/s): 100 indoor and 50 outdoor, drawn from 4 open datasets, each with a text prompt
that contains the goal and 3D ground truth (lidar, tracked people, obstacle maps) for scoring. Built by
qwen_robotics_open_dataset
(scripts/build_eval_suite.py; design and decisions in docs/eval_design.md, every audit decision in… See the full description on the dataset page: https://huggingface.co/datasets/Jinyan0924/qwen_robotics_nav_eval.
