datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aloha_sim_insertion_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_scripted.aloha_sim_transfer_cube_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_scripted.aloha_sim_insertion_scripted_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_scripted_image.aloha_sim_transfer_cube_scripted_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_scripted_image.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.aloha_sim_transfer_cube_scripted_rawdynamic_robot_bench_dr_scripted_14k
dynamic_robot_bench_dr_scripted_14k
14,400 scripted-expert demonstrations across the 72 evaluated task
families of dynamic-robot-bench — a
conveyor-belt dynamic-manipulation benchmark (Franka Panda + wrist camera, ManiSkill 3 / SAPIEN
GPU sim). One LeRobot v2.1 dataset: 200 episodes per family,
success-filtered, every domain-randomization knob on, and belt speed uniform over
0.10–0.40 m/s.
1,118,617 frames · 209 distinct language instructions · 20 fps
The belt speed… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_14k.BridgeData-V2-Scripted-Images
BridgeData V2 Image Triplets Dataset
This dataset contains image triplets from BridgeData V2 trajectories in ImageFolder format.
Derived From
This dataset is a derivative of the 30 GB scripted subset of BridgeData V2 from RAIL-Berkeley. All rights and original licensing apply.
Dataset Structure
initial_images/: Contains first frame images (initial state)
intermediate_images/: Contains intermediate frame images (frame 38)
final_images/: Contains final frame… See the full description on the dataset page: https://huggingface.co/datasets/VyoJ/BridgeData-V2-Scripted-Images.common-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.parc-track3-scripted-demos
PARC Track3 Scripted Demos
PARC2026 Track3(言語・構成的把持運搬)学習用の決定論的スクリプト方策デモ集。
Franka Emika Panda 7-DoF × LIBERO-plus キッチンシーン(MuJoCo、EGL、128px記録、
最大300ステップ/試行、制御20Hz)。
成功 = タスクdone かつ非注目物体の変位1mm以下。
生成コードはPARC2026プロジェクトの work/infra/remote/scripted_demo.py(+テスト)。
プランナーは決定論的:同一BDDL+同一.pruned_init+同一コード=同一デモ。
乱数源なし(random/np.random不使用)。物体配置はBDDL領域サンプリングによる。
バッチ一覧(timestamp dir)
dir
本数
方策
proxy smooth*
備考
20260828(3 dirs)
264
phases既定
k1 0.327
baseline… See the full description on the dataset page: https://huggingface.co/datasets/shimokawatoko/parc-track3-scripted-demos.anv_data_ke_kikuyu_scriptedaloha_sim_transfer_cube_scripted_image_rawaloha_sim_insertion_scripted_rawscripted_atomic_step_pose_0.6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 955,
"total_frames": 159935,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:955"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_pose_0.6.aloha_sim_insertion_scripted_image_rawso101_ball_cup_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 12,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/alexis779/so101_ball_cup_scripted.ur5_real_setup_scripted_demos_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "isaaclab_ur5",
"total_episodes": 200,
"total_frames": 86463,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DORLR/ur5_real_setup_scripted_demos_v1.scripted_atomic_step_train_frac0.2_largeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 512,
"total_frames": 107566,
"total_tasks": 1,
"total_videos": 1024,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:512"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.2_large.scripted_atomic_step_train_frac0.3_large_blindThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.3_large_blind.aloha_sim_transfer_cube_scripted_rawsquare_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.agentview": {
"dtype": "video",
"shape": [
84,
84,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/Leejungwook/square_scripted.dynamic_robot_bench_dr_scripted_10k
dynamic_robot_bench_dr_scripted_10k
10,000 scripted-expert demonstrations across all 100 dynamic task families of
dynamic-robot-bench — a conveyor-belt dynamic-manipulation benchmark (Franka
Panda + wrist camera, ManiSkill 3 / SAPIEN GPU sim). One LeRobot v2.1 dataset:
100 episodes per family, success-filtered, language-prompted per episode.
Collection configuration (identical for every family)
Scripted expert with per-step auto-derived speed caps, recorded as… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_10k.frankapickhard50_scripted_10demo_shoulder_100attempts_molmobottts-vc-mcv-scripted-v24.0-fy-nl-dii
tts-vc-mcv-scripted-v24.0-fy-nl-dii
This is a single-speaker speech dataset for West Frisian. It carries the Dii voice, a female voice. Source audio comes from Mozilla Common Voice scripted-speech prompts (release 24.0), read in Dutch/West Frisian. The audio was converted to the Dii voice identity using voice-conversion. The dataset contains 9676 recordings.
Related links
Models trained on this dataset:
OpenVoiceOS/phoonnx_fy-NL_dii_unicode
Dataset collection:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-vc-mcv-scripted-v24.0-fy-nl-dii.frankapickhard50_scripted_10demo_shoulder_12attempts_molmobotaloha_sim_insertion_scripted_rawslice_banana_franka_diffik.jointtarget_15hz_nominal_scripted_20260918square_scripted_20_demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.agentview": {
"dtype": "video",
"shape": [
84,
84,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/Leejungwook/square_scripted_20_demo.pick_place_franka_franka_umi_tacticle_scriptedThis dataset was created using LeRobot.
Dataset Description
Scripted-expert DexSuite data for pick_place with the Franka arm and the UMI tactile parallel gripper: 200 successful episodes from the waypoint planner with randomized object initialization, including dense fingertip tactile readings (Flexitac simulation). The wrist camera looks at the flat face of the fingers (2026-09-16 mount). All episodes are successful, recorded at 20 Hz with a static front camera and a… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/pick_place_franka_franka_umi_tacticle_scripted.scripted_atomic_train_frac_0.3_large_goal_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_train_frac_0.3_large_goal_annotation.
