datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aloha_sim_insertion_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_scripted.aloha_sim_transfer_cube_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_scripted.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.scriptsancient-scripts-datasets
Ancient Scripts Decipherment Datasets
Collated datasets for the paper:
Deciphering Undersegmented Ancient Scripts Using Phonetic Prior
Jiaming Luo, Frederik Hartmann, Enrico Santus, Regina Barzilay, Yuan Cao
Transactions of the Association for Computational Linguistics, 2021
arXiv:2010.11054
This repository gathers the training datasets used in the paper — both those hosted in the authors' GitHub repos and the external cited sources.
Repository Structure
data/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Nacryos/ancient-scripts-datasets.pali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.scripted_atomic_step_pose_0.6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 955,
"total_frames": 159935,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:955"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_pose_0.6.so101_ball_cup_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 12,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/alexis779/so101_ball_cup_scripted.ur5_real_setup_scripted_demos_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "isaaclab_ur5",
"total_episodes": 200,
"total_frames": 86463,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DORLR/ur5_real_setup_scripted_demos_v1.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.kaggle_scripts_new_format_subset
Dataset Card for "kaggle_scripts_new_format_subset"
More Information needed
scripted_atomic_step_train_frac0.2_largeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 512,
"total_frames": 107566,
"total_tasks": 1,
"total_videos": 1024,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:512"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.2_large.scripted_atomic_step_train_frac0.3_large_blindThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.3_large_blind.square_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.agentview": {
"dtype": "video",
"shape": [
84,
84,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/Leejungwook/square_scripted.dynamic_robot_bench_dr_scripted_10k
dynamic_robot_bench_dr_scripted_10k
10,000 scripted-expert demonstrations across all 100 dynamic task families of
dynamic-robot-bench — a conveyor-belt dynamic-manipulation benchmark (Franka
Panda + wrist camera, ManiSkill 3 / SAPIEN GPU sim). One LeRobot v2.1 dataset:
100 episodes per family, success-filtered, language-prompted per episode.
Collection configuration (identical for every family)
Scripted expert with per-step auto-derived speed caps, recorded as… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_10k.script-fidelity-benchmark
Script fidelity benchmark
Anonymous supplement for the paper "Script collapse in multilingual ASR:
A reference-free metric and 100-pair benchmark."
Script Fidelity Rate (SFR) measures the fraction of ASR hypothesis characters
that belong to the expected target script. WER measures word edits, while SFR
checks whether the output is written in the target orthography.
Related resources:
PyPI package: https://pypi.org/project/script-fidelity/
Hugging Face Evaluate metric:… See the full description on the dataset page: https://huggingface.co/datasets/themechanism/script-fidelity-benchmark.square_scripted_20_demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.agentview": {
"dtype": "video",
"shape": [
84,
84,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/Leejungwook/square_scripted_20_demo.scripted_atomic_train_frac_0.3_large_goal_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_train_frac_0.3_large_goal_annotation.pick_place_franka_franka_umi_tacticle_scriptedThis dataset was created using LeRobot.
Dataset Description
Scripted-expert DexSuite data for pick_place with the Franka arm and the UMI tactile parallel gripper: 200 successful episodes from the waypoint planner with randomized object initialization, including dense fingertip tactile readings (Flexitac simulation). The wrist camera looks at the flat face of the fingers (2026-09-16 mount). All episodes are successful, recorded at 20 Hz with a static front camera and a… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/pick_place_franka_franka_umi_tacticle_scripted.so101_cube_disk_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 12,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/alexis779/so101_cube_disk_scripted.lift_franka_franka_umi_tacticle_scriptedThis dataset was created using LeRobot.
Dataset Description
Scripted-expert DexSuite data for lift with the Franka arm and the UMI tactile parallel gripper: 200 successful episodes from the waypoint planner with randomized object initialization, including dense fingertip tactile readings (Flexitac simulation). The wrist camera looks at the flat face of the fingers (2026-09-16 mount). All episodes are successful, recorded at 20 Hz with a static front camera and a wrist… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/lift_franka_franka_umi_tacticle_scripted.stack_franka_franka_umi_tacticle_scriptedThis dataset was created using LeRobot.
Dataset Description
Scripted-expert DexSuite data for stack with the Franka arm and the UMI tactile parallel gripper: 200 successful episodes from the waypoint planner with randomized object initialization, including dense fingertip tactile readings (Flexitac simulation). The wrist camera looks at the flat face of the fingers (2026-09-16 mount). All episodes are successful, recorded at 20 Hz with a static front camera and a wrist… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/stack_franka_franka_umi_tacticle_scripted.push_franka_franka_umi_tacticle_scriptedThis dataset was created using LeRobot.
Dataset Description
Scripted-expert DexSuite data for push with the Franka arm and the UMI tactile parallel gripper: 200 successful episodes from the waypoint planner with randomized object initialization, including dense fingertip tactile readings (Flexitac simulation). The wrist camera looks at the flat face of the fingers (2026-09-16 mount). All episodes are successful, recorded at 20 Hz with a static front camera and a wrist… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/push_franka_franka_umi_tacticle_scripted.reach_franka_franka_umi_tacticle_scriptedThis dataset was created using LeRobot.
Dataset Description
Scripted-expert DexSuite data for reach with the Franka arm and the UMI tactile parallel gripper: 200 successful episodes from the waypoint planner with randomized object initialization, including dense fingertip tactile readings (Flexitac simulation). The wrist camera looks at the flat face of the fingers (2026-09-16 mount). All episodes are successful, recorded at 20 Hz with a static front camera and a wrist… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/reach_franka_franka_umi_tacticle_scripted.scripted_atomic_step_train_frac0.3_large_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 1328,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.3_large_image.isaaclab-ur7e-ball_topside-scripted_shoulderR_100This dataset was created using LeRobot.
Dataset Description
Merged from:
ICRA2027-CSI/isaaclab-ur7e-ball_top-scripted_shoulderR episodes 0-49 -> merged episodes 0-49 (50 episodes)
ICRA2027-CSI/isaaclab-ur7e-ball_side-scripted_shoulderR episodes 0-49 -> merged episodes 50-99 (50 episodes)
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0"… See the full description on the dataset page: https://huggingface.co/datasets/DexSteer/isaaclab-ur7e-ball_topside-scripted_shoulderR_100.assembly_peg_insert_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 128,
"total_frames": 4055,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 7,
"splits": {
"train": "0:128"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lukasskellijs/assembly_peg_insert_scripted.
