datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hy-Embodied-0.5-VLA-Data
Hy-Embodied-0.5-VLA
From Vision-Language-Action Models to a Real-World Robot Learning Stack
Tencent Robotics X × Tencent Hy Team
📖 Abstract
We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B
Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions.
Dataset Details
Dataset Description
Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM.
Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.vlabench_composite_ft_lerobot_videovlabench_primitive_ft_lerobot_video
VLABench Primitive Tasks Dataset - LeRobot v3.0
Dataset Description
This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially.
Compared with the v2.0 version and the RLDS version of the dataset, this release stores visual observations in a video-compressed format rather than as individual image files. This design provides significant advantages in both storage efficiency and data loading performance.… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_ft_lerobot_video.VR-egodex-annotation-converted-v6.0
VR-egodex-annotation-converted-v6.0
EgoDex converted from LeRobot v2.1 into the Layer-1 v0.6.0 annotation schema, with
per-clip narration included as language sidecars.
314,839 clips · 78,282,306 frames · 724.8 hours @ 30 fps · 129 tasks
100% narration coverage (1 sidecar per clip)
71 GB annotations + 2.3 GB narratives
Videos are NOT included. This release contains annotations and narration only. Source
video lives in griffinlabs/EgoDex-LeRobot-v3.0;
orig_id in the manifest… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egodex-annotation-converted-v6.0.mikasa-robo-vla-lerobotverl_vla_libero_collectedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 32,
"total_frames": 2990,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1e-06,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miical/verl_vla_libero_collected.vlabench_unifiedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 10977,
"total_frames": 3114872,
"total_tasks": 295,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:10977"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/vlabench_unified.train_ctf_eefnyush-vlai-l1-block-placement
NYUSH VLAI L1 — Block Placement
Data collected by ColinSkywalker.
Formats and branches
Branch
Representation
main (default)
Joint · LeRobot v2.1
joint-v2.1
Joint · LeRobot v2.1
Preview
Each GIF contains both camera views on one shared playback clock. Left: Agent View; right: Wrist Camera. The timestamp shows elapsed dataset time; playback is 2× speed.
Put the blue block into the red plate
Episode 16 · complete… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-vlai-l1-block-placement.train_ctftest_ctflingbot-vla-transfer
Lingbot VLA transfer — Hugging Face only
已完成实验:E5 allclean2epoch global64
作业25708998已完成17180步。此前的 e5-allclean-2epoch-gbs64-hf-only-b-20261009 是fresh复现包,
现已补发实际训练好的最终模型(FP32存储,6分片,约23.77GiB)。
模型固定revision:b3306e5b517d33815d7cafe392d5cc326af97bc2。
完成态实验与使用说明|机器可读索引。
原source包和旧manifest保留;本页下面的八H200实验仍是另一项尚未训练的新配置。
迁入端不需要GitHub账号、git clone或GitHub网络。所有接收、配置和训练入口都在固定HF包内。
新实验:E5八H200,micro24,GA2,6epoch
Release:… See the full description on the dataset page: https://huggingface.co/datasets/ar-mine/lingbot-vla-transfer.csgo-vla-stage1-5hz
CS:GO VLA Stage 1 Dataset (5Hz Chunked)
Vision-Language-Action dataset for Counter-Strike: Global Offensive with action chunking, converted from the TeaPearce CS:GO dataset.
Overview
Frame rate: 5Hz (every 3rd frame)
Action chunking: 3 actions per sample (~200ms coverage)
Total samples: ~1.8M chunks
Split: train / test following Diamond split
Map: Dust2 deathmatch
Action Format
<|action_start|> m1_x m1_y [keys1] ; m2_x m2_y [keys2] ; m3_x m3_y [keys3]… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/csgo-vla-stage1-5hz.VLADBench-reeval
VLADBench-reeval
We re-evaluated the VLADBench benchmark (Li et al., 2025, arXiv:2503.21505) against current SOTA VLMs under the original scoring criteria and prompts. See Eventual-Inc/VLADBench for the code, the companion site for interactive results, and the article for a summary of the findings.
Cost vs Score
Up and to the left is better. The dashed line is the cost-performance frontier which is defined by measuring the best TOTAL score available against its… See the full description on the dataset page: https://huggingface.co/datasets/Eventual-Inc/VLADBench-reeval.2chSTRUCTURE.md — подробности датасета, разбор схемы и семантика форума
⚡ 2ch — Russian Anonymous Imageboard Archive · 2009–2026
Датасет тредов Двача
Total: 1 455 924 threads · 238 479 526 posts
BoardThreadsPostsShare
/b/872 569114 896 80948.18%
/po/225 71035 024 43914.69%
/fag/48 56324 337 45610.21%
/sex/35 5576 611 2842.77%
/news/69 5005 031 2052.11%
/dev/18 4844 310 2051.81%
/mov/8 4463 309 6041.39%
/ftb/5 3242 574 1131.08%
/wm/4 3482 494 5441.05%
/wrk/5 5752 181 8450.91%
/sp/3… See the full description on the dataset page: https://huggingface.co/datasets/vladlinv/2ch.atari-vla-stage1-15hz
TESS-Atari Stage 1 (15Hz)
Human gameplay demonstrations from Atari games with action chunking, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~1.3M
Observation Rate
5 Hz
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Why Action Chunking?
VLA models run at ~5 Hz inference speed, but Atari runs at… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-15hz.crane_vlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"action": {
"dtype": "float32",
"shape": [
4
],
"names": [
"slew",
"boom_angle",
"extension",
"hoist"
]
},
"observation.state": {
"dtype": "float32"… See the full description on the dataset page: https://huggingface.co/datasets/Grigorij/crane_vla.vla0-context-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/mattpidden/vla0-context-dataset.so100_dsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 60,
"total_frames": 71734,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/so100_ds.vlabench_sub_dir_detlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 5000,
"total_frames": 567036,
"total_tasks": 128,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:5000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/globcy/vlabench_sub_dir_detla.atari-vla-stage1-5hz
TESS-Atari Stage 1 (5Hz)
Human gameplay demonstrations from Atari games, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~4M
Action Rate
5 Hz (1 action per observation)
Format
Lumine-style action tokens
Games Included
Alien, Asterix, BankHeist, Breakout, DemonAttack, Freeway, Frostbite, Hero, MsPacman, RoadRunner, Seaquest… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-5hz.vla0-real-world-dataset-v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/mattpidden/vla0-real-world-dataset-v3.omy_f3m_motor_feedback_vla_probe_rock_real_left_probe_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 20,
"total_frames": 10481,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_probe_rock_real_left_probe_right.omy_f3m_motor_feedback_vla_probe_rock_real_right_probe_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 30,
"total_frames": 16552,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_probe_rock_real_right_probe_right.my-vla-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sentinel_v2",
"total_episodes": 1,
"total_frames": 484,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/grboguz/my-vla-1.VLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame
Spa-Bench fine-tuning demonstrations — motion-trimmed release
This is a non-destructive, motion-trimmed derivative of the
canonical 1,200-episode Spa-Bench dataset.
It removes initial idle prefixes while preserving episode identity, prompt,
action/state alignment, and all five source camera streams at the public head.
Explore episodes in the LeRobot visualizer
Transformation
The baseline is the component-wise median of the first five action frames.
Motion onset… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame.vla_total_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 222,
"total_frames": 108018,
"total_tasks": 7,
"total_videos": 444,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:222"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dongkkka/vla_total_dataset.spa-bench-eval-rollouts-vla-0-partial
Spa-Bench Evaluation Rollouts — VLA-0 (Partial)
Open this dataset in the LeRobot visualizer
This repository contains physical Spa-Bench rollout trajectories with synchronized middle and wrist RGB video, robot state, action, timestamps, episode indices, and task indices. The native Hugging Face Data Studio viewer is enabled through the Parquet files declared above.
Dataset summary
Coverage: 72 scored familiar-condition recordings; withheld-composition evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Spa-Bench/spa-bench-eval-rollouts-vla-0-partial.SONIC-VLA-BonesSeed-V2
SONIC-VLA-BonesSeed-V2 — prompt → motion-token dataset with physical onset augmentation (Unitree G1)
A LeRobot v2.1 dataset pairing a language prompt + ego-view + proprioception with the 64-dim
FSQ motion_token of the GEAR-SONIC whole-body
controller. It is the training input for
wsagi/GR00T-N1.7-G1-SONIC-BonesSeed-V2,
whose headline is fixing cold-start onset (one-shot motions self-starting from a standing pose,
no bootstrap).
What's new vs V1 (wsagi/SONIC-VLA-BonesSeed):
V1… See the full description on the dataset page: https://huggingface.co/datasets/wsagi/SONIC-VLA-BonesSeed-V2.
