datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hy-Embodied-0.5-VLA-Data
Hy-Embodied-0.5-VLA
From Vision-Language-Action Models to a Real-World Robot Learning Stack
Tencent Robotics X × Tencent Hy Team
📖 Abstract
We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.VLADBenchLavalObjaverseDataset
Laval Objaverse Dataset
vLAR Group | SIGGRAPH Asia 2026
A large-scale, high-quality dataset for multi-view relighting.
📖 Dataset Summary
The Laval Objaverse Dataset is a comprehensive dataset designed for multi-view relighting and novel view synthesis tasks. It combines high-quality 3D assets from Objaverse with realistic, diverse illumination conditions… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/LavalObjaverseDataset.PhysInOnePhysInOne: Visual Physics Learning and Reasoning in One Suite
vLAR Group |
The Hong Kong Polytechnic University |
Syai Singapore |
Meta
CVPR 2026
🧭 Navigation
📌 Summary
🚀 Release Timetable
📦 Repositories & Downloads
📊 Data Splits
🧱 3D Assets
🛠️ Data Processing
🏆 Leaderboard Evaluation Data
🎞️ Rendered Data
1. Download Scripts
2. Install Dependencies… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/PhysInOne.eval-resultsVLA4CoDrive
Vision–Language–Action Dataset for Cooperative Autonomous Driving
VLA4CoDrive is a large-scale cooperative Vision–Language–Action (VLA) dataset designed to support autonomous driving under multi-vehicle cooperation. This work has been accepted to the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. This dataset was developed at
AI-SENDs Lab,
Clemson University, USA.
🔍 Overview
We introduce VLA4CoDrive, a cooperative… See the full description on the dataset page: https://huggingface.co/datasets/sayedpedramhaeri/VLA4CoDrive.vlabench_primitive_pretrain_lerobot
Datacard
This is the official VLABench primitive pretraining dataset converted to the
LeRobot format. The dataset contains language-conditioned manipulation
trajectories collected with a Franka Panda robot in VLABench simulation.
This LeRobot version is hosted at:
https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code:… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot.gpt-edit-simplermikasa-robo-vla-rldsVLAC-Cut-FullData
VLAC-Cut-FullData
VLAC-Cut-FullData is the full-data release for VLAC-Cut. It provides the complete raw-data archive set, benchmark-style JSON files, and a lightweight frame-extraction workflow for reproducing evaluation on the released benchmark protocol.
Contents
benchmark_style_all/
train/video_progress_benchmark_file.json
test_expert_seen/video_progress_benchmark_file.json
test_expert_unseen/video_progress_benchmark_file.json… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/VLAC-Cut-FullData.Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B
Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions.
Dataset Details
Dataset Description
Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM.
Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.HQ-Edit
Dataset Card for HQ-EDIT
HQ-Edit, a high-quality instruction-based image editing dataset with total 197,350 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3.
HQ-Edit’s high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/HQ-Edit.vlabench_composite_ft_lerobot_videosurg-vla-datasetvlabench-assetsvlaGPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.VLA_Instruction_TuningThis repository contains the VLA-IT dataset, a curated 650K-sample Vision-Language-Action Instruction Tuning dataset, and the SimplerEnv-Instruct benchmark. These are presented in the paper InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation. The dataset is designed to enable robots to integrate multimodal reasoning with precise action generation, preserving the flexible reasoning of large vision-language models while delivering leading manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/VLA_Instruction_Tuning.VLAC-Cut-Benchmark
Video Progress Benchmark
Paper ·
Code ·
Model ·
Benchmark
Overview
The Video Progress Benchmark (VPB) evaluates process-level task progress estimation for robot manipulation. It measures whether a model can capture advancement, stagnation, regression, and recovery throughout an execution video.
VPB is built from the held-out portion of the Progress Annotation Dataset. Unlike endpoint-only evaluations, VPB focuses on temporal task progress and supports analysis… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/VLAC-Cut-Benchmark.VLA-OS-Dataset
Dataset Card
This is the training dataset used in the paper VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models.
Source
Project Page: https://nus-lins-lab.github.io/vlaos/
Paper: https://arxiv.org/abs/2506.17561
Code: https://github.com/HeegerGao/VLA-OS
Model: https://huggingface.co/Linslab/VLA-OS
Usage
Ensure you have installed git lfs:
curl -s… See the full description on the dataset page: https://huggingface.co/datasets/Linslab/VLA-OS-Dataset.minecraft-vla-sftThis repository contains the dataset used for JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse.
Project Website
VLABench
VLABench VLM Evaluation Dataset
This dataset is the VLM evaluation split of VLABench, prepared for reproducible VLABench evaluation with PhysBrainEvalKit.
Source
Project page: https://vlabench.github.io/
Paper: https://arxiv.org/abs/2412.18194
Official code: https://github.com/OpenMOSS/VLABench
Directory layout
The dataset is organized by evaluation dimension and subtask:
vlm_evaluation_v1.0/
├── CommenSence/
├── Complex/
├── M&T/
├── PhysicsLaw/… See the full description on the dataset page: https://huggingface.co/datasets/VLyb/VLABench.LET-KUAVO-VLA-1.0-Dataset
LET-KUAVO-VLA-1.0-Dataset
vlabench_primitive_ft_lerobot_video
VLABench Primitive Tasks Dataset - LeRobot v3.0
Dataset Description
This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially.
Compared with the v2.0 version and the RLDS version of the dataset, this release stores visual observations in a video-compressed format rather than as individual image files. This design provides significant advantages in both storage efficiency and data loading performance.… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_ft_lerobot_video.collected_demos_trainingAIR-VLA_hdf5_datasetsdrifting-vla-v2-droidVL-Adapter-datasets
VL-Adapter Datasets
Processed CLIP-ResNet101 grid features and annotations for
VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks
(Sung, Cho, Bansal — CVPR 2022), code at ylsung/VL_adapter.
This repository replaces the original Google Drive download, which is no longer
available. It holds the same data in a Hub-native layout, plus a script that
rebuilds the exact datasets/ directory tree the training code expects.
Quick start — rebuild… See the full description on the dataset page: https://huggingface.co/datasets/ylsung/VL-Adapter-datasets.raw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.VLA_Arena_L0_L_lerobot_openpi
VLA-Arena Dataset (L0 - Large Variant)
About VLA-Arena
VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L0_L_lerobot_openpi.
