datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-Robotics-GR00T-X-Embodiment-Sim
PhysicalAI-Robotics-GR00T-X-Embodiment-Sim
Github Repo: Isaac GR00T N1
We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks.
Cross-embodied bimanual manipulation: 9k trajectories
Dataset Name
#trajectories
bimanual_panda_gripper.Threading
1000
bimanual_panda_hand.LiftTray
1000
bimanual_panda_gripper.ThreePieceAssembly
1000… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim.SAGE-10k
SAGE-10k
SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects.
🔑 Key Features
SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.PhysicalAI-Autonomous-Vehicles
PHYSICAL AI AUTONOMOUS VEHICLES
The PhysicalAI-Autonomous-Vehicles dataset provides one of the largest, geographically diverse collections of multi-sensor data empowering AV researchers to build the next generation of Physical AI based end-to-end driving systems. This dataset is ready for commercial/non-commercial AV use per the license agreement.
Data Collection Method
Automatic/Sensor
Labeling Method
Automatic/Sensor
This dataset has a total of 1700 hours of driving… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles.PhysicalAI-Robotics-Open-H-Embodiment
Dataset Description:
Open-H-Embodiment is a community‑driven dataset initiative building the open, shared foundation needed to train and evaluate AI autonomy models for surgical robotics and ultrasound.
This dataset is a multi-embodiment collection of LeRobot datasets of paired kinematics and video, across tasks such as tabletop exercises, clinical procedures, as well as simulations of healthcare robotics applications.
Maintainer / Hosting Organization:
NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Open-H-Embodiment.PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
Dataset Description:
PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.OpenMathInstruct-2
OpenMathInstruct-2
OpenMathInstruct-2 is a math instruction tuning dataset with 14M problem-solution pairs
generated using the Llama3.1-405B-Instruct model.
The training set problems of GSM8K
and MATH are used for constructing the dataset in the following ways:
Solution augmentation: Generating chain-of-thought solutions for training set problems in GSM8K and MATH.
Problem-Solution augmentation: Generating new problems, followed by solutions for these new problems.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMathInstruct-2.PhysicalAI-SmartSpaces
Physical AI Smart Spaces Dataset
Overview
Comprehensive, annotated dataset for multi-camera tracking and 2D/3D object detection. This dataset is synthetically generated with Omniverse and Cosmos Transfer.
This dataset consists of over 280 hours of video from across nearly 1,800 cameras from indoor scenes in warehouses, hospitals, retail, and more. The dataset is time synchronized for tracking humans, forklifts, pallet trucks and Autonomous Mobile Robots (AMRs)… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.HelpSteer2
HelpSteer2: Open-source dataset for training top-performing reward models
HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
This dataset has been created in partnership with Scale AI.
When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.PhysicalAI-Robotics-Locomanipulation-GRAIL
📢 News
[2026-07-15] Released task-general tracking policy checkpoints trained on the released data. Follow the tracking doc to use them to track our released motion data.
[2026-07-14] Updated data/pickup_table and data/pickup_ground. If you downloaded them before this date, please re-download.
Dataset Overview
Tabletop Pickup
Ground Pickup
Tabletop Manipulation
Ground Manipulation
Sitting
Curb
Slope… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL.OpenMathReasoning
OpenMathReasoning
OpenMathReasoning is a large-scale math reasoning dataset for training large language models (LLMs).
This dataset contains
306K unique mathematical problems sourced from AoPS forums with:
3.2M long chain-of-thought (CoT) solutions
1.7M long tool-integrated reasoning (TIR) solutions
566K samples that select the most promising solution out of many candidates (GenSelect)
Additional 193K problems sourced from AoPS forums (problems only, no solutions)
We used… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMathReasoning.Nemotron-CC-v2
Nemotron-Pre-Training-Dataset-v1 Release
Data Overview
This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models.
This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
PhysicalAI SDG-Warehouse
PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.OpenCodeInstruct
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
Dataset Description
We introduce OpenCodeInstruct, the largest open-access instruction tuning dataset, comprising 5 million diverse samples. OpenCodeInstruct is designed for supervised fine-tuning (SFT).
Technical Report - Discover the methodology and technical details behind OpenCodeInstruct.
Github Repo - Access the complete pipeline used to perform SFT.
This dataset is ready for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeInstruct.PhysicalAI-Autonomous-Vehicles-NuRec
task_categories:
- robotics
tags:
- physicalAI
Find the 1500+ scenes in the sample_set/26.04_release folder.
Dataset Description:
Neural reconstructed dataset that carries 3D reconstructed driving scenes. The scenes are about 20 second long and stored in form of usdz files, along with respective xodr map files, surface mesh. The reconstructions were generated using 6 camera views (front-wide 120 deg, front-tele 30 deg, cross right/left 120 deg and rear right/left… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles-NuRec.PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
Dataset Description:
The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research.
Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.PhysicalAI-Robotics-Manipulation-Kitchen-Demos
PhysicalAI-Robotics-Manipulation-Kitchen-Demos
We provide a 600 hours of human-teleoperated demonstrations across 316 different tasks, totalling 55k trajectories.
The datasets are collected using Franka Panda robot with an Omron mobile base.
The datasets follow the LeRobot format. Here is an overview of important elements of each dataset:
Click to expand dataset structure
lerobot/
├── meta/ # Metadata files describing the dataset
│ ├── info.json… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen-Demos.Open-SWE-Traces
Open-SWE-Traces: Advancing Distillation for Software Engineering Agents
🚨 What's New
[09/26] Release v1.2: Added new agent trajectories generated by Qwen3.8-27B for
mini-swe-agent. Trajectories for OpenCode and Claude Code harnesses will be released soon.
[08/26] Release v1.1: Added new agent trajectories generated by DeepSeek-V4-Flash and
Qwen3.6-27B across OpenHands,
SWE-agent, and mini-swe-agent harnesses.
[06/21] Release v1.0: Released 207k agent… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Open-SWE-Traces.HiLiftAeroML
HiLiftAeroML: High-Fidelity Computational Fluid Dynamics Dataset for High-Lift Aircraft Aerodynamics
Contact:
Neil Ashton (contact@caemldatasets.org)
Summary
This dataset provides the first-ever open-source high-fidelity CFD dataset of a high-lift aircraft for the purpose of AI surrogate model development. The dataset is composed of 1,800 samples, arising from 180 geometry variants of the NASA Common Research Model (CRM-HL) across 10 angles of attack (from 4° to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HiLiftAeroML.Nemotron-CC-v2.1
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1.PhysicalAI-Robotics-GR00T-Teleop-GR1
Introduction
TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments.
Project page: https://dreamdojo-world.github.io/
Paper: https://arxiv.org/abs/2602.06949
Code: https://github.com/NVIDIA/DreamDojo
How to Use
Check out https://github.com/NVIDIA/DreamDojo
Citation
@article{gao2026dreamdojo,
title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.PhysicalAI-Robotics-Manipulation-SingleArm
Dataset Description:
PhysicalAI-Robotics-Manipulation-SingeArm is a collection of datasets of automatic generated motions of a Franka Panda robot performing operations such as block stacking, opening cabinets and drawers. The dataset was generated in IsaacSim leveraging task and motion planning algorithms to find solutions to the tasks automatically [1, 3]. The environments are table-top scenes where the object layouts and asset textures are procedurally generated [2].This dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-SingleArm.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.OpenScienceReasoning-2
Dataset Description
OpenScienceReasoning-2 is a multi-domain synthetic dataset designed to improve general-purpose reasoning in large language models (LLMs). The dataset contains multiple-choice and open-ended question-answer pairs with detailed reasoning traces and spans across diverse scientific domains, including STEM, law, economics, and humanities. OpenScience aims to boost accuracy on advanced benchmarks such as GPQA-Diamond, MMLU-Pro and HLE via supervised finetuning or… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2.video-to-data-robot-dexterity-task-library-and-dataset
Video to Data: Robot Dexterity Task Library and Dataset
Dataset Description
This dataset contains samples of human demonstrations on manipulation tasks retargeted to bimanual Sharpa robot hands and episodes of robot executions that mimic the original human demonstrations. The former allows a Video to Data user to easily experiment with the Video to Data grounding pipeline, and the latter is an example of the grounded robot data that can be generated with the Video… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-to-data-robot-dexterity-task-library-and-dataset.SPEED-Bench
📒 Blog |
📄 Paper |
🤗 Data |
⚙️ Measurement Framework
SPEED-Bench (SPEculative Evaluation Dataset) is a unified benchmark designed to evaluate speculative decoding (SD) across diverse semantic domains and realistic serving regimes, using production-grade inference engines.
It measures both acceptance-rate characteristics and end-to-end throughput, enabling fair, reproducible, and robust comparisons between SD strategies.
SPEED-Bench introduces a benchmarking ecosystem for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SPEED-Bench.PhysicalAI-Robotics-GR00T-Teleop-Sim
Simulation GR1 Tabletop Task 1K Dataset
Dataset Description:
The PhysicalAI-Robotics-GR00T-Teleop-GR1 dataset consists of 1000 teleoperation trajectories in simulation using the GR1 robot with upper body control. The simulation setup mimics tabletop manipulation tasks and uses RGB observations with a virtual camera. The robot is equipped with simulated Fourier hands.
This dataset is ready for non-commercial use.
Dataset Owner(s):
NVIDIA GEAR… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim.Nemotron-CC-Math-v1
Nemotron-Pre-Training-Dataset-v1 Release
👩💻 Authors: Rabeeh Karimi Mahabadi, Sanjeev Satheesh
📘 Paper: Nemotron-cc-math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
📝 Blog: Nemotron-cc-math blog
Data Overview
We’re excited to introduce Nemotron-CC-Math - a large-scale, high-quality math corpus extracted from Common Crawl which was used in nemotron pre-training.
This dataset is built to preserve and surface high-value mathematical and code content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-Math-v1.Nemotron-ClimbLab
ClimbLab Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.OCR-Synthetic-Multilingual-v1
OCR-Synthetic-Multilingual-v1
Dataset Description
Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al.
This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OCR-Synthetic-Multilingual-v1.
