Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /HelpSteer2 HelpSteer2: Open-source dataset for training top-performing reward models HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses. This dataset has been created in partnership with Scale AI. When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.tabular10K<n<100K456 likes50k downloads2y agoHugging Face02nvidia /PhysicalAI-Robotics-GR00T-Teleop-Sim Simulation GR1 Tabletop Task 1K Dataset Dataset Description: The PhysicalAI-Robotics-GR00T-Teleop-GR1 dataset consists of 1000 teleoperation trajectories in simulation using the GR1 robot with upper body control. The simulation setup mimics tabletop manipulation tasks and uses RGB observations with a virtual camera. The robot is equipped with simulated Fourier hands. This dataset is ready for non-commercial use. Dataset Owner(s): NVIDIA GEAR… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim.tabular1M<n<10M19 likes18k downloads10mo agoHugging Face03nvidia /PhysicalAI-Robotics-GR00T-Teleop-GR1 Introduction TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments. Project page: https://dreamdojo-world.github.io/ Paper: https://arxiv.org/abs/2602.06949 Code: https://github.com/NVIDIA/DreamDojo How to Use Check out https://github.com/NVIDIA/DreamDojo Citation @article{gao2026dreamdojo, title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.tabular1M<n<10M29 likes14k downloads8mo agoHugging Face04nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B131 likes6.8k downloads1y agoHugging Face05nvidia /Granary Granary: Speech Recognition and Translation Dataset in 25 European Languages Granary is a large-scale, open-source multilingual speech dataset covering 25 European languages for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks. Overview Granary addresses the scarcity of high-quality speech data for low-resource languages by consolidating multiple datasets under a unified framework: 🗣️ ~1M hours of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Granary.tabularautomatic-speech-recognition100M<n<1B225 likes5.7k downloads4mo agoHugging Face06nvidia /aisimulate-fpm-dataset AISimulate FPM Dataset Forward Pass Model (FPM) libraries and independent latency measurements. This private repository is the input-data source for FPM Gym. Layout data/ <org>--<model>/<system>/<framework>/<framework-version>/<parallelism>/ manifest.json # flat metadata for the promoted snapshot fpm/ # optional self-benchmark FPM libraries and sidecars provenance/ # FPM producer configurations… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/aisimulate-fpm-dataset.tabularn<1K6 likes2.9k downloads10d agoHugging Face07nvidia /HelpSteer HelpSteer: Helpfulness SteerLM Dataset HelpSteer is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses. Leveraging this dataset and SteerLM, we train a Llama 2 70B to reach 7.54 on MT Bench, the highest among models trained on open-source datasets based on MT Bench Leaderboard as of 15 Nov 2023. This model is available on… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer.tabular10K<n<100K253 likes2.3k downloads2y agoHugging Face08nvidia /Nemotron-RL-Ultra-Training-Blends Dataset Description: This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.tabulartext-generation10K<n<100K20 likes1.8k downloads11d agoHugging Face09nvidia /PhysicalAI-Robotics-GR00T-Teleop-G1 Unitree G1 Fruits Pick and Place 1K Dataset Dataset Description: The PhysicalAI-Robotics-GR00T-Teleop-G1 dataset consists of1000 teleoperation trajectories of real robot data using Unitree G1, with upper body control. The robot chooses the correct fruit to pick and place on the plate according to the language prompt. A total of 4 fruits are used: Apple, Pear, Starfruit, Grape. The robot is equipped with the default realsense camera, and a pair of Unitree G1 Tri-fingers… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-G1.tabular100K<n<1M26 likes1.7k downloads1y agoHugging Face10nvidia /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K42 likes1.5k downloads11d agoHugging Face11nvidia /GR00T-N1.7-AppleToPlate Dataset Description: The GR00T-N1.7-AppleToPlate dataset is a multimodal collection of trajectories collected on a Unitree G1 humanoid robot. It supports a humanoid (G1) static loco-manipulation task in which the robot picks up an apple and places it on a plate. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for an apple pick-and-place task. Dataset Name # Trajectories G1 Static… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GR00T-N1.7-AppleToPlate.tabularrobotics100K<n<1M4 likes1.5k downloads3mo agoHugging Face12nvidia /Nemotron-RL-Agentic-SWE-Pivot-v1 Dataset Description: The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The dataset is a refactored version of the SWE-Gym and R2E-Gym datasets to support the NeMo Gym input format. This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.tabular10K<n<100K15 likes1.4k downloads11d agoHugging Face13nvidia /LIBERO_LeRobot_v3 LIBERO LeRobot v3 Dataset Summary nvidia/LIBERO_LeRobot_v3 is a LeRobotDataset v3.0 conversion of the LIBERO robot manipulation benchmark. LIBERO is designed for studying lifelong robot learning and knowledge transfer across language-conditioned manipulation tasks. This repository packages the LIBERO task suites as LeRobot-compatible datasets with Parquet state/action data, MP4 video observations, and LeRobot metadata. The dataset is organized as five top-level… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LIBERO_LeRobot_v3.tabular100K<n<1M9 likes1.4k downloads4mo agoHugging Face14nvidia /Arena-G1-Loco-Manipulation-Task Dataset Description: The Arena-G1-Loco-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (G1) loco-manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for box pick and place task. Dataset Name # Trajectories G1 Loco-Manipulation Task 50 This dataset is ideal for behavior cloning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-G1-Loco-Manipulation-Task.tabularrobotics10K<n<100K6 likes1.3k downloads10mo agoHugging Face15nvidia /miracl-vision MIRACL-VISION MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark. This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.image100K<n<1M13 likes1.2k downloads1y agoHugging Face16nvidia /hifitts-2 HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset Dataset Description This repository contains the metadata for HiFiTTS-2, a large scale speech dataset derived from LibriVox audiobooks. For more details, please refer to our paper. The dataset contains metadata for approximately 36.7k hours of audio from 5k speakers that can be downloaded from LibriVox at a 48 kHz sampling rate. The metadata contains estimated bandwidth, which can be used to infer the original… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/hifitts-2.tabular10M<n<100M34 likes1k downloads11mo agoHugging Face17nvidia /Nemotron-Cascade-2-RL-data Dataset Description: The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data. This dataset is ready for commercial use. The dataset contains the following subset: IF-RL Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.tabular10K<n<100K54 likes996 downloads7mo agoHugging Face18nvidia /PhysicalAI-Robotics-Manipulation-Kitchen PhysicalAI Robotics Manipulation in the Kitchen Dataset Description: PhysicalAI-Robotics-Manipulation-Kitchen is a dataset of automatic generated motions of robots performing operations such as opening and closing cabinets, drawers, dishwashers and fridges. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen.tabularrobotics100K<n<1M14 likes828 downloads1y agoHugging Face19nvidia /compute-eval Dataset Card for ComputeEval ComputeEval is a benchmark for evaluating LLM-generated CUDA code on correctness and performance. Each problem provides a self-contained programming challenge — spanning kernels, runtime APIs, memory management, parallel algorithms, GPU libraries, and template/DSL kernel authoring — with a held-out test harness for functional validation and optional performance benchmarks for measuring GPU execution time against a baseline. Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/compute-eval.tabulartext-generation1K<n<10K30 likes810 downloads26d agoHugging Face20nvidia /aerial-isac-pusch-hest Aerial ISAC PUSCH Channel Estimates Uplink PUSCH DMRS channel estimates from a 5G CBRS cell running indoors on the NVIDIA Aerial testbed, paired with camera-derived floor positions of a person walking through the cell: 1.19 million estimates over 18 runs, nine with a person in the area and nine recorded empty as a background reference. Plan view in the label coordinate frame, with the recorded track of a clear run and of an obstacle run. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/aerial-isac-pusch-hest.tabularobject-detection100K<n<1M5 likes772 downloads1mo agoHugging Face21nvidia /PhysicalAI-Robotics-Manipulation-ObjectsPhysicalAI-Robotics-Manipulation-Objects is a dataset of automatic generated motions of robots performing operations such as picking and placing objects in a kitchen environment. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with Kinova Gen3 arms. The environments are kitchen scenes where the furniture and appliances were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Objects.tabularrobotics100K<n<1M8 likes637 downloads1y agoHugging Face22nvidia /aerial-isac-srs-iq Aerial ISAC SRS I/Q Raw uplink Sounding Reference Signal (SRS) I/Q captured on the NVIDIA Aerial 5G testbed, paired with a synchronized video and camera-derived ground truth for two pedestrians and a car moving through the sensing area. The labeled span in real time: camera view, Range-Doppler map, and range/velocity tracks. Also available as isac_rd_demo.mp4. Dataset Description: This dataset provides synchronized multi-modal recordings designed for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/aerial-isac-srs-iq.tabularobject-detectionn<1K11 likes610 downloads2mo agoHugging Face23nvidia /Arena-G1-Static-PickNPlace-Task Dataset Description: The Arena-G1-Static-PickNPlace-Task dataset is a multimodal collection of trajectories generated in Isaac Lab. It supports humanoid (G1) loco-manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for an apple pick-and-place task. Dataset Name # Trajectories G1 Static PickNPlace Task 200 This dataset is ideal for behavior… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-G1-Static-PickNPlace-Task.tabularrobotics10K<n<100K3 likes603 downloads3mo agoHugging Face24nvidia /Nemotron-RL-Instruction-Following-MultiTurnChat-v1 Dataset Description: The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.tabular1K<n<10K6 likes512 downloads12d agoHugging Face25dacorvo /funes-nvidia-Open-SWE-Traces-200-sessions Funes recall store — NVIDIA Open-SWE-Traces (200 sessions) A 200-session subset: a small funes recall store over 200 resolved==1 trajectories (the largest by content) of nvidia/Open-SWE-Traces, for quick evaluation while the full store is being built. Source nvidia/Open-SWE-Traces (CC-BY-4.0) Sessions 200 (largest resolved) Chunks 155361 Embedding model BAAI/bge-small-en-v1.5 (384-dim) Harnesses openhands, sweagent Full store… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-nvidia-Open-SWE-Traces-200-sessions.tabular100K<n<1M0 likes414 downloads9d agoHugging Face26nvidia /Arena-GR1-Manipulation-PlaceItemCloseDoor-Task Dataset Description: The Arena-GR1-Manipulation-PlaceItemCloseDoor-Task dataset is a multimodal collection of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation tasks in the IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, and action) needed to train and evaluate generalist robot policies for a sequential task (e.g. putting object into a fridge and closing the door). Dataset Name # Trajectories GR1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-PlaceItemCloseDoor-Task.tabularrobotics10K<n<100K2 likes315 downloads7mo agoHugging Face27nvidia /Arena-GR1-Manipulation-Task Dataset Description: The Arena-GR1-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for opening microwave task. Dataset Name # Trajectories GR1 Manipulation Task 50 This dataset is ideal for behavior cloning, policy learning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-Task.tabularrobotics1K<n<10K4 likes301 downloads10mo agoHugging Face28nvidia /PhysicalAI-GR00T-Tuned-Tasks Dataset Description: This dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) tabletop manipulation tasks for industrial settings. Each dataset entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for tasks like pouring nuts or sorting pipes by color. Dataset Name # Trajectories Exhaust-Pipe-Sorting-task 1000 Nut-Pouring-task 1000 This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-GR00T-Tuned-Tasks.tabularrobotics100K<n<1M3 likes289 downloads1y agoHugging Face29nvidia /Arena-GR1-Manipulation-Task-v3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "GR1", "total_episodes": 50, "total_frames": 4928, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-Task-v3.tabularrobotics1K<n<10K0 likes243 downloads10mo agoHugging Face30nvidia /Nemotron-RLHF-GenRM-v1 Dataset Description: This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking. The dataset is composed of: Preference data focused on diverse domains A synthetic safety blend The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.tabularreinforcement-learning100K<n<1M5 likes224 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.