datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.Nemotron-Content-Safety-Audio-Dataset
Nemotron Content Safety Audio Dataset
Dataset Description
The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories.
LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval
VoMP: Predicting Volumetric Mechanical Properties
Dataset Description:
The Pre-Processed 3D Dataset is a dataset that is composed of 4 individual 3D asset datasets which are processed to render them from multiple views, voxelize the assets, and propagate VLM annotations for material properties.
We release pre-processed data derived from the 3D assets, specifically: voxels, rendered images, and LLM-annotated material descriptions.
This dataset is for research and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval.dspace-tuv-data
Stationary-start TUV test clips
This dataset contains 86 clips with verified zero-speed starts and requested event durations below 35 seconds, plus one explicitly selected 35-second CCRb clip, across CCRs (stationary target), CCRm (moving target), and CCRb (braking target).
CCRs: 31, CCRm: 50, CCRb: 5
Start here
Read clips.csv to select a clip by scenario, speed, or duration. All paths in the manifest are relative to this dataset root. files.csv lists each payload… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/dspace-tuv-data.nvidia-kagglenvidia-nemotron-model-reasoning-dataset-turkish
Nemotron Reasoning Challenge - Turkish
Turkish translation of the training data from NVIDIA's Nemotron Model Reasoning Challenge
Each row is a reasoning puzzle framed in an "Alice's Wonderland" setting. Given a few input/output examples, the model needs to figure out the hidden rule and apply it to a new input.
Category
Rows
Description
bit
1602
Hidden bit manipulation rule on 8-bit binary numbers
grav
1597
Falling distance with a modified gravitational constant… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/nvidia-nemotron-model-reasoning-dataset-turkish.NVIDIA-Nemotron-Model-Reasoning-Challengenvidia-qaNvidia Documentation Question and Answer pairs
Q&A dataset for LLM finetuning about the NVIDIA about SDKs and blogs
This dataset is obtained by generating Q&A pairs from a few NVIDIA websites such as development kits and guides. This data can be used to fine-tune any LLM for indulging knowledge about NVIDIA into them.
Source: https://www.kaggle.com/datasets/gondimalladeepesh/nvidia-documentation-question-and-answer-pairs
ko-nvidia-qnanvidia_finalnvidia_shortnvidia_cleanednvidia_original_cleaned
