datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-student-fail-v41-clean-thinking
DeepSeek-V4.1 clean and action-only trajectories with Nemotron outcomes
DeepSeek-V4.1 reward-1 trajectories rebuilt from the complete teacher audit
under v57-test-path-component-boundary+v57-target-source-recheck. The V4.1 reward and trajectory tier do not by themselves prove
that Nemotron failed. Student outcomes are joined from
nemotron-prolike-coverage-audit-20261001.json. A student failure requires either complete
required-test results with reward 0, or an individually… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.fineinstructions_nemotron
✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions
This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline.
The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details.
Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.Nemotron-RL-Agentic-SWE-Pivot-v1
Dataset Description:
The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The dataset is a refactored version of the SWE-Gym and R2E-Gym datasets to support the NeMo Gym input format.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
22,927,812,461 (22.9B)
Trainable tokens
22,927,812,461 (22.9B)
Documents
21,377,358
Shards
180
UTF-8 bytes
77,994,866,327
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.Nemotron-Cascade-2-RL-data
Dataset Description:
The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data.
This dataset is ready for commercial use.
The dataset contains the following subset:
IF-RL
Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.grpo-qwen3-1.7b-nemotron-leetcode-clean-3.2k-bs32-n8-verl091-epoch2-146102-rollouts
Coding RL rollouts
grpo_Qwen3-1.7B_Nemotron-LeetCode-clean-3.2k_bs32_n8_seqs16_32k_epoch2_verl091
One verified gzip JSONL shard per training step; 256 responses per shard.
LCB binary grading after thinking, without an EOS gate.
Nemotron-RL-Instruction-Following-MultiTurnChat-v1
Dataset Description:
The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.nemotron-cc-v2.1-hq-dqa-qwen3-tokens
Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3-8B tokenizer
Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1
(High-Quality-DQA subset) and tokenized with the Qwen/Qwen3-8B tokenizer.
In the source data each row is a web document whose tail carries synthetic QA pairs marked
Question: / Answer:. Here that document is split into its original prose (context) and
the individual QA pairs, each tokenized separately. The Question: / Answer: marker keywords… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens.grpo-qwen3-1.7b-nemotron-leetcode-clean-3.2k-bs32-n8-verl091-146102-rollouts
Coding GRPO rollouts
grpo_Qwen3-1.7B_Nemotron-LeetCode-clean-3.2k_bs32_n8_seqs16_32k_1epoch_verl091
One verified gzip JSONL shard per training step; 256 responses per shard.
LCB binary grading after thinking, without an EOS gate.
Ornith-1.5-35B-A3B-Nemotron-v2-100M
Ornith 1.5 35B A3B Nemotron v2 100M
This dataset contains 108,729 English conversations with
108,729 regenerated assistant turns and 100,014,884 generated
assistant completion tokens. 100M refers to the completion-token target, not the
number of examples.
The prompt mix is a deterministic sample from
nvidia/Nemotron-Post-Training-Dataset-v2.
It covers the source dataset's chat, code, math, and STEM subsets. Every assistant turn
was regenerated with ornith-ai/Ornith-1.5-35B-A3B;… See the full description on the dataset page: https://huggingface.co/datasets/jzinno/Ornith-1.5-35B-A3B-Nemotron-v2-100M.nemotron-r1-en-ar-midtrain
nemotron-r1-en-ar-midtrain
Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.OpenReasoning-Nemotron-7B_eval_8179
mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
79.0
98.8
89.0
81.7
60.1
62.5
50.6
46.8
68.7
13.3
49.6
59.7
AIME24
Average Accuracy: 79.00% ± 1.42%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179.pretrain-nemotron-math-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
1,446,296,439 (1.4B)
Trainable tokens
1,446,296,439 (1.4B)
Documents
42,379
Shards
23
UTF-8 bytes
4,965,563,314
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.fineinstructions_nemotron
✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions
This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline.
The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details.
Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/Bas95/fineinstructions_nemotron.nemotron-mc-en-ar-midtrain
nemotron-mc-en-ar-midtrain
Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.Nemotron-RLHF-GenRM-v1
Dataset Description:
This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking.
The dataset is composed of:
Preference data focused on diverse domains
A synthetic safety blend
The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.Nemotron-RL-math-advanced_calculations
Dataset Description:
The Nemotron-RL-math-advanced_calculations is a dataset designed to test a model's ability to solve complex, multi-step math problems in a multi-step agentic environment. It involves counterintuitive calculations with varying levels of function composition.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-math-advanced_calculations.NVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresnemotron-nano-eval-logs-and-scoresnemotron_cc_v2_hq_packed4096_200shard
Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train)
Documents from nvidia/Nemotron-CC-v2
High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.nemotron-sft-balanced-2b-v1
Nemotron SFT Dataset
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Statistics
Total Samples: 200,000
Total Tokens: 1,252,287,904
Average Tokens per Sample: 6261.4
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Strategy: balanced
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Avg Tokens/Sample
Stage-1/math
20,000
151,546,125
20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.nemotron-sft-general-focused-stage1-2-ChatML-V3
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 496,385
Total Tokens: 1,114,218,401
Average Tokens per Sample: 2244.7
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V3.generations-nemotron-nano-9b-v2-simnpo-gentle-bm25-10bnemotron_cc_v2_hq_packed4096
Nemotron-CC-v2 High-Quality, packed to 4096 tokens
5% subset of nvidia/Nemotron-CC-v2
High-Quality documents, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id
input_ids… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096.Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.jailbreak-llama-3.3-nemotron-49b-v1.5Nemotron-Research-Reasoning-Qwen-1.5B_eval_569a
