datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Post-Training-Dataset-v1
Nemotron-Post-Training-Dataset-v1 Release
This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5.
Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model).
Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.Nemotron-Post-Training-Dataset-v2
Nemotron-Post-Training-Dataset-v2 Release
Data Overview
This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning.
NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.Llama-Nemotron-Post-Training-Dataset
Llama-Nemotron-Post-Training-Dataset-v1.1 Release
Update [4/8/2025]:
v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉
Data Overview
This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.gorilla-openfunctions-v1Luciole-PostTraining-Dataset-1.1
Table of Contents
Dataset Description
Curation Rationale
Bias, Risks, and Limitations
Data Subsets
Sample Metadata
Downloading the Data
Available Configurations
Loading Examples
Accessing Data Through the Directory Hierarchy
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.noncanonical-post-trainingNemotron-Post-Training-Dataset-v2-decontaminated
Decontamination
This dataset is a decontaminated version of nvidia/Nemotron-Post-Training-Dataset-v2.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Nemotron-Post-Training-Dataset-v2-decontaminated.Llama-Nemotron-Post-Training-Dataset-SFT-CoT-OnlyChallenge-phase1-dataset
Post-training for Robotics Foundation Models — Challenge Phase 1 Dataset
This is the public Phase 1 dataset for the RSS 2026 Workshop & Challenge on Post-training for Robotics Foundation Models.
Workshop website: https://posttraining-for-robotics.github.io/
The dataset contains real-robot bimanual manipulation trajectories on three benchmark tasks, collected on a bimanual YAM follower teleoperated by a GELLO leader arm. Every frame is timestamp-aligned across joint state, action… See the full description on the dataset page: https://huggingface.co/datasets/Posttraining-RFM-RSS2026/Challenge-phase1-dataset.Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format
This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
into the ShareGPT format while preserving the original splits and columns.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"messages": [
{"role": "user", "content": "User message"},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.mix-context-post-training-128k
Mix-Context Post-Training Dataset for 128K Context Extension
Overview
Mix-Context Post-Training 128K is a dataset designed specifically for post-training context window extension of pretrained LLMs.
It targets the stage after base pretraining, where a model is adapted to operate over much longer contexts (up to 128K tokens) while preserving short-context behavior. The dataset mixes short- and long-context packed sequences with a controlled length distribution to support:… See the full description on the dataset page: https://huggingface.co/datasets/ghostcc3/mix-context-post-training-128k.Salesforce-xlam-function-calling-60kSwallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.gliner-biomed-post-training
GLiNER-BioMed post-training dataset
This dataset, used for the post-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@article{yazdani2026gliner,
author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas},
title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-post-training.webui-training-datallama_nemotron_post_training_sft_sciencePuzzle-KD-Nemotron-Post-Training-Dataset-v2
Puzzle-KD-Nemotron-Post-Training-Dataset-v2 Release
Dataset Overview
The Puzzle-KD-Nemotron-Post-Training-Dataset-v2 dataset is a curated and filtered subset of NVIDIA’s Nemotron-Post-Training-Dataset-v2.
The original dataset was released by NVIDIA in August 2025 as part of the NVIDIA Nemotron Nano 9B model family and was designed to
improve post-training alignment for reasoning, math, code, STEM, and chat capabilities across multiple languages.
It contains synthetic… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Puzzle-KD-Nemotron-Post-Training-Dataset-v2.Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT
This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
converted to ShareGPT format and merged into a single dataset.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"original_split": "code|math|science|chat|safety",
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT.oracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-datasmall-mind-post-training-data
small-mind-companion — post-training data
Every corpus used to post-train a ~2B vision-language model
(google/gemma-4-E2B-it) for long-horizon personalised companion dialogue, in the order it was
used: LoRA SFT → LoRA DPO → on-policy distillation.
Part of the OneBee Datasets
collection.
Contents
Path
Rows
Schema
sft/v0/{train,val}.jsonl
202 / 23
messages
sft/v1/{train,val}.jsonl
2232 / 248
messages
dpo/v0/{train,val}.jsonl
200 / 23
prompt, chosen… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/small-mind-post-training-data.Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT
This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
converted to ShareGPT format and merged into a single dataset.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"original_split": "code|math|science|chat|safety",
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT.SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.AutoIF-instruct-61k-with-funcsLlama-Nemotron-Post-Training-Dataset-SFT-math-FI
Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset.
The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model.
Translation Process
The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.posttraining-eval-resultsNemotron-Post-Training-Dataset-v1
Nemotron-Post-Training-Dataset-v1 Release
This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5.
Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model).
Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency.… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/Nemotron-Post-Training-Dataset-v1.oracle_military_submarine_post_hoc_unmixed_fd-training-dataosim-post-training
SOUL
This CMU-LTI mirror hosts the post-training data used for ODYSSIM releases. It mirrors the original sunweiwei/Soul dataset layout under the CMU-LTI organization.
SOUL is the data suite for human behavior simulation used in
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO),
spanning conversation, social simulation, social cognition, role-play, and human-centric evaluation.
📄 Paper: https://arxiv.org/abs/2605.20506
💻 Code:… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/osim-post-training.typhoon-s-instruct-post-training
Typhoon-S Instruct Post-Training
Dataset Summary
This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths.
The dataset follows a two-part mixture philosophy:
Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-instruct-post-training.AutoIF-instruct-61k
