Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Post-Training-Dataset-v1 Nemotron-Post-Training-Dataset-v1 Release This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5. Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.text10M<n<100M194 likes14k downloads1y agoHugging Face02nvidia /Nemotron-Post-Training-Dataset-v2gated Nemotron-Post-Training-Dataset-v2 Release Data Overview This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning. NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.text1M<n<10M156 likes6.5k downloads1y agoHugging Face03nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M708 likes6.4k downloads1y agoHugging Face04Post-training-Data-Flywheel /gorilla-openfunctions-v1text10K<n<100K0 likes2.2k downloads2y agoHugging Face05OpenLLM-France /Luciole-PostTraining-Dataset-1.1 Table of Contents Dataset Description Curation Rationale Bias, Risks, and Limitations Data Subsets Sample Metadata Downloading the Data Available Configurations Loading Examples Accessing Data Through the Directory Hierarchy Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.text1M<n<10M6 likes1.6k downloads5d agoHugging Face06brendanlong /noncanonical-post-training0 likes1.4k downloads17d agoHugging Face07openeurollm /Nemotron-Post-Training-Dataset-v2-decontaminated Decontamination This dataset is a decontaminated version of nvidia/Nemotron-Post-Training-Dataset-v2. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Nemotron-Post-Training-Dataset-v2-decontaminated.text1M<n<10M1 likes1k downloads6mo agoHugging Face08di-zhang-fdu /Llama-Nemotron-Post-Training-Dataset-SFT-CoT-Only0 likes944 downloads1y agoHugging Face09Posttraining-RFM-RSS2026 /Challenge-phase1-dataset Post-training for Robotics Foundation Models — Challenge Phase 1 Dataset This is the public Phase 1 dataset for the RSS 2026 Workshop & Challenge on Post-training for Robotics Foundation Models. Workshop website: https://posttraining-for-robotics.github.io/ The dataset contains real-robot bimanual manipulation trajectories on three benchmark tasks, collected on a bimanual YAM follower teleoperated by a GELLO leader arm. Every frame is timestamp-aligned across joint state, action… See the full description on the dataset page: https://huggingface.co/datasets/Posttraining-RFM-RSS2026/Challenge-phase1-dataset.video1K<n<10K2 likes934 downloads4mo agoHugging Face10MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 into the ShareGPT format while preserving the original splits and columns. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "messages": [ {"role": "user", "content": "User message"}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.text10M<n<100M41 likes734 downloads1y agoHugging Face11ghostcc3 /mix-context-post-training-128k Mix-Context Post-Training Dataset for 128K Context Extension Overview Mix-Context Post-Training 128K is a dataset designed specifically for post-training context window extension of pretrained LLMs. It targets the stage after base pretraining, where a model is adapted to operate over much longer contexts (up to 128K tokens) while preserving short-context behavior. The dataset mixes short- and long-context packed sequences with a controlled length distribution to support:… See the full description on the dataset page: https://huggingface.co/datasets/ghostcc3/mix-context-post-training-128k.text-generation10K<n<100K3 likes674 downloads9mo agoHugging Face12Post-training-Data-Flywheel /Salesforce-xlam-function-calling-60ktext10K<n<100K0 likes644 downloads2y agoHugging Face13tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes595 downloads8mo agoHugging Face14anthonyyazdaniml /gliner-biomed-post-training GLiNER-BioMed post-training dataset This dataset, used for the post-training stage of the GLiNER-BioMed models, was introduced in the paper GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @article{yazdani2026gliner, author = {Yazdani, Anthony and Stepanov, Ihor and Teodoro, Douglas}, title = {{GLiNER-BioMed}: a suite of efficient… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-post-training.text10K<n<100K0 likes368 downloads3mo agoHugging Face15post-train /webui-training-dataimage1K<n<10K0 likes365 downloads7mo agoHugging Face16hamishivi /llama_nemotron_post_training_sft_sciencetext100K<n<1M0 likes356 downloads1y agoHugging Face17nvidia /Puzzle-KD-Nemotron-Post-Training-Dataset-v2 Puzzle-KD-Nemotron-Post-Training-Dataset-v2 Release Dataset Overview The Puzzle-KD-Nemotron-Post-Training-Dataset-v2 dataset is a curated and filtered subset of NVIDIA’s Nemotron-Post-Training-Dataset-v2. The original dataset was released by NVIDIA in August 2025 as part of the NVIDIA Nemotron Nano 9B model family and was designed to improve post-training alignment for reasoning, math, code, STEM, and chat capabilities across multiple languages. It contains synthetic… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Puzzle-KD-Nemotron-Post-Training-Dataset-v2.text100K<n<1M11 likes342 downloads11mo agoHugging Face18MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 converted to ShareGPT format and merged into a single dataset. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "original_split": "code|math|science|chat|safety", "messages": [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT.text1M<n<10M3 likes324 downloads2y agoHugging Face19surrogate-base-model /oracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-data0 likes275 downloads1mo agoHugging Face20arjhinety /small-mind-post-training-data small-mind-companion — post-training data Every corpus used to post-train a ~2B vision-language model (google/gemma-4-E2B-it) for long-horizon personalised companion dialogue, in the order it was used: LoRA SFT → LoRA DPO → on-policy distillation. Part of the OneBee Datasets collection. Contents Path Rows Schema sft/v0/{train,val}.jsonl 202 / 23 messages sft/v1/{train,val}.jsonl 2232 / 248 messages dpo/v0/{train,val}.jsonl 200 / 23 prompt, chosen… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/small-mind-post-training-data.text-generation1K<n<10K0 likes227 downloads22d agoHugging Face21MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 converted to ShareGPT format and merged into a single dataset. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "original_split": "code|math|science|chat|safety", "messages": [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT.text1M<n<10M3 likes220 downloads2y agoHugging Face22amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes210 downloads10mo agoHugging Face23Post-training-Data-Flywheel /AutoIF-instruct-61k-with-funcstext10K<n<100K8 likes207 downloads2y agoHugging Face24LumiOpen /Llama-Nemotron-Post-Training-Dataset-SFT-math-FI Llama-Nemotron-Post-Training-Dataset-SFT-math-FI This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset. The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model. Translation Process The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.texttext-generation1M<n<10M1 likes197 downloads3mo agoHugging Face25benchpress /posttraining-eval-results0 likes197 downloads5mo agoHugging Face26nick007x /Nemotron-Post-Training-Dataset-v1 Nemotron-Post-Training-Dataset-v1 Release This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5. Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency.… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/Nemotron-Post-Training-Dataset-v1.text10M<n<100M0 likes196 downloads6mo agoHugging Face27model-organisms-for-real /oracle_military_submarine_post_hoc_unmixed_fd-training-data0 likes176 downloads4mo agoHugging Face28cmu-lti /osim-post-training SOUL This CMU-LTI mirror hosts the post-training data used for ODYSSIM releases. It mirrors the original sunweiwei/Soul dataset layout under the CMU-LTI organization. SOUL is the data suite for human behavior simulation used in Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO), spanning conversation, social simulation, social cognition, role-play, and human-centric evaluation. 📄 Paper: https://arxiv.org/abs/2605.20506 💻 Code:… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/osim-post-training.texttext-generation10K<n<100K1 likes172 downloads4mo agoHugging Face29typhoon-ai /typhoon-s-instruct-post-training Typhoon-S Instruct Post-Training Dataset Summary This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths. The dataset follows a two-part mixture philosophy: Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-instruct-post-training.texttext-generation100K<n<1M0 likes167 downloads8mo agoHugging Face30Post-training-Data-Flywheel /AutoIF-instruct-61ktext10K<n<100K18 likes165 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.