datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.Nemotron-Post-Training-Dataset-v1
Nemotron-Post-Training-Dataset-v1 Release
This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5.
Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model).
Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.CryoLithe-training-datasetThe training Dataset for CryoLithe Models
The dataset contains selected tilt series, tilt angles, and corresponding cryo-CARE+IsoNet and Icecream reconstructions using odd/even pairs. For EMPIAR-11058
Icecream reconstructions were obtained by splitting across angles.
Whenever available, we also provide dose-fractionated tilt series.
Dataset format:
Files ending with '.rawtlt' or '.tlt' correspond to the tilt angles.
Files ending with '_corrected.mrc' correspond to cryo-CARE+IsoNet… See the full description on the dataset page: https://huggingface.co/datasets/sada-group/CryoLithe-training-dataset.Nemotron-Post-Training-Dataset-v2
Nemotron-Post-Training-Dataset-v2 Release
Data Overview
This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning.
NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.lora-training-datasetsLlama-Nemotron-Post-Training-Dataset
Llama-Nemotron-Post-Training-Dataset-v1.1 Release
Update [4/8/2025]:
v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉
Data Overview
This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.VideoGPT-plus_Training_Datasete2e-stream-slam-training-dataset
e2e-stream-slam training assets
Reproducibility bundle for the V4 SLAMFormer ablation suite.
Code: https://github.com/SlamMate/e2e-semantic-SLAM/tree/submap (commit 3195a7a)
Contents
Checkpoints
File
Size
Role
checkpoints/v1_paper_ckpt10.pth
3.6 GB
SLAMFormer paper base ckpt (10 ep on the paper datasets). PRETRAINED init for V3 Scale Token training.
checkpoints/v3_scale_token_ckpt2.pth
3.8 GB
V3 Scale Token epoch-2 (3 ep, 3×A6000… See the full description on the dataset page: https://huggingface.co/datasets/qizhangslam/e2e-stream-slam-training-dataset.SII_self_evovling_02_training_datasetBitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.gaussian_training_datasets
Gaussian Training Datasets (COLMAP) for msplat
COLMAP-format multi-view scenes for training 3D Gaussian Splatting models,
packaged for msplat — a Metal-native 3DGS
trainer for Apple Silicon. Also includes pre-trained .ply splats under
tested_outputs/.
All scenes are redistributed from third-party datasets. Full credit goes to
their original authors — see Licensing & credits and please
cite the original papers. This repo only repackages them in COLMAP layout for
convenience.… See the full description on the dataset page: https://huggingface.co/datasets/alexmkwizu/gaussian_training_datasets.Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.GEWDiff_training_dataset
GEWDiff Training & Evaluation Dataset
📘 Overview
The GEWDiff Training & Evaluation Dataset is derived from the EnMAP Champion and MDAS hyperspectral datasets.It is designed for image enhancement, super-resolution, restoration, and generative remote sensing tasks.The dataset includes Low-Quality (LQ) low-resolution images, corresponding Ground-Truth (GT) high-resolution images, and optional structure information such as masks and edges (partially provided;… See the full description on the dataset page: https://huggingface.co/datasets/zhu-xlab/GEWDiff_training_dataset.solaris-training-dataset
Solaris Training Dataset
Project Page | Paper | Github
The training dataset collected via SolarisEngine to train the Solaris multiplayer world model for Minecraft.
Dataset Info
The dataset contains 12.64 M frames across two players. It was recorded at 20 FPS. The observation (video) dimensions are: 1280 × 720.
The action space is presented below:
Action key
Type
Description
forward
bool/sustainedPlayer moving forward (W).
back
bool/sustained
Player… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/solaris-training-dataset.All-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.song_dataset_training_20s_cleanedLlama-Nemotron-Post-Training-Dataset-SFT-CoT-OnlyChatTS-Training-Dataset
ChatTS-Training Data
This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model.
Datasets
align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256.
align_random: Alignment training dataset with random sequence lengths between 64 and 1024.
sft: SFT dataset generated with Time Series Evol-Instruct.
ift: Instruction following dataset.
dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.akd-video-training-datasetNemotron-Post-Training-Dataset-v2-decontaminated
Decontamination
This dataset is a decontaminated version of nvidia/Nemotron-Post-Training-Dataset-v2.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Nemotron-Post-Training-Dataset-v2-decontaminated.Elastic-Forcing-training-dataset
Elastic-Forcing training datasets
wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351.
Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales.
wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded).
Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT
Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format
This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1
into the ShareGPT format while preserving the original splits and columns.
Format
Each example contains all original fields plus a messages array:
{
"input": "original input text",
"output": "original output text",
... (other original columns) ...,
"messages": [
{"role": "user", "content": "User message"},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.Linny-Training-Dataset-Syntheticcomma-training-dataset-v1.0-chunkslegal-training-dataset
JudicialMind Legal Training Dataset
A large-scale, multilingual query–passage corpus for training and evaluating
legal information-retrieval and question-answering systems.
3.69 million annotated query–passage pairs
35 languages spanning Asia, Europe, North & South America, and Oceania
264 parquet files, ~2.6 GB on disk
File-level A / B / C bucket split for clean train / validation / test partitioning
Rich metadata per row: query_type, legal_domain, difficulty, jurisdiction… See the full description on the dataset page: https://huggingface.co/datasets/judicialmind/legal-training-dataset.Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.
