datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.Llama-Nemotron-Post-Training-Dataset
Llama-Nemotron-Post-Training-Dataset-v1.1 Release
Update [4/8/2025]:
v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉
Data Overview
This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.VideoGPT-plus_Training_DatasetSII_self_evovling_02_training_datasetAll-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.ChatTS-Training-Dataset
ChatTS-Training Data
This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model.
Datasets
align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256.
align_random: Alignment training dataset with random sequence lengths between 64 and 1024.
sft: SFT dataset generated with Time Series Evol-Instruct.
ift: Instruction following dataset.
dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.Elastic-Forcing-training-dataset
Elastic-Forcing training datasets
wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351.
Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales.
wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded).
Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.dnd-35-training-dataset
D&D 3.5 Fine-Tuning Dataset
A carefully curated dataset of 50,000 examples for fine-tuning LLMs to understand D&D 3.5 mechanics.
Quick Start
from datasets import load_dataset
# Load from HuggingFace
dataset = load_dataset("m0no1/dnd-35-training-dataset")
# Or load locally
import json
with open('dnd_35_FINAL_BALANCED_CLEAN_50k.jsonl', 'r') as f:
data = [json.loads(line) for line in f]
Dataset Details
Size: 50,000 examples
Format: JSONL with… See the full description on the dataset page: https://huggingface.co/datasets/m0no1/dnd-35-training-dataset.SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.agent-training-dataset
🤖 Agent Training Dataset — Legendary Edition
The most comprehensive open-source dataset for training AI agents that actually work.
Built by Adewale David and his AI buddy.
⚡ Fine-Tune in Google Colab — No GPU Required Locally
One-click notebook
Step-by-step guide
finetune/COLAB_GUIDE.md
Evaluate your model
finetune/notebooks/evaluate_model.ipynb
Colab free tier (T4):Use Qwen2.5-3B-Instruct — trains in ~5 hrsColab Pro (L4/A100): Use… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/agent-training-dataset.Bitext-customer-support-llm-chatbot-training-dataset-spanish
Spanish Customer Support LLM Chatbot Training Dataset
Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset.
This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models.
Dataset Details
Dataset Description
This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.UniME-V2-Training-Datasets
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
Tiancheng Gu*,
Kaicheng Yang*,
kaichen Zhang,
Xiang An,
Ziyong Feng, Yueyi Zhang,
Weidong Cai,
Jiankang Deng,
Lidong Bing
🛠️ Implementation
git clone https://github.com/deepglint/UniME-v2.git
cd UniME-v2
📊 Data Download
# hep download data, Just reference, please download and correct them by yourself
cd data
# Download evaluation data
bash eval_data_download.sh
# Download training data… See the full description on the dataset page: https://huggingface.co/datasets/TianchengGu/UniME-V2-Training-Datasets.infosec-dataset-training
InfoSec Dataset Training
网络安全/信息安全领域训练数据集集合,汇聚多个开源安全数据集,提供统一的下载、转换和格式适配工具链。
数据集概览
数据集
来源
格式
语言
条目数
说明
cybersecurity_hq
自建
Alpaca
中文
20
网络安全基础问答
cybersecurity_sharegpt_chinese
ystemsrx/Cybersecurity-ShareGPT-Chinese
ShareGPT
中文
32,008
网络安全多轮对话
cybersecurity_chinese_mixed_v2
qingmian/CyberSecurity-Chinese-Mixed-V2
ShareGPT
中文
16,004
网络安全混合对话
trendyol_cybersecurity
Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Alpaca
英文
53,201
网络安全指令微调… See the full description on the dataset page: https://huggingface.co/datasets/lxcxjxhx/infosec-dataset-training.ru-open-llama-training-datasetsLlama-Nemotron-Post-Training-Dataset
Llama-Nemotron-Post-Training-Dataset-v1.1 Release
Update [4/8/2025]:
v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉
Data Overview
This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/Llama-Nemotron-Post-Training-Dataset.gaiasky-training-dataset
Gaia Sky Expert Dataset
This dataset is designed for fine-tuning Large Language Models to become experts in the Gaia Sky ecosystem. It covers 3D astronomical visualization, Java engine architecture, Python scripting API, and GLSL shader logic.
Dataset Structure
The repository is organized into two primary configurations:
1. Distilled (Instruction-Tuned)
File: train.jsonl
Format: {"instruction": "...", "output": "...", "source_file": "..."}
Description:… See the full description on the dataset page: https://huggingface.co/datasets/Langurmonkey/gaiasky-training-dataset.glyphmatics-complete-training-dataset
GlyphMatics Complete Training Dataset
Canonical synthetic training data for GlyphMatics / SigilAGI.
Covers
glyph encoding
glyph decoding
semantic compression
reconstruction
Alpha/Beta/Gamma mapping
SigilAGI routing
VIL normalization
GIIBL lattice blocks
RC3 cube encoding
Quantum Glyph states
mobile deployment planning
safety-aware symbolic transformation
Dataset Viewer
The public dataset viewer is configured only for:
data/train.jsonl
data/validation.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/glyphmatics-complete-training-dataset.nvidia_Llama-Nemotron-Post-Training-Dataset__science_2000KGFactExplainer-Training-Dataset
KGFactExplainer Training Dataset
This repository contains the generated training and validation data used
to fine-tune the Qwen3-8B model for the evidence-path generation component
of KGFactExplainer.
Dataset Description
The dataset contains teacher-generated supervision for generating
traversal-compatible evidence paths between summary facts and
source-document knowledge graphs.
The data were generated as part of the KGFactExplainer experimental
pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/nirasha/KGFactExplainer-Training-Dataset.hypencoder-msmarco-training-datasetThe MSMARCO training data used to train the models from Hypencoder: Hypernetworks for Information Retrieval
.
Dataset Overview
This dataset is based on the MSMARCO Passage dataset and includes all the queries which have a positive passage in the original dataset (there are additional queries with no positive passages which we do not use). Each query has the known positive passage as well as 200 additional passages. These additional passages may be unlabeled positives or negatives.… See the full description on the dataset page: https://huggingface.co/datasets/jfkback/hypencoder-msmarco-training-dataset.KAG-Thinker-training-datasetxyrus-cosmic-training-dataset-complete
🌌 Xyrus Cosmic Complete Training Dataset (Harmony Format)
Overview
The COMPLETE training dataset for Xyrus Cosmic GPT-OSS:20B, including all expansions and variations.
📊 Dataset Statistics
Total Unique Examples: 1781
Format: Harmony (GPT-OSS chat format)
Splits: Train (1424) / Val (178) / Test (179)
Dataset Components
xyrus_training_dataset.jsonl: 309 examples
xyrus_augmented_dataset.jsonl: 391 examples
xyrus_sdg_dataset.jsonl: 135 examples… See the full description on the dataset page: https://huggingface.co/datasets/ToddLLM/xyrus-cosmic-training-dataset-complete.Training-Ai-Islamic-Dataset
🕌 Training AI Islamic Dataset
18.7M passages from classical Islamic books spanning 1,400 years of scholarship.
Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training.
📊 Dataset Structure
collections/: Categorized Islamic passages compressed in JSONL format.
metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.All-CVE-Records-Training-Dataset-archive
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/All-CVE-Records-Training-Dataset-archive.zignet-training-dataset
ZigNet Training Dataset
Curated dataset of Zig programming examples for LLM fine-tuning
This dataset was created for the ZigNet project to train language models on Zig programming language patterns, idioms, and documentation.
Dataset Structure
Files
data/training/
├── dataset-train.jsonl # 9,629 examples (70%)
├── dataset-validation.jsonl # 2,063 examples (15%)
├── dataset-test.jsonl # 2,064 examples (15%)
└── dataset-stats.json # Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fulgidus/zignet-training-dataset.customer-service
Customer Service Conversations Dataset
This dataset contains 100 realistic customer service conversations between customers and support agents. Each dialogue is 11 turns long and covers a variety of common issues such as late deliveries, billing errors, account problems, and more. It is ideal for training and evaluating AI assistants, chatbots, and customer support models.
Dataset Structure
Each conversation is stored as a JSON object with the following fields:
id:… See the full description on the dataset page: https://huggingface.co/datasets/ai-training-datasets/customer-service.Llama-Nemotron-Post-Training-Dataset-v1Astral-1.5-Post-Training-Dataset-SFT
Astral 1.5 Post-Training Dataset
A albeit smaller, yet higher-quality reasoning dataset combining mathematics, code, and general stem used in the training of the Astral 1.5 model family.
Dataset Description
This dataset merges four datasets to create a high quality 25 thousand example dataset. With the size of the dataset, we rely on the principle that quality > quantity leads to better model performance.
Dataset Composition
Setup
General STEM:… See the full description on the dataset page: https://huggingface.co/datasets/LucidityAI/Astral-1.5-Post-Training-Dataset-SFT.training_dataset_8K_en
