datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KodCode-V1-SFT-R1
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.KodCode-V1-SFT-4o
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-4o.pa-warm-start-sft-xl-50b-mix
geodesic-research/pa-warm-start-sft-xl-50b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.pa-warm-start-sft-heavy-25b-mix
geodesic-research/pa-warm-start-sft-heavy-25b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-heavy-25b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-heavy-25b-mix.barbet-long-context-sft
Barbet long-context SFT
Release a9fe3ba5b4c2869855e4e75118262797572d695c146eaa8fed7a180ec3544381 preserves 6633 active records. This is one joint
assistant-only SFT dataset; no Barbet model training has been run.
The skill-prefill migration has revised 1245
of 1254 records from its fixed base snapshot.
Revisions replace their original records in the explicit shard lists above. Old
bundles and releases remain available at their pinned commits. Additional records
from other… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/barbet-long-context-sft.Eurus-2-7B-SFT_eval_2e29
mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
2.3
21.0
30.6
11.0
11.4
10.4
6.8
1.5
2.1
1.3
4.1
4.4
AIME24
Average Accuracy: 2.33% ± 0.67%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
30
2
3.33%
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.sorrel-sft-voicepa-warm-start-sft-xl-smokeDolci-Think-SFT-translated
Dolci-Think-SFT-translated
Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations.
Columns
Each row is a translated conversation plus the result of a post-translation quality filter:
id — source record id.
messages — the translated conversation (list of {content, role}).
filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tts-realspeech-sft-en-de
LAION TTS Real-Speech SFT — English + German, emotion-balanced
1,947,272 real recorded utterances — no synthetic voices — selected from freely-licensed corpora
and balanced across 40 emotions x 2 languages. 6,996 hours,
313,844,544 MOSS frames (3,766,134,528 audio tokens), 79,337,527 aligned words.
Each row is a self-contained TTS example: a corrected procedural caption, the transcript with
word-level timestamps, the original audio, and the target MOSS-Audio-Tokenizer-v2 codes.… See the full description on the dataset page: https://huggingface.co/datasets/laion/tts-realspeech-sft-en-de.pa-warm-start-sft-xl-calibrationpa-warm-start-sft-heavy-25b-mix-longpa-warm-start-sft-xl-50b-mix-metagaming-filteredSmolDataEnvs-sft
🛠️ SmolDataEnvs: SFT
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
4,677 worked examples of an agent doing data science the right way. Each row is a complete,
verified-correct trajectory: read the question, poke at the data with a shell tool… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-sft.turkce-sft-qa-3.7m
🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti
3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde
tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri
setinden geldiğini taşır.
English: A merged, row-level deduplicated and quality-filtered collection of
24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries
its source dataset, source URL and original license.
🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.0399-tv-valid-clean-sft-tokenized-llmjp4-8bChinese-DeepSeek-R1-Distill-data-110k-SFT
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下:
Math:共计36568个样本,
Exam:共计2432个样本,
STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.ReasonXL-SFT
ReasonXL: A Multilingual Cross-Domain Reasoning Corpus
ReasonXL is a large-scale multilingual reasoning corpus spanning five languages, with 2,538,450 positionally aligned examples per language (12,692,250 rows total). It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains.
Data Generation
English source samples were drawn from 10 existing reasoning datasets, filtered and… See the full description on the dataset page: https://huggingface.co/datasets/toroe/ReasonXL-SFT.smolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.gigaverbo-v2-sft
GigaVerbo-v2 SFT: A Large-Scale Portuguese Instruction-Tuning Dataset
Dataset Summary
GigaVerbo-v2 SFT is a large-scale instruction-tuning dataset designed for supervised fine-tuning of language models in Portuguese. The dataset comprises approximately 2.1 billion tokens (~4.4 GB) across 4 million instruction-following examples, organized into 12 distinct task categories. It is entirely composed of high-quality, LLM-generated data that has been carefully curated and… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-sft.openswe-success-sft-v1
Open-SWE-Traces successes — cleaned and tokenized for MiniCPM5 SFT (v1)
Successful software-engineering agent trajectories from
nvidia/Open-SWE-Traces (revision
f8fb5b3d2c787f85f8a00f5fe04fe3f1a11088ef), filtered, validated and pre-tokenized with the native
MiniCPM5-2B-Midtrain tokenizer and chat template
(revision 0a45344e) for supervised fine-tuning with a 131,072-token context.
Split
Trajectories
Unique tasks
Repositories
Input tokens
Supervised tokens
Longest… See the full description on the dataset page: https://huggingface.co/datasets/LingweiGu/openswe-success-sft-v1.VR-X-SFT-RL
VR-X: Visual Reasoning Benchmark for UniVR
VR-X contains three independent data blocks:
SFT data organized by capability.
VR-X-RL data for visual-reasoning reinforcement learning.
VR-X-Eval held-out evaluation data.
VR-X-RL and VR-X-Eval are independent from SFT and must be loaded separately.
Public repository paths use anonymous source codes; no source-to-code mapping is
published.
Repository layout
.
├── Robot Manipulation/ # SFT only
│ └── RM-###/
│… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/VR-X-SFT-RL.fable-5-sft-traces
Fable-5 SFT Traces
Author / maintainer: kelexine (github.com/kelexine)
A cleaned, anonymised, schema-normalised derivative of
Kelexine/Fable-5-traces
— agentic traces from Fable-5 (claude-fable-5), the model now publicly
known as Claude Mythos — Anthropic's top-of-family frontier model at time
of collection.
The dataset supports three fine-tuning shapes off a single JSONL with no
preprocessing required:
Mode
Fields used
Full SFT (thinking + response)
messages or… See the full description on the dataset page: https://huggingface.co/datasets/kelexine/fable-5-sft-traces.pa-warm-start-sft-xl-1b-smokereadall-sft-stage-a-b
ReadAll / ReadTwice SFT:Stage A + Stage B
当前 ReadAll 模型的 SFT 数据和可移植训练包。训练链为:
Qwen/Qwen2.5-7B-Instruct → Stage A (ReadTwice step286) → Stage B (ReadAll union step448)
阶段
训练行数
验证行数
Parquet 分片
全局 batch
学习率
1 epoch 更新数
A
73,416
1,676
15 + 1
256
1e-5
286
B
57,287
无独立验证集
29
128
5e-7
448
这些计数是 SFT 消息样本行数,包含 SKIM / UPDATE / FINAL,并非独立问题数。45 个 Parquet 共 1,030,251,959 字节。原始数据分片和 manifest 原样保留,逐一核对原始 SHA-256;没有删列、重新筛选或重新生成。
Stage A 使用普通多轮 assistant-token SFT、12,288 token… See the full description on the dataset page: https://huggingface.co/datasets/Xirui1208/readall-sft-stage-a-b.combined_gsm8k_math_dataset_dapo_math_17k_Qwen3-4B_ntokens2048_sftbhasha-sft
Bhasha SFT
Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual
Large Language Models. The dataset contains collation of over 13 million instances of
instruction-response data for 3 Indian languages (Hindi, Gujarati, Bengali) and English having both human annotated and synthetic data.
Curated by: Soket AI Labs
Language(s) (NLP): [English, Hindi, Bengali, Gujarati]
License: [cc-by-4.0, apache-2.0, mit]… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/bhasha-sft.UMM-Reflection-SFT-Data
UMM-Reflection SFT Data
The reflection-SFT data of
UMM-Reflection (Learning
Native Reflection in Unified Models). It trains
UMM-Reflection-BAGEL-SFT.
Research use only, non-commercial. The rows are derived from datasets
with different licenses, some of them non-commercial. Each row records its
source dataset and license in source_dataset and source_license, and
each row follows the terms of its source. See LICENSE.md.
Contents
Part
Rows
Shards
Size… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/UMM-Reflection-SFT-Data.
