datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-CoT
Dataset Card for NuminaMath CoT
Dataset Summary
Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT.Scaffold-CoT
Scaffold-CoT
Structured chain-of-thought training data with 3,726,548 examples in 76 JSONL shards.
Fields
Every row has exactly four top-level fields:
Field
Contents
metadata
domain, subdomain, difficulty, length_bucket
input
Ordered user messages as {index, content} objects
cot
Ordered {index, type, content} events, including reasoning, tool calls, and tool results
output
Ordered final assistant answers as {index, content} objects
The index… See the full description on the dataset page: https://huggingface.co/datasets/Specific-Labs/Scaffold-CoT.CoT-Collection"""
_LICENSE = "CC BY 4.0"
_HOMEPAGE = "https://github.com/kaistAI/CoT-Collection"
_LANGUAGES = {
"en": "English",
}
# _ALL_LANGUAGES = "all_languages"
class CoTCollectionMultiConfig(datasets.BuilderConfig):Math-CoT-44k-Qwen3-32b-n32-16384-with-logprob-and-entropy
Qwen3-32B Math n32 16384 (44k Queries)
This dataset contains multi-sampled rollout traces from Qwen3-32B on around 44k math queries.
For each query, the model is rolled out 32 times with a maximum generation length of 16384 tokens.
Each response is annotated with answer correctness (acc_reward), and includes token-level statistics (action_entropy, action_log_probs) for further analysis and research.
Resources
Paper: Rethinking Generalization in Reasoning SFT: A… See the full description on the dataset page: https://huggingface.co/datasets/jasonrqh/Math-CoT-44k-Qwen3-32b-n32-16384-with-logprob-and-entropy.cot-distill-lora-subspace
CoT distillation across interpolated teachers: teacher CoT, splits and student generations
This repository holds the data half of the paper on what a student's weights keep when its teacher stops thinking. The paper is [citation pending]. The LoRA adapters and analysis outputs are in the companion model repository, kimtaeyeong1229/cot-distill-lora-subspace.
Question. A teacher is built by mixing a reasoning model (λ = 0) and a short-answer instruct model (λ = 1) with ratio λ… See the full description on the dataset page: https://huggingface.co/datasets/kimtaeyeong1229/cot-distill-lora-subspace.stage3-final-mixture-cot50
Stage 3 Final Mixture — 50% CoT Compression
This is a deterministic capability-preserving rewrite of
leonli66/stage3-final-mixture for LCLM Stage-3 post-training.
Only the reasoning_data and dolci_think subsets change. Their
compression_prompt is the ordinary prompt. A deterministic 50% arm keeps
the complete assistant target as ordinary SFT; the other arm wraps the inferred
reasoning prefix in <|memory_start|>...<|memory_end|> while keeping the final
answer trainable. All… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-final-mixture-cot50.Aesir-Character-CoT-roleplay
Overview
Think with your role.
Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character.
Continue updating until money run out, I will try to update this dataset in near future
Stats
1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content)
~14,349 assistant turns, each with full character-POV reasoning
Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.cot-faithfulness-code-bug
Code bug faithfulness trajectories
The olmo3-v3 configuration contains 1,232 saved trajectories and judge annotations. The reasoning column wraps available reasoning text in <thinking> and </thinking> tokens. System and user prompts are separate.
from datasets import load_dataset
traces = load_dataset("shiv96/cot-faithfulness-code-bug", "olmo3-v3", split="train")
Cot-Drop
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.CoT-chemistry-SFT
CoT-chemistry-SFT
Full chemistry chain-of-thought (CoT) dataset for supervised fine-tuning (SFT), generated by o4-mini.
This is the complete 1,606-example dataset. A 100-example public preview is available at Arminzd/CoT-O4_mini.
Dataset Details
Examples: 1,606
Generated by: o4-mini
Purpose: SFT training for chemistry tool-calling agents (tool-n1 project)
Fields
Field
Description
uid=3154455(arminzd) gid=3154455(arminzd)… See the full description on the dataset page: https://huggingface.co/datasets/Arminzd/CoT-chemistry-SFT.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.Ascend-COT-v2-packed
AscendKernelGen/Ascend-COT-v2-packed
AscendKernelGen/Ascend-CoT-v2-packed contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-packed.AIME25-CoT-CN
Sci-Bench-AIME25'
This repo is a branch of Sci Bench made by IPF team. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path.
Brief intro
💻 Overview
A brief template and final report will be posted in Isaac's Blog
And the markdown template can be found in data/I_2
❓ Why we do this?
The multi-lingual datasets are scarce, while the CoT of Math is even less, no matter whether the CoT or the solution contains pictures… See the full description on the dataset page: https://huggingface.co/datasets/IPF/AIME25-CoT-CN.gsm8k-cot-120b
🚀 GSM8K-Teacher-CoT-120B (2025)
High-Quality Short Chain-of-Thought Distillation Dataset
Plain Text • No LaTeX • No ChatML • Deterministic Final Answers
This dataset provides high-quality short Chain-of-Thought (CoT) reasoning generated by OpenAI gpt-oss-120b on the GSM8K benchmark.
It is designed for small reasoning models (7B–14B).
The dataset is:
✔ plain text
✔ concise and deterministic
✔ fully normalized
✔ free of LaTeX, ChatML, XML, Markdown
✔ optimized for tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/HAD653/gsm8k-cot-120b.cot-hidden-state-trajectories
CoT Hidden-State Trajectories
Chain-of-thought traces and generation-time hidden-state activations from
11 open-weight language models, on Codeforces (competitive programming),
Hendrycks MATH, and SATBench (Boolean satisfiability).
This dataset accompanies the paper Reasoning Models Don't Just Think
Longer, They Move Differently (arXiv:2605.15454).
The paper asks whether reasoning-trained models follow different
hidden-state paths than matched instruction-tuned baselines, after… See the full description on the dataset page: https://huggingface.co/datasets/gjoelbye/cot-hidden-state-trajectories.Ascend-COT-v1
AscendKernelGen/Ascend-COT-v1
AscendKernelGen/Ascend-CoT-v1 contains a small subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v1.Ascend-COT-v2-json
AscendKernelGen/Ascend-COT-v2-json
AscendKernelGen/Ascend-CoT-v2-json contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-json.en-vocab-en-mnemonics-cotregulatory-compliance-cot-trial
🛑 Most Compliance Failures Are Not Missing-Rule Failures. They Are Reasoning Failures.
The model cites the right regulation — and still reaches the wrong conclusion.
This dataset trains ordered reasoning_steps with explicit conclusions, so the chain of custody from rule → application → decision is auditable, not implied.
Regulatory Compliance & Legal CoT Dataset — Official Open‑Source Evaluation Package (50 Rows Subset) by springofwindslabs
Full production volumes (1,000-row… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/regulatory-compliance-cot-trial.zebra-cot-mistral-small-3.2-24b-preprocessed
Zebra-CoT Preprocessed — Mistral Hackathon 2026
Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct.
Format
text: formatted as [INST] question [/INST] <think> reasoning </think> answer
image: PIL JPEG image for the corresponding visual task
Usage
Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning.
Hackathon
Created for Mistral Hackaton 2026 — Fine-tuning track with W&B.
cqa-creative-writing-expert-cot-preview
CQA: Creative Quality Alignment — Research-Grade Schema v2
English
This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.Multilingual-CoT-Collection"""
_LICENSE = "CC BY 4.0"
_HOMEPAGE = "https://github.com/kaistAI/CoT-Collection"
_LANGUAGES = {
"ko": "Korean",
"fr": "French",
"ru": "Russian",
"ja": "Japanese",
"zh": "Chinese",
}
# _ALL_LANGUAGES = "all_languages"
class CoTCollectionMultiConfig(datasets.BuilderConfig):Atlas-Think-Cot-12M
Atlas-Think-Cot-12M
Atlas-Think-Cot-12M is a large-scale, high-quality reasoning dataset curated for mathematical problem-solving, code generation, and scientific thinking. This dataset emphasizes step-by-step solutions and detailed reasoning, with a major share of mathematical problems guiding its structure and composition.
Mixture of Mathematics, Coding, and Science. [ <:think>/cot ]
Quick Start with Hugging Face Datasets🤗
pip install -U datasets… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Atlas-Think-Cot-12M.Ascend-CoT-v3-json
Ascend-CoT-v3-json
Ascend-CoT-v3-json is an Ascend C / CANN supervised fine-tuning dataset for custom operator development. It contains cleaned CoT-style samples for Ascend C kernel implementation, tiling logic, CANN API usage, debugging, and operator-development reasoning.
The release is organized into two final SFT subsets in one dataset repository.
Related Artifacts
Paper: AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-CoT-v3-json.CoTLab_data
CoTLab: Medical Chain-of-Thought & Reasoning Benchmarks
A standardized benchmark suite for evaluating Chain-of-Thought (CoT) reasoning, faithfulness, and clinical domain knowledge in large language models.
Component Datasets & Upstream Sources
This repository aggregates standardized evaluation splits, MCQ questions, and reasoning pairs from the following open medical benchmarks:
MedQA (USMLE): Clinical board examination questions (Jin et al., Applied Sciences… See the full description on the dataset page: https://huggingface.co/datasets/huseyincavus/CoTLab_data.CoT-XLangRU:CoT-XLang — это многоязычный датасет, состоящий из текстовых примеров с пошаговыми рассуждениями (Chain-of-Thought, CoT) на различных языках, включая английский, русский, японский и другие. Он используется для обучения и тестирования моделей в задачах, требующих пояснений решений через несколько шагов. Датасет включает около 2,419,912 примеров, что позволяет эффективно обучать модели, способные генерировать пошаговые рассуждения.
Рекомендация:Используйте датасет для обучения моделей… See the full description on the dataset page: https://huggingface.co/datasets/Egor-3926/CoT-XLang.AIME25-CoT-CN
Sci-Bench-AIME25'
This repo is a branch of Sci Bench made by IPF team-SnailAILab. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path.
📚 Cite
If you use the Sci-Bench-AIME25 (IPF/AIME25-CoT-CN) dataset in your research, please cite:
@dataset{zhang2025scibench_aime25,
title = {{Sci-Bench-AIME25}: A Multi-Modal Chain-of-Thought Dataset for Advanced Tool-Intergrated Mathematical Reasoning},
author = {Zhang, Haoxiang and Wang, Siyuan… See the full description on the dataset page: https://huggingface.co/datasets/SnailAILab/AIME25-CoT-CN.salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.atcoder_cot
Dataset Card for Atcoder-CoT
Dataset Description
Atcoder-CoT is a proof-of-concept dataset designed to demonstrate how a dataset like the one found here can be used to generate synthetic datasets for training reasoning models, particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation. It leverages human-created and debugged solutions, combined with LLM-generated text to create conversational turns. The approach can also be easily adapted to simulate human… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_cot.olympiad-math-cot
Olympiad Math — CoT Distillation Dataset
Chain-of-Thought solutions for olympiad-level math problems,
distilled from stronger models (Claude, GPT via OpenRouter)
on top of human-authored problem+answer pairs.
Used to fine-tune local 9B models (GLM-Z1-9B, Qwen3.5-9B) via LoRA SFT.
Dataset Files
File
Examples
Description
data/sft_train.jsonl
22,990
Main SFT set — deduplicated good solutions
data/dpo_pairs.jsonl
4,393
DPO pairs — chosen (complete) vs… See the full description on the dataset page: https://huggingface.co/datasets/NecroMOnk/olympiad-math-cot.
