datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-student-fail-v41-clean-thinking
DeepSeek-V4.1 clean and action-only trajectories with Nemotron outcomes
DeepSeek-V4.1 reward-1 trajectories rebuilt from the complete teacher audit
under v57-test-path-component-boundary+v57-target-source-recheck. The V4.1 reward and trajectory tier do not by themselves prove
that Nemotron failed. Student outcomes are joined from
nemotron-prolike-coverage-audit-20261001.json. A student failure requires either complete
required-test results with reward 0, or an individually… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.UltraData-SFT-2605-no-think-8k-32k
UltraData-SFT-2605 · no_think · 8k–32k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 8k–32k range.
This is the medium-length tier intended for standard long-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.epic-thinking
Source
Rows
glaiveai/reasoning-v1-20m
1 999 793
BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples CC
1 267 534
PrimeIntellect/INTELLECT-3-SFT openreasoning_science
1 000 000
PrimeIntellect/INTELLECT-3-SFT am_chat
852 816
nvidia/Nemotron-Cascade-SFT-Stage-1 general
583 612
open-thoughts/OpenThoughts2-1M
541 898
PrimeIntellect/SYNTHETIC-1-SFT-Data
474 810
allenai/Dolci-Think-SFT-7B
334 908
allenai/Dolci-Think-SFT-32B
327 491
GeneralReasoning/GeneralThought-430K
291 946… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epic-thinking.thinking-cap-tier-curricula-complete
Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge:
Zero <|pad|> batch residues: 100% eliminated across all files.
Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.UltraData-SFT-2605-no-think-32k-200k
UltraData-SFT-2605 · no_think · 32k–200k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 32k–200k range.
This is the long-context tier intended for extended-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
this repo — fxmeng/UltraData-SFT-2605-no-think-32k-200k
32k–200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-32k-200k.grade_school_math_thinkingthinking-cap-tier-raw-traces
Thinking Cap Tier Raw Traces (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized:
Zero batch-padding residues (<|pad|>): Completely purged across all records.
Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.thinking-cap-tier-lima-dense
Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge:
Zero <|pad|> batch residues: 100% eliminated across all records.
Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions.
Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.open_parallel_think_cot_update_wo_answer
open_parallel_think_cot_update_wo_answer
This dataset is derived from haowu89/open_parallel_think_cot_update.
Transformation applied:
For every example, for every string item inside context, remove the final sentence.
The intent is to strip the trailing answer-bearing sentence while keeping the earlier reasoning trajectory.
Generated on 2026-04-15.
thinking-steering-vectorsgrug-think
grug-think
grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good.
big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket.
what in box
100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.llm-jp-4-thinking-sft-data-chatmlllm-jpのデータセットllm-jp-4-thinking-sft-dataを、
ChatML形式に変換したものです。
ライセンス
各サンプルのライセンスは、元データセットカードに記載された各データソースのライセンスに従います。
本リポジトリは、元となったデータ全体に対して新たなライセンスを付与するものではありません。
利用する場合は、対応する元データソースのライセンス条件を確認してください。
Nieto-2022-ThinkingOutLoudOpenAccessEEGBasedBCIDatasetInnerSpeech
Thinking out loud: an open-access EEG-based BCI dataset for inner speech recognition
This is an unofficial mirror of OpenNeuro dataset ds003626, version
2.1.2. It is not affiliated with or endorsed by the dataset authors, their
institutions, or OpenNeuro.
Source and documentation
Original dataset: OpenNeuro ds003626 v2.1.2
Article: Nieto et al., Scientific Data (2022)
Original dataset documentation: README
Official analysis code: N-Nieto/Inner_Speech_Dataset
The… See the full description on the dataset page: https://huggingface.co/datasets/raei/Nieto-2022-ThinkingOutLoudOpenAccessEEGBasedBCIDatasetInnerSpeech.claude_opus_4.8_max_thinking_5k_v2
Claude Opus 4.8 MAX THINKING — Distillation Dataset
5,000 high-quality examples designed to distill the maximum-effort reasoning, honest analysis, production software engineering, and agentic capabilities of Claude Opus 4.8.
Overview
This dataset captures Opus 4.8’s signature strengths:
Deep, structured, high-effort reasoning
Honest communication about trade-offs and uncertainties
Excellent production software engineering judgment
Strong agentic workflow design… See the full description on the dataset page: https://huggingface.co/datasets/11-47/claude_opus_4.8_max_thinking_5k_v2.ablation_nemotron_no_thinking
Dataset: ablation_nemotron_no_thinking
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/ablation_nemotron_no_thinking/stage_1/tmp.
HQ-Chat-2k
🧠 HQ-Chat-2K — High-Quality Conversational & Instruction-Tuning Dataset
2,000 carefully curated, high-quality conversation and instruction examples for fine-tuning Small Language Models (SLMs) and compact LLMs from ~500M to 3B parameters.
HQ-Chat-2K is a high-quality conversational and instruction-tuning dataset designed specifically for training and fine-tuning small to medium-sized Large Language Models (LLMs).
The dataset contains 2,000 curated user–assistant examples… See the full description on the dataset page: https://huggingface.co/datasets/ThinkNet/HQ-Chat-2k.gpt-oss-120b-mandarin-thinking-eval-logs-and-scoresgpt-oss-20b-mandarin-thinking-eval-logs-and-scoresOpenThought3-Qwen3-4BOpenThought3-Qwen3-4B
OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format.
Data Creation and Cleaning
This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.optimal_thinking_benchThis dataset is released as part of OptimalOptimalThinkingBench research project.
IMPORTANT: This is only a subset of OptimalThinkingBench that does not contain the math problems. To download the full dataset, please refer to our project materials here for more details.
Loading the dataset with transformers
This dataset is built using Llama-4-Maverick and Reasoning-Gym. Details on how to generate this dataset can be found in OptimalOptimalThinkingBench paper.
Minimal example below… See the full description on the dataset page: https://huggingface.co/datasets/facebook/optimal_thinking_bench.gsm8k-thinkingGSM8K with a reasoning/thinking/reflecting done by:
new: llama3.1-405b
old: chatgpt-4o-lastest
the new one is not complete. [at 50~% as of now]
medra-thinking-768CT-RATE-Thinking
CT-RATE-Thinking: Reasoning-Augmented CT Report Dataset
🎉🎉🎉 Our paper was accepted at the 28th conference of The Medical Image Computing and Computer Assisted Intervention Society (MICCAI). See you in Daejeon, Korea, September 23–27, 2025.CT-RATE-Thinking is a reasoning-augmented dataset derived from CT-RATE, containing chain-of-thought VQA pairs and report-level thinking narratives for 3D chest CT volumes.
It was generated as part of the μ²Tokenizer project… See the full description on the dataset page: https://huggingface.co/datasets/AlpachinoNLP/CT-RATE-Thinking.Think-Bench
THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
Official repository for "THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models".
For more details, please refer to the project page with dataset exploration and visualization tools.
[Paper] [Github] [ModelScope Dataset] [Visualization]
👀 About Think-Bench
Reasoning models have made remarkable progress in complex tasks… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuan218/Think-Bench.GLM-5.2-FP8-nemotron-codealpaca-thinking
GLM-5.2-FP8 Nemotron-CodeAlpaca Thinking Dataset
820,790 single-turn conversations generated by zai-org/GLM-5.2-FP8
with thinking enabled.
Prompt source
Rows (public)
Nemotron-Post-Training-Dataset-v2
800,944
CodeAlpaca-20k (corrected prompts, instruction + "\n\n" + input)
19,846
Total
820,790
Generation: temperature=1.0, top_p=0.95, max_tokens=24576, thinking
enabled. The CodeAlpaca prompts here include the input field.
Relationship to… See the full description on the dataset page: https://huggingface.co/datasets/JessieWei/GLM-5.2-FP8-nemotron-codealpaca-thinking.TeichAI-thinking-reasoning-x
TeichAI Thinking & Reasoning Datasets
A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled.
These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows.
Schema
Each row in the dataset has the following fields:
question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication.
question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.thinkflow-vla-features-b2Dataset_of_Russian_thinkingRu
RTD
Описание:Russian Thinking Dataset — это набор данных, предназначенный для обучения и тестирования моделей обработки естественного языка (NLP) на русском языке. Датасет ориентирован на задачи, связанные с генерацией текста, анализом диалогов и решением математических и логических задач.
Основная информация:
Сплит: train
Количество записей: 147.046
Цели:
Обучение моделей пониманию русского языка.
Создание диалоговых систем с естественным взаимодействием.… See the full description on the dataset page: https://huggingface.co/datasets/qwqeqw/Dataset_of_Russian_thinking.thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3
ThinkingCap Condensed — Qwen3.8 / GLM-5.2 / Kimi-K3
Condensed ThinkingCap-style reasoning traces for SFT.
1,985 traces: each row pairs a full multi-turn teacher trace (Qwen3.8-Max,
GLM-5.2 or Kimi K3, via
r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation)
with a condensed TC-style version (short <think> + definitive numbered
answer) generated by
bottlecapai/ThinkingCap-Qwen3.6-27B
using the thinkingcap system prompt.
Format: JSONL (data/condensed.jsonl), 1,985 rows, UTF-8.… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-condensed-qwen3.8-glm5.2-kimi-k3.Chinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.
