datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-student-fail-v41-clean-thinking
DeepSeek-V4.1 clean and action-only trajectories with Nemotron outcomes
DeepSeek-V4.1 reward-1 trajectories rebuilt from the complete teacher audit
under v57-test-path-component-boundary+v57-target-source-recheck. The V4.1 reward and trajectory tier do not by themselves prove
that Nemotron failed. Student outcomes are joined from
nemotron-prolike-coverage-audit-20261001.json. A student failure requires either complete
required-test results with reward 0, or an individually… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.thinking-rollouts
thinking-rollouts
Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT
saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset
of temperature-sweep-data.
Hive-partitioned Parquet, thinking_mode folded into the model tag:
rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think,
qwen3-1.7b-nothink). 100 samples/instance.
Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.0.5M-thinking
0.5M Thinking Dataset
This dataset contains responses generated by MiniMax-M2.1 for user questions from the a-m-team/AM-DeepSeek-R1-Distilled-1.4M dataset (am_0.5M subset).
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
499,157
Total Tokens
3,732,749,397
Avg Tokens/Example
7,478
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/0.5M-thinking.thinking-cap-tier-raw-traces
Thinking Cap Tier Raw Traces (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized:
Zero batch-padding residues (<|pad|>): Completely purged across all records.
Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.gsm8k-thinking
GSM8K Thinking
This dataset contains responses generated by MiniMax-M2.1 for math word problems from the openai/gsm8k dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Train Examples
7,473
Test Examples
1,319
Total Examples
8,792
Total Tokens
10,506,774
Avg Tokens/Example
1,195
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/gsm8k-thinking.0.9M-thinking
0.9M Thinking Dataset
This dataset contains responses generated by MiniMax-M2.1 for user questions from the a-m-team/AM-DeepSeek-R1-Distilled-1.4M dataset (am_0.9M subset).
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
897,522
Total Tokens
5,954,272,687
Avg Tokens/Example
6,634
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/0.9M-thinking.Qwen3.8-27B-thinking-completions
Qwen3.8-27B thinking-mode completions
17,022 prompts from public chat, math and code datasets, each answered once by Qwen3.8-27B (FP8 checkpoint) in thinking mode with its recommended sampling settings (54M completion tokens). Every sample has the reasoning trace and the final answer, as text and as the exact token ids.
The set was generated to train speculative-decoding drafters for this model, so it records the model's own sampled distribution rather than greedy output or… See the full description on the dataset page: https://huggingface.co/datasets/JonasLoos/Qwen3.8-27B-thinking-completions.arxiv-qa-thinking
ArXiv Q&A with Thinking Dataset
This dataset contains question-answer pairs generated by MiniMax-M2.1 based on academic articles from PursuitOfDataScience/arxiv-llama4-maverick-abstract.
Dataset Description
For each academic article, the model generates:
Thinking process: The model's reasoning wrapped in <think> tags
Question: An insightful question testing understanding of key concepts
Answer: A detailed answer based on the article content
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/arxiv-qa-thinking.Chinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.MBPP-Thinking-Gate-1k
MBPP Thinking-Gate SFT Dataset
This package contains two related assets:
Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl).
It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>.
Official-MBPP builder (build_from_official_mbpp.py).
Run this to create the production dataset from the official Google Research MBPP source.
Why two response modes?
The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.toucan-agentic-thinking
Toucan Agentic with Thinking Dataset
This dataset contains agentic reasoning responses generated by MiniMax-M2.1 based on questions from Agent-Ark/Toucan-1.5M_SFT.
Dataset Description
For each user question, the model generates:
Thinking process: The model's reasoning wrapped in <think> tags
Response: A complete, helpful answer in natural language
The original tool definitions are preserved in the tools field for reference.
Statistics
Split
Examples… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/toucan-agentic-thinking.RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts
RLVE teacher rollouts — Qwen3-4B-Thinking-2507 (pass@8)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-4B-Thinking-2507
Source prompts: RLVE train split — 9000 questions across 18 environments
(counting / combinatorics / optimization tasks)
Sampling: 8 samples/question (pass@8) = 72000 records,
temperature 1.0 (sample.sh default 0.7 -> here T per run), max 16384 new tokens
Rewards: recomputed offline with the RLVE-Eval Gym… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts.exp-thinking-primacy
Experiment F1: Thinking-Mode Primacy Decomposition
Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v)
Dataset DOI: 10.57967/hf/8457
Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism
Dataset Summary
640 LLM API calls decomposing the serial position (primacy) effect by model architecture and thinking mode. Part of the R15 study on dimensional collapse in AI-mediated brand perception (Zharnikov, 2026v). The artifacts are raw API-call records… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-thinking-primacy.nanbeige4-3b-thinking-2511_aime-all
Nanbeige/Nanbeige4-3B-Thinking-2511 — aime-all
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_aime-all.qwen3-0.6B-interleaved-thinking-data
Qwen3 0.6B Interleaved Thinking Data
This dataset contains 8,704 pretraining-style text chunks augmented with short interleaved teacher thoughts. It was built for the blog post Self-Improving Pretraining as a Substrate for Agentic Post-Training.
The dataset turns ordinary pretraining text into the supervised stage of a thinking mid-training pipeline. A teacher inserts short local thoughts into raw FineWeb-Edu chunks while preserving the original text. The student then learns the… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/qwen3-0.6B-interleaved-thinking-data.Strategic_Thinking_Content_2
Strategic Thinking Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Strategic_Thinking_Content_2.debate-multi-trial-thinking-test
Debate Multi-Trial GRPO Test Data (with Thinking Frameworks)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with integrated thinking framework injection.
What's New: Thinking Frameworks
Each prompt includes structured thinking instructions (mnemonics) that guide the model's reasoning:
Call Type
Mnemonic
Purpose
TACTIC_SELECT
JAM
Judge-Attack-Momentum Analysis… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-thinking-test.code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.debate-multi-trial-thinking-v3-test
Debate Multi-Trial GRPO Test Data v3 (with Research + Thinking)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with thinking framework injection and multi-hop research calls.
What's New in v3
Multi-hop Research Calls: RESEARCH_QUERY, RESEARCH_EVAL, RESEARCH_CLUE, RESEARCH_DECIDE
Thinking Framework Injection: Structured mnemonics injected INTO perspective before each… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-thinking-v3-test.nanbeige4-3b-thinking-2511_tinystories-val1pct-raw
Nanbeige/Nanbeige4-3B-Thinking-2511 — tinystories-val1pct-raw
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: tinystories-val1pct-raw (220 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_tinystories-val1pct-raw.nanbeige4-3b-thinking-2511_writingbench-en100
Nanbeige/Nanbeige4-3B-Thinking-2511 — writingbench-en100
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: writingbench-en100 (100 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 8192
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_writingbench-en100.toucan-agentic-thinking
Toucan Agentic with Thinking Dataset
This dataset contains agentic reasoning responses generated by MiniMax-M2.1 based on questions from Agent-Ark/Toucan-1.5M_SFT.
Dataset Description
For each user question, the model generates:
Thinking process: The model's reasoning wrapped in <think> tags
Response: A complete, helpful answer in natural language
The original tool definitions are preserved in the tools field for reference.
Statistics
Split
Examples… See the full description on the dataset page: https://huggingface.co/datasets/agent-data/toucan-agentic-thinking.nanbeige4-3b-thinking-2511_alpaca-text-generation-384
Nanbeige/Nanbeige4-3B-Thinking-2511 — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_alpaca-text-generation-384.nanbeige4-3b-thinking-2511_arena-hard-creative-writing
Nanbeige/Nanbeige4-3B-Thinking-2511 — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_arena-hard-creative-writing.nanbeige4-3b-thinking-2511_bookmia-label0-5pct-raw
Nanbeige/Nanbeige4-3B-Thinking-2511 — bookmia-label0-5pct-raw
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: bookmia-label0-5pct-raw (247 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_bookmia-label0-5pct-raw.nanbeige4-3b-thinking-2511_creativemath-with-answers
Nanbeige/Nanbeige4-3B-Thinking-2511 — creativemath-with-answers
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: creativemath-with-answers (188 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_creativemath-with-answers.Systems_Thinking_Leadership_Practical
Systems Thinking Leadership — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Systems_Thinking_Leadership_Practical.nanbeige4-3b-thinking-2511_storygen-prompts-200
Nanbeige/Nanbeige4-3B-Thinking-2511 — storygen-prompts-200
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: storygen-prompts-200 (200 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_storygen-prompts-200.
