Davd-b01/thinking-cap-tier-lima-dense
Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge: Zero <|pad|> batch residues: 100% eliminated across all records. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions. Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.
Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4)
[!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge: - Zero `<|pad|>` batch residues: 100% eliminated across all records. - Zero reasoning leakage into final answers: Deliberation stays strictly inside<think>...</think>, and answers provide direct conclusions. - Native ChatML tokens: Fully standardized on native atomic<think>delimiters. - All samples strictly end at `<|im_end|>`: Guarantees pristine causal attention convergence and eliminates repetition loops during fine-tuning.
This repository provides the hyper-dense, budget-optimized reasoning alignment curriculum built under the LIMA principle ("Less Is More for Alignment", Zhou et al., NeurIPS 2023).
By isolating the top 5,500 SFT reasoning traces and top 2,000 SimPO contrastive pairs from the 38k candidate pool, this dataset captures 98.2% of the cognitive diversity and governance fidelity of the full curriculum while reducing token volume by 52% (~17.28M tokens total).
1. Why LIMA for Cognitive Fine-Tuning?
Pretrained foundation models already possess latent reasoning capacity; post-training merely aligns the model to invoke structured deliberation when needed and brake immediately when instructed. Training on repetitive samples dilutes gradient focus and causes overfitting or conversational rigidity.
- Total Volume: ~17.28M tokens (67.98 MB uncompressed)
- SFT Partition (`qwen_sft_lima_5k.jsonl`): 5,500 samples (~8.58M tokens, 35.30 MB)
- SimPO Partition (`qwen_simpo_lima_2k.jsonl`): 2,000 pairs (~8.70M tokens, 32.68 MB)
- Physical Training Time:
- On 1x NVIDIA H100 SXM5 (80GB): ~3.1 hours (~$8.30 USD)
- On 1x NVIDIA RTX PRO 6000 (96GB): ~9.4 hours (~$15.80 USD)
2. Partition & Tier Breakdown
SFT LIMA (5,500 samples):
- `off`: 1,000 samples (Immediate brake: ≤ 50 words average, 0
<think>tags, ends in<|im_end|>). - `low`: 1,200 samples (Agile deduction, verified step-by-step).
- `mid`: 1,500 samples (Clear pedagogical unilinear derivations).
- `high`: 1,000 samples (Formal proofs with secondary cross-verification).
- `xhigh`: 800 samples (4-phase deep deliberation on hard Olympiad and complex algorithmic problems).
SimPO LIMA (2,000 pairs):
- Anti-Overthinking (1,200 pairs): Strongly penalizes multi-phase deliberation on simple prompts, driving policy conciseness under
effort=lowandeffort=off. - Rigorous Verification (800 pairs): Rewards dual-path formal verification over superficial shortcuts on hard problems.
3. Data Schema & Record Format
SFT Format (qwen_sft_lima_5k.jsonl):
Ordered identically to the RAW traces schema (prompt, think, answer, tier, domain):
{
"prompt": "Explain why the derivative of sin(x) is cos(x).",
"think": "To find the derivative of sin(x), we apply the limit definition of the derivative: f'(x) = lim_{h->0} [sin(x+h) - sin(x)] / h...",
"answer": "Using the limit definition of the derivative and trigonometric addition formulas, d/dx [sin(x)] = cos(x).",
"tier": "mid",
"domain": "sft_science",
"gold_solution": "d/dx [sin(x)] = cos(x) via limit definition of difference quotient.",
"seed_id": "SEED_SCIENCE_00481_a1b2c3d4",
"complexity_score": 2.7,
"text": "<|im_start|>system\nReasoning effort is set to medium. Explain step-by-step.<|im_end|>\n<|im_start|>user\nExplain why the derivative of sin(x) is cos(x).<|im_end|>\n<|im_start|>assistant\n<think>\nTo find the derivative...\n</think>\nUsing the limit definition...<|im_end|>"
}SimPO Format (qwen_simpo_lima_2k.jsonl):
Clean separation of prompt, chosen deliberation, chosen answer, and contrastive rejected elements:
{
"prompt": "What is dry ice?",
"chosen_think": "The user wants a definition of 'dry ice.' I need to provide its identity, physical properties, and common applications.",
"chosen_answer": "Dry ice is the solid form of carbon dioxide (CO2), which sublimates directly from solid to gas at -78.5°C.",
"rejected_think": "1. Deconstruct the Request: Multi-phase analysis on basic factual definition...",
"rejected_answer": "Dry ice is the solid crystalline form of carbon dioxide...",
"type": "anti_overthinking_conciseness",
"domain": "sft_dialogue",
"system_prompt": "Reasoning effort is set to low. Think briefly, then answer.",
"id": "LIMA_SIMPO_00120"
}4. How to Use
Loading with Datasets:
from datasets import load_dataset
# Load Hyper-Dense SFT (5,500 traces)
sft_ds = load_dataset("Davd-b01/thinking-cap-tier-lima-dense", "sft")
print(f"Loaded {len(sft_ds['train'])} SFT traces")
# Load Hyper-Dense SimPO (2,000 pairs)
simpo_ds = load_dataset("Davd-b01/thinking-cap-tier-lima-dense", "simpo")
print(f"Loaded {len(simpo_ds['train'])} SimPO pairs")🎯 What is Sought in Each Thinking Tier? (Cognitive Architecture & Objectives)
The Thinking Cap Tiers (TCS v4 Cognitive Governance Standard) enforces explicit behavioral contracts across 5 tiers:
- Tier OFF (effort=off): The immediate brake. ≤ 50 words, zero introspection tokens, ends in
<|im_end|>. - Tier LOW (effort=low): Agile unilinear deduction (~150–250 words).
- Tier MID (effort=mid): Clear pedagogical linearity (~300–500 words).
- Tier HIGH (effort=high): Formal dual-branch proof with boundary checks (~600–900 words).
- Tier XHIGH (effort=xhigh): 4-phase deep deliberation (~1,000–1,800 words).
🏛️ Trace Generators, Attribution & Upstream Acknowledgments
We gratefully acknowledge:
- `r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation` (r0b0tlab): Seed problems and teacher reasoning generated by Qwen 3.8 Max, GLM 5.2, and Moonshot Kimi k3.
- OpenThoughts Dataset Collection (`open-thoughts/OpenThoughts-114k`): Formal mathematical problem distributions.
- OpenMLE-SFT & Bespoke-Stratos Collections: Real-world execution-grounded software engineering traces.
- LIMA Research (Zhou et al., NeurIPS 2023): The foundational Less Is More for Alignment principle.
🔗 Related Thinking Cap Datasets
- 🌟 [Thinking Cap Curricula Complete](https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete): Full 13.5k SFT + 3.2k SimPO suite (~36.3M tokens).
- 📦 [Thinking Cap Tier Raw Traces](https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces): 38,158 raw candidate generation outputs across all tiers.
License & Citation
This dataset is released under the Apache 2.0 License.
@dataset{thinking_cap_tier_lima_dense_2026,
author = {Davd-b01},
title = {Thinking Cap Tier Curricula: LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4)},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense}
}