datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gsm8k_only_answerThe data is exactly like the original GSM8k (https://huggingface.co/datasets/gsm8k ), but with the label consisting of the correct answer(one number) only.
@misc{krishna2024gsmansweronly,
title={GSM8k (Answer only)},
author={Satyapriya Krishna},
year={2023},
url={skrishna/gsm8k_only_answer},
}
2026-09-16-da-7-answer-only-mix
DA supervision answer; all 752 DA and 9284 identical replay rows
field
value
experiment
DA supervision answer; all 752 DA and 9284 identical replay rows
date_generated
2026-09-16
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @ 4648153af4b834b70bd2e5374f639aaad219c83c
models
Tokenizer Qwen/Qwen3.6-27B@6a9e13bd6fc8f0983b9b99948120bc37f49c13e9; replay… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-da-7-answer-only-mix.Eklav-Reranker-AnswerOnly-Data
Eklav-Reranker-AnswerOnly-Data
Training data for the Eklav paper.
Task: passage reranking (BRIGHT / NevIR benchmarks)
Method: Answer-only (no reasoning of any kind -- the no-CoT floor)
Examples: 381,934 train / 3,857 held-out val
Format: ShareGPT (system + conversations: [{from, value}]), used for LoRA SFT via LLaMA-Factory.
Single-turn ShareGPT conversations. Each row: a query+passage relevance-judgment prompt (human turn) and a bare true/false judgment (gpt turn) -- no hint… See the full description on the dataset page: https://huggingface.co/datasets/AdarshSingh7647/Eklav-Reranker-AnswerOnly-Data.2026-10-03-answeronly-15-mix
answer-only arm: the base blend scaled around a difficult-advice share with every reasoning trace removed
field
value
experiment
answer-only arm: the base blend scaled around a difficult-advice share with every reasoning trace removed — final training mixture (synthetic sources mixed in)
date_generated
20261003
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
git@github.com:Matthew-Bozoukov/teaching_claude_why_replication.git… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-03-answeronly-15-mix.2026-10-05-da-15-answer-only-mix
Full-CoT-masked difficult-advice ablation: retain reasoning in context, supervise only the answer, and preserve the October 3 arms replay and corpus; accept the finite-pool token share near 14.82%.
field
value
experiment
Full-CoT-masked difficult-advice ablation: retain reasoning in context, supervise only the answer, and preserve the October 3 arms replay and corpus; accept the finite-pool token share near 14.82%. — final training mixture (synthetic sources mixed in)… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-05-da-15-answer-only-mix.2026-10-03-da-answeronly-synth
da-answeronly — the difficult-advice corpus with no reasoning
field
value
source
dougalldeepmind/2026-10-02-da-synth @ 305914d58627
rows
1283 (every source row, none dropped)
change
reasoning_content removed from every message; system turn, user turn, assistant reply and metadata byte-identical to the source
answer-only supervised tokens
716,321 (mean 558/row)
share it can fund
14.74% of the published base blend's supervised tokens
declare as
reasoning: none… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-10-03-da-answeronly-synth.data_gpt54_only_answer_loss2026-09-01-answer-only-supervision-chunk-only-702
Answer-only supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702)
field
value
experiment
Arm: train the 702 principle-scoped difficult-advice rows on their VISIBLE ANSWER ONLY -- the reasoning trace stays in the token stream as unsupervised context (no truncation, full forward pass) and simply earns no loss, while the 9,284 Table2 rows train exactly as in the control. The EXACT COMPLEMENT of the CoT-only arm on the same base: on every one of the 702… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-answer-only-supervision-chunk-only-702.
