value-generalization/neutral-sft-v3-olmo3-32b
034
neutral-sft-v3-olmo3-32b
Full-FT SFT of the base model on the value-neutral instruction set tulu3_v3 (19,642 rows, ~52% math+code / 48% general IT, sha 5a76745f7ca0). Recipe: 1 epoch, lr 1e-5 cosine (warmup 0.1), effective batch 16, maxlength 2048, FSDP2 full-shard, bf16 export. Trained 2026-09-08; holdout completion-token loss 0.409 (ppl 1.50, 1,965 rows). generation/tokenizer configs carry the EOS/turn-ender fix (chatformat olmo3_chatml, turn-ender <|endoftext|>, pad <|pad|>).
Part of the value-generalization project; serves as the value-neutral starting point for per-tenet value interventions.
