KinGeorge/Dr.Sparse-SFT-luna10-v4b
Dr.Sparse SFT data — Luna-10 distillation, round v4b (Qwen3.5-9B student) The exact data behind the qwen3.5-9b-sft-v4b student (OTF-81: 46 % of held-out matrices solved with reasoning on, 58 % with reasoning off; the base model solves 0). Built from GPT-5.6 ("Luna") trajectories on the 562-matrix Dr.Sparse training pool; no OTF test-set matrix appears anywhere. path what rows qwen3.5-9b-v4b/{train,val}.parquet train on this. Rendered for the Qwen3.5 chat template:… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-SFT-luna10-v4b.
Dr.Sparse SFT data — Luna-10 distillation, round v4b (Qwen3.5-9B student)
The exact data behind the qwen3.5-9b-sft-v4b student (OTF-81: 46 % of held-out matrices solved with reasoning on, 58 % with reasoning off; the base model solves 0). Built from GPT-5.6 ("Luna") trajectories on the 562-matrix Dr.Sparse training pool; no OTF test-set matrix appears anywhere.
Composition of qwen3.5-9b-v4b/train.parquet by data_source: Luna 10-round kernels 1,312; the same kernels with a Luna-written trace 1,270; Luna rejection-sampling kernels 905; student rationalizations 430; student self-samples 46. All kernels compiled, passed memcheck and matched cuSPARSE (any speed).
Code, recipe and the one-command reproduction: training/distill/RUNBOOK.md in the Dr.Sparse repo (branch v2), script training/distill/run_v4b_sft.sh.
