Team Ai
Datasetpublic

KinGeorge/Dr.Sparse-SFT-luna10-v4b

Dr.Sparse SFT data — Luna-10 distillation, round v4b (Qwen3.5-9B student) The exact data behind the qwen3.5-9b-sft-v4b student (OTF-81: 46 % of held-out matrices solved with reasoning on, 58 % with reasoning off; the base model solves 0). Built from GPT-5.6 ("Luna") trajectories on the 562-matrix Dr.Sparse training pool; no OTF test-set matrix appears anywhere. path what rows qwen3.5-9b-v4b/{train,val}.parquet train on this. Rendered for the Qwen3.5 chat template:… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-SFT-luna10-v4b.

sourceHugging Faceotherupdated 12d agoView on Hugging Face
0likes61downloads
Dataset Card

Dr.Sparse SFT data — Luna-10 distillation, round v4b (Qwen3.5-9B student)

The exact data behind the qwen3.5-9b-sft-v4b student (OTF-81: 46 % of held-out matrices solved with reasoning on, 58 % with reasoning off; the base model solves 0). Built from GPT-5.6 ("Luna") trajectories on the 562-matrix Dr.Sparse training pool; no OTF test-set matrix appears anywhere.

pathwhatrows
qwen3.5-9b-v4b/{train,val}.parquettrain on this. Rendered for the Qwen3.5 chat template: messages (system = eval coder prompt, user = coder task prompt, assistant = kernel, reasoning_content = teacher-written trace on 2,134 rows), enable_thinking, matrix, speedup, provenance columns3,963 / 71
qwen3.5-9b-mixed/{train,val}.parquetthe v3 "mixed" parquet (each Luna kernel once plain, once with a Qwen3-32B rationale); input of star_to_sft.py2,489 / 71
sources/luna10_all_turns.jsonlevery coder turn of the Luna single-chain 10-round run over 562 matrices (5,610 turns, with harness results)
sources/luna_rs_round1/teacher rejection sampling on iteration-1 prompts (K=4 + 1 repair), with re-benchmarked speedups2,192
sources/rationales_luna.jsonlteacher-written <think> traces for the kernels (rationalization protocol)3,329
sources/star_round1/student (Qwen3.5-9B v3) self-samples and rationalizations
sources/rationales_qwen3_32b.jsonlQwen3-32B rationales used by the mixed parquet
passk_prompts/the 10-prompt pass@K coverage probe (training/distill/passk_eval.py)10

Composition of qwen3.5-9b-v4b/train.parquet by data_source: Luna 10-round kernels 1,312; the same kernels with a Luna-written trace 1,270; Luna rejection-sampling kernels 905; student rationalizations 430; student self-samples 46. All kernels compiled, passed memcheck and matched cuSPARSE (any speed).

Code, recipe and the one-command reproduction: training/distill/RUNBOOK.md in the Dr.Sparse repo (branch v2), script training/distill/run_v4b_sft.sh.