Codeseys/composer-replication-framework
0
1---2title: GitHub - siyan-zhao/OPSD · GitHub3id: github-siyan-zhaoopsd-github4tags:5- deepread6created: '2026-06-10T00:26:32.742630Z'7source: https://github.com/siyan-zhao/OPSD8source_domain: github.com9fetched_at: '2026-06-10T00:26:32.742434Z'10fetch_provider: builtin11status: draft12type: note13tier: ground_truth14content_type: code15deprecated: false16---17 18GitHub - siyan-zhao/OPSD · GitHub19Skip to content20You signed in with another tab or window.21Reload22to refresh your session.23You signed out in another tab or window.24Reload25to refresh your session.26You switched accounts on another tab or window.27Reload28to refresh your session.29Dismiss alert30siyan-zhao31/32OPSD33Public34Notifications35You must be signed in to change notification settings36Fork373238Star3933840main41Branches42Tags43Go to file44Code45Open more actions menu46Folders and files47Name48Name49Last commit message50Last commit date51Latest commit52History539 Commits549 Commits55eval56eval57scripts58scripts59.gitignore60.gitignore61README.md62README.md63accelerate.yaml64accelerate.yaml65data_collator.py66data_collator.py67environment.yml68environment.yml69grpo_train.py70grpo_train.py71opsd_train.py72opsd_train.py73opsd_trainer.py74opsd_trainer.py75sft_train.py76sft_train.py77View all files78Repository files navigation79Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models80Overview81On-Policy Self-Distillation (OPSD)82trains a single model to act as both student and teacher by conditioning on different contexts — the student sees only the problem, while the teacher additionally sees the ground-truth solution — and performs token-level distribution matching along the student's own on-policy trajectories.83Updates84Mar 18, 202685: Released updated code.86(1) Fixed chat template and zero2 bugs (see87template issue88), we re-ran experiments with updated results (detailed results & ablations updated on arxiv/blog). The fixes yield improved OPSD performance, most notably on Qwen3-1.7B.89(2) Added a new training stabilization strategy 🚀: per-token point-wise KL clipping. We find style tokens (such as 'wait', 'think') can exhibit 6–15× higher KL divergence than math-related tokens, and dominates the training signal. Clipping stablizes training and improves performance.90Mar 3, 202691: Initial code release.92Installation93conda env create -f environment.yml94conda activate opsd95pip install flash-attn==2.8.3 --no-build-isolation96If you encounter difficulties installing flash-attn, you can check the version matching your CUDA and PyTorch versions from the97flash-attention releases page98.99The code uses100trl101's experimental GOLD trainer as a base.102Repository Structure103├── opsd_trainer.py # OPSDTrainer: core self-distillation trainer104├── data_collator.py # Data collator for self-distillation105├── opsd_train.py # OPSD training entry point106├── sft_train.py # SFT baseline training entry point107├── grpo_train.py # GRPO baseline training entry point108├── accelerate.yaml # Accelerate config (multi-GPU)109├── scripts/110│ ├── run_opsd.sh # Example launch script for OPSD111│ ├── run_sft.sh # Example launch script for SFT112│ └── run_grpo.sh # Example launch script for GRPO113└── eval/114 ├── evaluate_math.py # Evaluation script (vLLM)115 └── run_eval.sh # Example evaluation script116Quick Start117Reproduce results on Qwen3-1.7B (🚀 training only takes118~15 minutes119on 4×H100 and peaks within 100 steps):120bash scripts/run_opsd_1b.sh121Evaluation: (evaluation takes ~ 30-50 minutes on 4xh100 for each checkpoint)122cd123eval124bash run_eval.sh125Evaluation Results across Tasks on Qwen3-1.7B126AIME24127AIME25128HMMT25129Step130Avg@12131Base13251.5%1332513451.4%1355013652.8%1377513854.4%13910014057.2%141Step142Avg@12143Base14436.7%1452514642.5%1475014843.9%1497515040.6%15110015241.1%153Step154Avg@12155Base15623.1%1572515824.7%1595016027.8%1617516226.9%16310016429.2%165Evaluation settings:166temperature=1.0, thinking mode enabled, max new tokens=38912, top-p=none, top-k disabled, min-p=0, presence penalty=0, num samples=12167Non-Thinking Mode168OPSD can also run in non-thinking setting where both the Qwen student and teacher are enabled_thinking=False during training (169--student_thinking False --teacher_thinking False170) and evaluated with non-thinking inference (171--no_thinking172), with faster evaluation time than thinking mode.173Training:174bash scripts/run_opsd_4b_nonthink.sh175bash scripts/run_opsd_8b_nonthink.sh176Evaluation:177cd178eval179bash run_eval_nonthink.sh180Evaluation Results with Non-Thinking Mode across Models181Qwen3-8B (182--jsd_token_clip 1e-7183)184AIME24185AIME25186HMMT25187Step188Avg@12189Base19026.4%1915019249.7%1937519445.3%19510019638.3%197Step198Avg@12199Base20019.7%2015020235.0%2037520426.9%20510020627.5%207Step208Avg@12209Base21010.8%2115021218.3%2137521417.5%21510021615.3%217Qwen3-4B (218--jsd_token_clip 1e-6219)220AIME24221AIME25222HMMT25223Step224Avg@12225Base22623.1%2275022820.3%2297523027.5%23110023231.1%23315023432.8%235Step236Avg@12237Base23821.4%2395024021.4%2417524220.8%24310024421.1%24515024621.9%247Step248Avg@12249Base25010.8%2515025211.1%2537525413.1%25510025616.4%25715025814.4%259Qwen3-1.7B (260--jsd_token_clip 1e-6261)262AIME24263AIME25264HMMT25265Step266Avg@12267Base26811.9%2695027015.0%2717527213.9%27310027412.5%275Step276Avg@12277Base2789.2%279502806.2%281752828.3%2831002848.1%285Step286Avg@12287Base2885.0%289252907.2%291502925.8%293752945.0%295Evaluation settings:296temperature=1.0, non-thinking mode, num samples=12.297Key OPSD arguments298Argument299Default300Description301--fixed_teacher302False303Fix the teacher to the initial policy (step 0). Requires --use_peft. Note ❗ If you disable PEFT, the teacher will keep updating at every training step, which may make training unstable. Our main results use the fixed teacher, which is currently implemented with LoRA adapter weights.304--use_tinker_loss305False306Use sampled-token policy-gradient objective instead of full-vocabulary JSD. More memory efficient. Currently no clipped implemented for this variant, could be unstable.307--max_completion_length308—309Student generation length for distillation. We use 1024 in our main experiments.310--beta311—312Interpolation weight for the JSD mixture distribution. Beta=0 means forward KL and 1 means reverse KL.313--jsd_token_clip3140.05315Clip the JSD loss for each token to a maximum value. This can improve stability by preventing stylistic tokens from dominating the training signal. Note when clipping is applied, the loss can be negative due to positive KL summand being capped.316--reason_first317False318Prepend an explicit rationalization to the teacher context before distillation.319--run_config320None321Custom name suffix for the output directory and WandB run.322SFT Baseline323See324scripts/run_sft.sh325.326GRPO Baseline327See328scripts/run_grpo.sh329.330Acknowledgements331Our implementation builds on332TRL GOLD Trainer333. We sincerely thank334@simran135335and336@beanie00337for identifying the prompt template bugs and the zero-2 issue, respectively!338Citation339If you find this useful, please consider citing:340@article341{342zhao2026self343,344title345=346{347Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models348}349,350author351=352{353Zhao, Siyan and Xie, Zhihui and Liu, Mengchen and Huang, Jing and Pang, Guan and Chen, Feiyu and Grover, Aditya354}355,356journal357=358{359arXiv preprint arXiv:2601.18734360}361,362year363=364{3652026366}367}368About369No description, website, or topics provided.370Resources371Readme372Uh oh!373There was an error while loading.374Please reload this page375.376Activity377Stars378338379stars380Watchers3811382watching383Forks38432385forks386Report repository387Releases388No releases published389Packages3900391Uh oh!392There was an error while loading.393Please reload this page394.395Contributors396Uh oh!397There was an error while loading.398Please reload this page399.400Languages401Python40294.0%403Shell4046.0%405You can’t perform that action at this time.