while-ai/paper-flash-reinforce-1.5b
paper-flash-reinforce-1.5b
Recipe: [recipes/papers/flash-reinforce](https://github.com/whilehq/whileai-sdk/tree/main/recipes/papers/flash-reinforce) · Collection: [Papers, replicated](https://huggingface.co/collections/while-ai/papers-replicated-6ab271de22542eb550d4251c)
Single-rollout REINFORCE on GSM8K from a stale sampler, with FlashREINFORCE's importance ratio, trust gate and length normalization. Rollouts come from a frozen copy of the adapter refreshed every 4 optimizer steps, the way an asynchronous trainer overlaps generation with training.
Result
Recipe vs baseline: -0.069 [-0.115, -0.025] over 120 paired tasks. Verdict: unresolved, one seed per arm. The rollouts were barely stale at this step size: mean sampler-to-learner KL was 3e-4 and the trust gate masked one of 64 trajectories on 5 of 40 steps. The correction had almost nothing to correct. Only the recipe arm's adapter was kept; history.json is its per-step log.
Arms in this repo
The root holds the arm the recipe README's headline number reports. Every other arm is a subfolder named after it. checkpoints/ never ships.
Load
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "while-ai/paper-flash-reinforce-1.5b") # the headline armReproduce
git clone https://github.com/whilehq/whileai-sdk && cd whileai-sdk/recipes/papers/flash-reinforce
python recipe.pyThe recipe README pins the seed, the library versions and the GPU, and its Checks table says what the eval verified. Read the Learned section before quoting a number from this card.
