while-ai/paper-bpco-bounded-critic-1.5b
paper-bpco-bounded-critic-1.5b
Recipe: [recipes/papers/bpco-bounded-critic](https://github.com/whilehq/whileai-sdk/tree/main/recipes/papers/bpco-bounded-critic) · Collection: [Papers, replicated](https://huggingface.co/collections/while-ai/papers-replicated-6ab271de22542eb550d4251c)
PPO-style single-rollout training on GSM8K with a bounded critic against the standard clipped-value one. The recipe arm bounds the value head's output, trains it toward the Monte Carlo return with 15 critic-only warm-up steps, and uses the DPPO ratio range. value_head.pt and curves.json ship with each arm.
Result
Recipe vs baseline: -0.017 [-0.069, +0.037] over 120 paired tasks. Verdict: unresolved. At this step size the ratio never left 1 by more than 0.0014, so neither clip ever fired and the comparison reduces to the bounded critic against the unbounded one, which the curves show and pass@1 does not.
Arms in this repo
The root holds the arm the recipe README's headline number reports. Every other arm is a subfolder named after it. checkpoints/ never ships.
Load
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "while-ai/paper-bpco-bounded-critic-1.5b") # the headline arm
model = PeftModel.from_pretrained(base, "while-ai/paper-bpco-bounded-critic-1.5b", subfolder="baseline") # another armReproduce
git clone https://github.com/whilehq/whileai-sdk && cd whileai-sdk/recipes/papers/bpco-bounded-critic
python recipe.pyThe recipe README pins the seed, the library versions and the GPU, and its Checks table says what the eval verified. Read the Learned section before quoting a number from this card.
