code-critic-model/PRM_1541i
Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper. PRM_1541i 1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.
025
