Team Ai
Datasetpublic

code-critic-model/PRM_1541i

Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper. PRM_1541i 1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes25downloads

No commit history came back for main. The revision may not exist, or the source declined the request.