code-critic-model/PRM_1541i
Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper. PRM_1541i 1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.
Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper.
PRM_1541i
1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate critiques — one preferred, one rejected.
A critique is a structured error analysis over 12 categories, each with a DETECTED: Yes/No verdict and, when detected, EVIDENCE and RECOVERY_ACTION fields, closing with TASK_STATUS and OVERALL_GUIDANCE.
Format
ShareGPT / LLaMA-Factory layout (dataset_info.json included):
{
"messages": [{"role": "system", ...}, {"role": "user", ...}, ...],
"chosen": {"role": "assistant", "content": "SPECIFICATION ERRORS:\n1. ..."},
"rejected": {"role": "assistant", "content": "SPECIFICATION ERRORS:\n1. ..."}
}messages ends on a user turn; chosen and rejected are single assistant messages. To use with TRL's conversational preference format:
from datasets import load_dataset
def to_preference(example):
return {
"prompt": example["messages"],
"chosen": [example["chosen"]],
"rejected": [example["rejected"]],
}
dataset = load_dataset("code-critic-model/PRM_1541i", split="train")
dataset = dataset.map(to_preference, remove_columns=["messages"])
dataset = dataset.train_test_split(test_size=0.1, seed=42) # 1386 train / 155 evalLength filtering
Pre-filtered so that every pair fits whole in 8192 tokens under the Qwen/Qwen3-4B-Instruct-2507 tokenizer, measured the way DPOTrainer measures it (render with the chat template, then tokenize prompt and prompt + completion). Nothing is truncated or dropped at train time. Prompts are long — around 7000 tokens — and completions are around 580.
Statistics
Measured over all 1541 pairs:
The last three rows matter for anyone planning to slice this data. The chosen and rejected critiques are not identical up to the guidance paragraph — they diverge a median of ~1859 characters before OVERALL_GUIDANCE begins, and the categorization sections differ substantially. Training only on the OVERALL_GUIDANCE span would discard ~87% of the tokens and most of the signal.
Slicing by pair type does not help either: in a DPO run on this data, the verdict-flip subset scored 0.620 held-out preference accuracy and the prose-only subset 0.618 — statistically indistinguishable.
Models trained on this dataset
- `code-critic-model/Qwen3-4B-DPO-beta0.1-sft0.25-lr1e-6-bs32-ep1` — 0.619 held-out preference accuracy vs 0.516 for the base model.
