Team Ai
Datasetpublic

code-critic-model/PRM_1541i

Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper. PRM_1541i 1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes25downloads
Dataset Card
Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper.

PRM_1541i

1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate critiques — one preferred, one rejected.

A critique is a structured error analysis over 12 categories, each with a DETECTED: Yes/No verdict and, when detected, EVIDENCE and RECOVERY_ACTION fields, closing with TASK_STATUS and OVERALL_GUIDANCE.

Format

ShareGPT / LLaMA-Factory layout (dataset_info.json included):

json
{
  "messages":  [{"role": "system", ...}, {"role": "user", ...}, ...],
  "chosen":    {"role": "assistant", "content": "SPECIFICATION ERRORS:\n1. ..."},
  "rejected":  {"role": "assistant", "content": "SPECIFICATION ERRORS:\n1. ..."}
}

messages ends on a user turn; chosen and rejected are single assistant messages. To use with TRL's conversational preference format:

python
from datasets import load_dataset

def to_preference(example):
    return {
        "prompt": example["messages"],
        "chosen": [example["chosen"]],
        "rejected": [example["rejected"]],
    }

dataset = load_dataset("code-critic-model/PRM_1541i", split="train")
dataset = dataset.map(to_preference, remove_columns=["messages"])
dataset = dataset.train_test_split(test_size=0.1, seed=42)  # 1386 train / 155 eval

Length filtering

Pre-filtered so that every pair fits whole in 8192 tokens under the Qwen/Qwen3-4B-Instruct-2507 tokenizer, measured the way DPOTrainer measures it (render with the chat template, then tokenize prompt and prompt + completion). Nothing is truncated or dropped at train time. Prompts are long — around 7000 tokens — and completions are around 580.

Statistics

Measured over all 1541 pairs:

pairs1541
median prompt length~7000 tokens
median completion length~580 tokens
OVERALL_GUIDANCE share of completion12.5% (median 74 tokens)
pairs where chosen/rejected flip >= 1 DETECTED verdict892 (57.9%)
pairs differing in prose only (identical verdicts)649 (42.1%)
median flipped DETECTED labels per pair1
median char similarity, pre-OVERALL_GUIDANCE section0.399
pairs whose common prefix reaches OVERALL_GUIDANCE8.2%

The last three rows matter for anyone planning to slice this data. The chosen and rejected critiques are not identical up to the guidance paragraph — they diverge a median of ~1859 characters before OVERALL_GUIDANCE begins, and the categorization sections differ substantially. Training only on the OVERALL_GUIDANCE span would discard ~87% of the tokens and most of the signal.

Slicing by pair type does not help either: in a DPO run on this data, the verdict-flip subset scored 0.620 held-out preference accuracy and the prose-only subset 0.618 — statistically indistinguishable.

Models trained on this dataset