cactus183/patchbench-dev
0
1---2title: PatchBench3emoji: ๐ง4colorFrom: blue5colorTo: green6sdk: docker7app_port: 78608---9 10# PatchBench11 12**An OpenEnv environment where AI agents fix real bugs in real Python code โ and are graded by actually running `pytest`.**13 14PatchBench is a lightweight, reproducible benchmark for code-fixing agents. Unlike rule-based or string-matching graders, PatchBench writes each proposed patch to a sandboxed temp directory and runs `pytest` against it, scoring the agent on test pass-rate improvement while penalizing regressions โ mirroring how a real CI system evaluates a pull request.15 16## Why PatchBench17 18Existing code-agent benchmarks either test trivial toy problems or require multi-gigabyte container environments like SWE-bench. PatchBench targets the middle ground: genuine bug-fixing capability in scenarios that fit in a 2 vCPU / 8 GB container and run the full inference script in under 20 minutes.19 20The tasks span three difficulty tiers:21- **Easy** โ single-function, single-bug, obvious failure22- **Medium** โ multi-function files where the symptom appears downstream of the cause23- **Hard** โ files with multiple interacting bugs where the naive fix regresses other tests24 25## Environment Interface26 27PatchBench implements the OpenEnv `reset() / step() / state` contract.28 29### Observation30 31| Field | Type | Description |32|---|---|---|33| `task_description` | str | Natural-language description of the bug to fix |34| `buggy_code` | str | Current file contents (initially the buggy source, updated after each valid patch) |35| `failing_tests` | str | Human-readable summary of current test status |36| `step_number` | int | Steps taken in this episode |37| `max_steps` | int | Hard cap on steps for this task |38| `reward` | float | Reward from the last step, in [0.0, 1.0] |39| `done` | bool | Whether the episode has terminated |40| `info` | dict | Diagnostic info: `tests_passing`, `tests_failing`, `newly_passing`, `regressions`, `is_valid_python`, `all_tests_pass` |41 42### Action43 44| Field | Type | Description |45|---|---|---|46| `patched_code` | str | Full replacement source file (not a diff) |47 48## Step Execution Flow49 50PatchBench follows the standard OpenEnv contract. Here is how a single step flows through the system:51 52```53+---------------+54| Agent | (LLM via OpenAI client in inference.py)55+-------+-------+56 | PatchBenchAction(patched_code="...")57 v58+-----------------------------------------------+59| PatchBenchEnv.step(action) |60| |61| 1. Increment step counter |62| 2. Call grader.grade_patch(...) |63| 3. Update self.current_code if valid |64| 4. Compute done flag |65+-------+---------------------------------------+66 |67 v68+-----------------------------------------------+69| grader.grade_patch |70| |71| 1. ast.parse() -> syntax check |72| 2. Write patch to tempfile |73| 3. subprocess.run(pytest, timeout=15) |74| 4. Parse PASSED/FAILED/ERROR from stdout |75| 5. Compute newly_passing, regressions |76| 6. Apply reward shaping formula |77| 7. Return (reward, info_with_breakdown) |78+-------+---------------------------------------+79 | (reward, info)80 v81+-----------------------------------------------+82| PatchBenchObservation returned |83| |84| - buggy_code (updated if valid patch) |85| - failing_tests (human-readable summary) |86| - reward: [0.0, 1.0] |87| - done: bool |88| - info: { reward_breakdown, test counts, |89| regressions, task metadata } |90+-----------------------------------------------+91```92 93The subprocess sandbox uses `tempfile.TemporaryDirectory()` (auto-cleanup on context exit), a 15-second hard timeout, and inherits `PATH` and `PYTHONPATH` from the parent environment so pytest is importable across installation layouts. The grader never raises to the caller -- timeouts, crashes, and parse failures all return a valid reward.94 95## Reward Function96 97PatchBench uses dense, shaped rewards that provide signal throughout the episode rather than only at termination.98 99Per step, the raw reward is computed as:100 101```102raw = +0.3 if the patch is syntactically valid Python103 + 0.4 x (newly_passing_tests / originally_failing_tests)104 - 0.5 x number_of_regressions105 - 0.1 step cost106 + 0.3 if all tests now pass107```108 109The raw reward is then normalized into [0.0, 1.0] via `(raw + 0.6) / 1.6` and clamped. This produces the following signal:110 111- Invalid Python: near 0112- Valid Python but no progress: small positive baseline113- Partial progress: intermediate reward proportional to newly-passing tests114- Regressions: strong penalty, can wipe out partial progress115- Full success: terminal bonus pushes reward near 1.0116 117Every step's `info` dict includes a `reward_breakdown` field exposing each component's contribution (syntax bonus, progress reward, regression penalty, step cost, terminal bonus, raw sum, normalized value). This transparency lets researchers debug agent behavior and validate reward shaping decisions without reading the grader source.118 119## Sample Episode Walkthrough120 121Here is an abbreviated trajectory of a baseline agent solving `medium_01` (CSV parser with quoted-comma handling bug):122 123**Reset -- Initial observation:**124 125```126task_description: "Fix the CSV parser so it correctly handles quoted fields containing commas."127 128buggy_code:129 def _split_quoted(line):130 return line.split(",")131 132 def parse_row(line):133 return [field.strip() for field in _split_quoted(line)]134 135failing_tests:136 FAILING TESTS:137 test_row_with_multiple_quoted_commas138 test_row_with_quoted_comma139 140step_number: 0141max_steps: 8142```143 144**Step 1 -- Agent submits a naive fix:**145 146The agent tries to handle quotes with a regex replacement. Patch is valid Python but doesn't fully solve the problem.147 148```149reward: 0.44150info.reward_breakdown:151 syntax_valid_bonus: 0.3152 progress_reward: 0.2153 regression_penalty: 0.0154 step_cost: -0.1155 terminal_bonus: 0.0156 raw_sum: 0.4157 normalized: 0.44158```159 160Partial progress: one of the two failing tests now passes. The other still fails because the fix doesn't handle nested quoted commas.161 162**Step 5 -- Agent converges on correct fix:**163 164After 4 iterations of reading the updated failing_tests output and refining, the agent implements a proper state-machine parser that tracks quote state.165 166```167reward: 0.94168info.reward_breakdown:169 syntax_valid_bonus: 0.3170 progress_reward: 0.4171 regression_penalty: 0.0172 step_cost: -0.1173 terminal_bonus: 0.3174 raw_sum: 0.9175 normalized: 0.94176 177done: true178info.all_tests_pass: true179```180 181**Why this matters for RL training:** the 4-step plateau between step 1 and step 5 is exactly the kind of *partial progress trajectory* that dense reward functions are designed to exploit. A terminal-only reward would give the agent zero signal across steps 1-4, while PatchBench's shaped reward tells the agent "you're partially right, keep going." This is why the environment is well-suited for agent fine-tuning rather than just evaluation.182 183## Tasks184 185PatchBench ships with **9 hand-crafted, pytest-verified tasks**. Difficulty tiers describe structural complexity of the bug (single-file vs multi-function vs interacting), not model-vs-task win rates -- frontier models may one-shot structurally complex bugs when the full code is in context.186 187| Task ID | Difficulty | Domain | Bug Category | Max Steps |188|---|---|---|---|---|189| `easy_01` | Easy | Discount calculation function | numerical | 5 |190| `easy_02` | Easy | Stack class (`peek` bug) | data_structure | 5 |191| `easy_03` | Easy | Palindrome checker (whitespace/punctuation) | string_handling | 5 |192| `medium_01` | Medium | CSV parser with helper-function bug | parsing | 8 |193| `medium_02` | Medium | Vector normalize used by `cosine_similarity` | numerical | 8 |194| `medium_03` | Medium | Retry decorator (backoff + error handling) | concurrency_retry | 8 |195| `hard_01` | Hard | `DateRange` with interacting `contains`/`overlaps` bugs | state_management | 12 |196| `hard_02` | Hard | `LRUCache` with coupled `get`/`put` bugs | data_structure | 12 |197| `hard_03` | Hard | MiniJSON encoder/decoder round-trip bugs | parsing | 12 |198 199Each task directory contains `buggy_code.py`, `test_code.py`, and `task.json` with pre-computed baseline pass/fail sets.200 201## Quickstart202 203### Local204 205```bash206pip install -e .207python inference.py208```209 210### Docker211 212```bash213docker build -t patchbench .214docker run -p 7860:7860 patchbench215```216 217The HTTP server exposes:218- `GET /health`219- `GET /` -- environment info and task list220- `POST /reset` -- body `{"seed": int?, "task_id": str?}`221- `POST /step` -- body `{"action": {"patched_code": "..."}}`222- `GET /state`223 224### Environment Variables225 226| Variable | Purpose | Default |227|---|---|---|228| `API_BASE_URL` | LLM inference endpoint | `https://router.huggingface.co/v1` |229| `MODEL_NAME` | Model identifier | `Qwen/Qwen2.5-72B-Instruct` |230| `HF_TOKEN` | API key for inference | (required) |231 232## Baseline Scores233 234We ran the baseline inference script against **Qwen/Qwen2.5-72B-Instruct** via the HuggingFace Router on all 9 tasks.235 236| Task | Difficulty | Steps Used | Final Score | Success |237|---|---|---|---|---|238| easy_01 | Easy | 1 / 5 | 0.94 | yes |239| easy_02 | Easy | 1 / 5 | 0.94 | yes |240| easy_03 | Easy | 1 / 5 | 0.94 | yes |241| medium_01 | Medium | 8 / 8 | 0.50 | yes |242| medium_02 | Medium | 1 / 8 | 0.94 | yes |243| medium_03 | Medium | 1 / 8 | 0.94 | yes |244| hard_01 | Hard | 1 / 12 | 0.94 | yes |245| hard_02 | Hard | 1 / 12 | 0.94 | yes |246| hard_03 | Hard | 2 / 12 | 0.65 | yes |247 248**Mean score across all tasks:** 0.86249 250Full run logs: [baseline_logs_full.txt](baseline_logs_full.txt)251 252### Observations253 254Qwen 2.5 72B one-shots 7 of 9 tasks (score 0.94 -- the one-step ceiling imposed by the step cost and terminal bonus formula). This is expected: with full file content and failing test names in context, frontier models have strong single-file debugging capability and do not need multi-step reasoning for localized bugs.255 256The signal of the environment emerges in two places:257 2581. **medium_01** (CSV quoted-comma parser): Qwen required all 8 steps with a 7-step plateau at 0.44 before converging at 0.94. This is the canonical dense-reward trajectory: partial progress, stuck state, hypothesis revision, convergence. Terminal-only reward functions would give zero signal across steps 1-7; PatchBench's shaped reward provides training signal at every step.259 2602. **hard_03** (MiniJSON encoder/decoder symmetry): Qwen made a partial fix on step 1 (score 0.35) that patched the encoder but broke the round-trip, then converged on step 2 (score 0.94) after seeing the new failing-tests output. This confirms the environment rewards *reading* the feedback loop, not just the initial guess.261 262### What this tells us about benchmark design263 264The 0.94 ceiling for one-shot solutions is intentional -- the step cost (-0.1) discourages infinite iteration, so even a perfect single-step patch cannot reach 1.0. In practice this means:265 266- **Evaluation use case:** PatchBench best differentiates models in the 0.40-0.94 band, where multi-step reasoning matters.267- **Training use case:** The dense reward curve across medium_01's trajectory (0.44 x 7 then 0.94) is precisely the signal an RL fine-tuning loop can exploit to teach iterative debugging behavior.268 269A future version will introduce cross-file patches and tighter step budgets to push more tasks into the training-signal band for frontier models.270 271## Architecture272 273```274patchbench/275โโโ inference.py Baseline inference script (OpenAI client)276โโโ openenv.yaml OpenEnv metadata277โโโ Dockerfile Container build278โโโ pyproject.toml279โโโ requirements.txt280โโโ README.md281โโโ patchbench/ Environment package282โ โโโ __init__.py283โ โโโ models.py Pydantic Observation / Action284โ โโโ environment.py PatchBenchEnv: reset / step / state285โ โโโ grader.py Subprocess-sandboxed pytest grader286โ โโโ tasks/ 9 hand-crafted task scenarios287โ โโโ easy/288โ โโโ medium/289โ โโโ hard/290โโโ server/291 โโโ app.py FastAPI HTTP wrapper for HF Space292```293 294## Design Notes295 296**Sandboxing.** The grader writes patched code to a fresh `tempfile.TemporaryDirectory()` and runs `pytest` as a subprocess with a 15-second timeout. Timeouts, crashes, and parse failures are all caught -- the grader never raises to the caller.297 298**Determinism.** Baseline pass/fail sets for every task are pre-computed and stored in `task.json`. The same action always produces the same reward.299 300**Efficiency.** No task takes more than 12 steps. With three tasks and a typical LLM latency, the full inference script completes well under the 20-minute budget.301 302## License303 304MIT305 