rootxhacker/patchbench
PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.
PatchBench
A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green.
Status: v0 spec. Rows are being built.
Row schema
Two configs prevent leakage:
- `tasks` (public eval surface): everything an agent sees — issue, context, tests.
- `gold` (harness-only):
fail_to_pass,pass_to_pass,gold_patch,env_setup_cmds.
Verification contract (every released row)
- Pre-fix: apply
test_filesto the repo atbase_commit→ everyfail_to_passtest fails; everypass_to_passtest passes. - Post-fix: apply
gold_patch→ everyfail_to_passandpass_to_passtest passes. - Determinism: each check is run twice; any row with a flaky test is dropped.
- No-leak check: issue text must not name the exact function/line to change (auto-screen + spot manual review).
A model's fix is scored: resolved iff all fail_to_pass pass AND all pass_to_pass still pass. Rows also record per-test results so partial behavior (fix works but broke something) is visible — that is the "introduced a new bug" signal.
Sourcing plan (target ~1,500–2,000 rows)
Context expansion is PatchBench's own contribution: no upstream dataset ships sliced 500–1500-line windows. We materialize the repo at base_commit, compute the failing-test dependency closure, and slice the modules inside it, filtering out rows whose closure exceeds the window.
Harness
Language-native runner (no Docker required, adapted from the SWE-bench grading loop + microsoft/RepoLaunch per-language build/test recipes): clone at base_commit → env_setup_cmds → apply test_files (+ candidate patch) → run test_cmds → parse per-test results → grade against F2P/P2P. Per-repo env caching makes verification tractable.
License audit
Upstream dataset licenses: SWE-bench-Live (MIT), SWE-Gym (MIT), SWE-smith (MIT). Multi-SWE-bench / PrimeIntellect rows are CC0/“other” (ByteDance community terms) — those rows are re-verified from original permissively-licensed repos before inclusion, or excluded.
