Team Ai
Datasetpublic

rootxhacker/patchbench

PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.

sourceHugging Faceupdated 27d agoView on Hugging Face
0likes283downloads
Dataset Card

PatchBench

A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green.

Status: v0 spec. Rows are being built.

Row schema

FieldTypeIn `tasks` configDescription
idstring✅patchbench-{lang}-{source}-{nnnn}
languagestring✅python \javascript \typescript \go \java \rust
sourcestring✅real (mined bug-fix PR) \synthetic (injected bug)
originstring✅Source dataset, e.g. swe-bench-live-multilang, swesmith-go
repostring✅GitHub repo the code came from
base_commitstring✅Commit the buggy state is taken at
issue_textstring✅Issue/PR report the model must diagnose from
context_fileslist of {path, content}✅~500–1500 lines total incl. the buggy file(s), callers, and related modules
buggy_fileslist of string✅Paths the fix is expected to touch
test_fileslist of {path, content}✅Test patch materialized as files
test_cmdslist of string✅Commands that run the tests for this row (language-native)
env_setup_cmdslist of string❌ (harness only)Install/build commands needed before test_cmds
fail_to_passlist of string❌Tests that MUST fail pre-fix and pass post-fix
pass_to_passlist of string❌Tests that MUST stay passing post-fix (the regression check)
gold_patchstring❌Reference fix, for harness/self-eval only

Two configs prevent leakage:

  • —`tasks` (public eval surface): everything an agent sees — issue, context, tests.
  • —`gold` (harness-only): fail_to_pass, pass_to_pass, gold_patch, env_setup_cmds.

Verification contract (every released row)

  1. 1.Pre-fix: apply test_files to the repo at base_commit → every fail_to_pass test fails; every pass_to_pass test passes.
  2. 2.Post-fix: apply gold_patch → every fail_to_pass and pass_to_pass test passes.
  3. 3.Determinism: each check is run twice; any row with a flaky test is dropped.
  4. 4.No-leak check: issue text must not name the exact function/line to change (auto-screen + spot manual review).

A model's fix is scored: resolved iff all fail_to_pass pass AND all pass_to_pass still pass. Rows also record per-test results so partial behavior (fix works but broke something) is visible — that is the "introduced a new bug" signal.

Sourcing plan (target ~1,500–2,000 rows)

LanguageReal sourceSynthetic top-upTarget
PythonSWE-Gym + SWE-bench Verified/trainSWE-smith-py~550
RustSWE-bench-Live rust + rustbenchSWE-smith-rs~300
JavaSWE-bench-Live java + Multi-SWE-bench javaSWE-smith-java~300
GoSWE-bench-Live go + Multi-SWE-bench goSWE-smith-go~250
JS/TSSWE-bench-Live js+ts + Multi-SWE-benchSWE-smith-js/ts~400

Context expansion is PatchBench's own contribution: no upstream dataset ships sliced 500–1500-line windows. We materialize the repo at base_commit, compute the failing-test dependency closure, and slice the modules inside it, filtering out rows whose closure exceeds the window.

Harness

Language-native runner (no Docker required, adapted from the SWE-bench grading loop + microsoft/RepoLaunch per-language build/test recipes): clone at base_commit → env_setup_cmds → apply test_files (+ candidate patch) → run test_cmds → parse per-test results → grade against F2P/P2P. Per-repo env caching makes verification tractable.

License audit

Upstream dataset licenses: SWE-bench-Live (MIT), SWE-Gym (MIT), SWE-smith (MIT). Multi-SWE-bench / PrimeIntellect rows are CC0/“other” (ByteDance community terms) — those rows are re-verified from original permissively-licensed repos before inclusion, or excluded.