Team Ai
Datasetpublic

rootxhacker/patchbench

PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.

sourceHugging Faceupdated 29d agoView on Hugging Face
0likes124downloads
java.jsonl4 linesDownload Raw Back to candidates
1version https://git-lfs.github.com/spec/v12oid sha256:0b0b6f6ca7b09865bff8fffc5fa4ed35ca372c8b205185d1ff3fb3ef8a82f62a3size 189896454