Team Ai
Datasetpublic

rootxhacker/patchbench

PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.

sourceHugging Faceupdated 27d agoView on Hugging Face
0likes283downloads

rootxhacker/patchbench · main · files are served by the source, never re-hosted here