Team Ai
Datasetpublic

rootxhacker/patchbench

PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.

sourceHugging Faceupdated 28d agoView on Hugging Face
0likes284downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
rootxhacker/patchbench · Team Ai