rootxhacker/patchbench
PatchBench A multi-language benchmark for evaluating whether an LLM can fix a bug correctly and without introducing new bugs. Each row presents a real issue report plus a large code context (~500–1500 lines); the model under test produces a patch; a per-row test suite then checks that (a) previously failing tests now pass and (b) the rest of the suite stays green. Status: v0 spec. Rows are being built. Row schema Field Type In tasks config Description… See the full description on the dataset page: https://huggingface.co/datasets/rootxhacker/patchbench.
0124
1version https://git-lfs.github.com/spec/v12oid sha256:0b0b6f6ca7b09865bff8fffc5fa4ed35ca372c8b205185d1ff3fb3ef8a82f62a3size 189896454 