Precise-Debugging-Benchmarking/PDB-Wild
PDB-Wild: Precise Debugging Benchmarking โ multi-line and repository-level bugs ๐ Paper ยท ๐ป Code ยท ๐ Project page ยท ๐ Leaderboard PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +โฆ See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.
PDB-Wild: Precise Debugging Benchmarking โ multi-line and repository-level bugs
๐ Paper ยท ๐ป Code ยท ๐ Project page ยท ๐ Leaderboard
PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
- Source datasets: BigCodeBench + LiveCodeBench (the 256 examples of PDB-Multi) and SWE-smith repositories (228 examples)
- Sibling datasets: PDB-Single ยท PDB-Single-Full ยท PDB-Multi
TL;DR
Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level precision (were unnecessary lines touched?) and bug-level recall (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit.
Statistics
- Total examples: 484
- Per source dataset:
bigcodebench: 37livecodebench: 219swebench(SWE-smith repositories): 228 โ 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines- Bug count distribution:
bug_count = 1: 220bug_count = 2: 164bug_count = 3: 100
Schema
SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see `dataset/swesmith` in the code repository.
Loading
from datasets import load_dataset
ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test")
example = ds[0]
print(example["buggy_code"])
print(example["gt_solution"])gt_diff is a JSON-encoded string; decode with json.loads(example["gt_diff"]).
License
MIT.
