Team Ai
Datasetpublic

Precise-Debugging-Benchmarking/PDB-Wild

PDB-Wild: Precise Debugging Benchmarking โ€” multi-line and repository-level bugs ๐Ÿ“„ Paper  ยท  ๐Ÿ’ป Code  ยท  ๐ŸŒ Project page  ยท  ๐Ÿ† Leaderboard PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.

sourceHugging Facemitupdated 2d agoView on Hugging Face
0likes35downloads
Dataset Card

PDB-Wild: Precise Debugging Benchmarking โ€” multi-line and repository-level bugs

๐Ÿ“„ Paper  ยท  ๐Ÿ’ป Code  ยท  ๐ŸŒ Project page  ยท  ๐Ÿ† Leaderboard

PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.

TL;DR

Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level precision (were unnecessary lines touched?) and bug-level recall (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit.

Statistics

  • โ€”Total examples: 484
  • โ€”Per source dataset:
  • โ€”bigcodebench: 37
  • โ€”livecodebench: 219
  • โ€”swebench (SWE-smith repositories): 228 โ€” 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines
  • โ€”Bug count distribution:
  • โ€”bug_count = 1: 220
  • โ€”bug_count = 2: 164
  • โ€”bug_count = 3: 100

Schema

fieldtypenotes
task_idstringunique identifier per buggy variant
source_datasetstringprovenance of the underlying program (bigcodebench, livecodebench, swebench = SWE-smith repository)
source_modelstringbug-generator model
task_promptstringnatural-language description of the task / fix target
gt_solutionstringverified correct program (for SWE-smith: the full content of target_file)
buggy_codestringprogram with injected bug(s)
gt_diffstring (JSON){line_no: {type, original, modified}} mapping โ€” the fix
bug_countintnumber of independent bug blocks (range: {1, 2, 3})
bug_type, bug_subtypestring \nullOrthogonal Defect Classification label (populated for bug_count == 1)
gt_lengthintline count of gt_solution
editable_lines, deletable_lines, frozen_linesint \nullhandler-derived line counts
is_buggybool \nulltrue for single-bug examples, null for composed multi-bug examples
repo, image_name, target_filestring \nullSWE-smith repository, Docker image and file path (SWE-smith examples only)

SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see `dataset/swesmith` in the code repository.

Loading

python
from datasets import load_dataset
ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test")
example = ds[0]
print(example["buggy_code"])
print(example["gt_solution"])

gt_diff is a JSON-encoded string; decode with json.loads(example["gt_diff"]).

License

MIT.