Precise-Debugging-Benchmarking/PDB-Wild
PDB-Wild: Precise Debugging Benchmarking โ multi-line and repository-level bugs ๐ Paper ยท ๐ป Code ยท ๐ Project page ยท ๐ Leaderboard PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +โฆ See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.
035
1---2language:3- en4- code5license: mit6task_categories:7- text-generation8size_categories:9- n<1K10tags:11- code12- debugging13- benchmark14configs:15- config_name: default16 data_files:17 - split: test18 path: data/test-*.parquet19---20 21# PDB-Wild: Precise Debugging Benchmarking โ multi-line and repository-level bugs22 23๐ [Paper](https://arxiv.org/abs/2604.17338) ยท 24๐ป [Code](https://github.com/Bill1235813/PDB) ยท 25๐ [Project page](https://precise-debugging-benchmark.github.io/) ยท 26๐ [Leaderboard](https://precise-debugging-benchmark.github.io/leaderboard.html)27 28`PDB-Wild` is the **multi-line and repository-level bug set** of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (`gt_diff`) that encodes the minimal correct fix.29 30- **Source datasets:** [BigCodeBench](https://huggingface.co/datasets/bigcode/bigcodebench) + [LiveCodeBench](https://huggingface.co/datasets/livecodebench/execution) (the 256 examples of [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)) and [SWE-smith](https://github.com/SWE-bench/SWE-smith) repositories (228 examples)31- **Sibling datasets:** [PDB-Single](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single) ยท [PDB-Single-Full](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full) ยท [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)32 33## TL;DR34 35Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level **precision** (were unnecessary lines touched?) and bug-level **recall** (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit.36 37## Statistics38 39- **Total examples:** 48440- **Per source dataset:**41 - `bigcodebench`: 3742 - `livecodebench`: 21943 - `swebench` (SWE-smith repositories): 228 โ 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines44- **Bug count distribution:**45 - `bug_count = 1`: 22046 - `bug_count = 2`: 16447 - `bug_count = 3`: 10048 49## Schema50 51| field | type | notes |52|---|---|---|53| `task_id` | string | unique identifier per buggy variant |54| `source_dataset` | string | provenance of the underlying program (`bigcodebench`, `livecodebench`, `swebench` = SWE-smith repository) |55| `source_model` | string | bug-generator model |56| `task_prompt` | string | natural-language description of the task / fix target |57| `gt_solution` | string | verified correct program (for SWE-smith: the full content of `target_file`) |58| `buggy_code` | string | program with injected bug(s) |59| `gt_diff` | string (JSON) | `{line_no: {type, original, modified}}` mapping โ the fix |60| `bug_count` | int | number of independent bug blocks (range: {1, 2, 3}) |61| `bug_type`, `bug_subtype` | string \| null | Orthogonal Defect Classification label (populated for `bug_count == 1`) |62| `gt_length` | int | line count of `gt_solution` |63| `editable_lines`, `deletable_lines`, `frozen_lines` | int \| null | handler-derived line counts |64| `is_buggy` | bool \| null | `true` for single-bug examples, null for composed multi-bug examples |65| `repo`, `image_name`, `target_file` | string \| null | SWE-smith repository, Docker image and file path (SWE-smith examples only) |66 67SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see [`dataset/swesmith`](https://github.com/Bill1235813/PDB/tree/main/dataset/swesmith) in the code repository.68 69## Loading70 71```python72from datasets import load_dataset73ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test")74example = ds[0]75print(example["buggy_code"])76print(example["gt_solution"])77```78 79`gt_diff` is a JSON-encoded string; decode with `json.loads(example["gt_diff"])`.80 81## License82 83MIT.84 