Team Ai
Datasetpublic

Precise-Debugging-Benchmarking/PDB-Wild

PDB-Wild: Precise Debugging Benchmarking โ€” multi-line and repository-level bugs ๐Ÿ“„ Paper  ยท  ๐Ÿ’ป Code  ยท  ๐ŸŒ Project page  ยท  ๐Ÿ† Leaderboard PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild.

sourceHugging Facemitupdated 3d agoView on Hugging Face
0likes35downloads
README.md84 linesDownload Raw Back to root
1---2language:3- en4- code5license: mit6task_categories:7- text-generation8size_categories:9- n<1K10tags:11- code12- debugging13- benchmark14configs:15- config_name: default16  data_files:17  - split: test18    path: data/test-*.parquet19---20 21# PDB-Wild: Precise Debugging Benchmarking โ€” multi-line and repository-level bugs22 23๐Ÿ“„ [Paper](https://arxiv.org/abs/2604.17338) &nbsp;ยท&nbsp;24๐Ÿ’ป [Code](https://github.com/Bill1235813/PDB) &nbsp;ยท&nbsp;25๐ŸŒ [Project page](https://precise-debugging-benchmark.github.io/) &nbsp;ยท&nbsp;26๐Ÿ† [Leaderboard](https://precise-debugging-benchmark.github.io/leaderboard.html)27 28`PDB-Wild` is the **multi-line and repository-level bug set** of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (`gt_diff`) that encodes the minimal correct fix.29 30- **Source datasets:** [BigCodeBench](https://huggingface.co/datasets/bigcode/bigcodebench) + [LiveCodeBench](https://huggingface.co/datasets/livecodebench/execution) (the 256 examples of [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)) and [SWE-smith](https://github.com/SWE-bench/SWE-smith) repositories (228 examples)31- **Sibling datasets:** [PDB-Single](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single) ยท [PDB-Single-Full](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full) ยท [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)32 33## TL;DR34 35Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level **precision** (were unnecessary lines touched?) and bug-level **recall** (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit.36 37## Statistics38 39- **Total examples:** 48440- **Per source dataset:**41  - `bigcodebench`: 3742  - `livecodebench`: 21943  - `swebench` (SWE-smith repositories): 228 โ€” 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines44- **Bug count distribution:**45  - `bug_count = 1`: 22046  - `bug_count = 2`: 16447  - `bug_count = 3`: 10048 49## Schema50 51| field | type | notes |52|---|---|---|53| `task_id` | string | unique identifier per buggy variant |54| `source_dataset` | string | provenance of the underlying program (`bigcodebench`, `livecodebench`, `swebench` = SWE-smith repository) |55| `source_model` | string | bug-generator model |56| `task_prompt` | string | natural-language description of the task / fix target |57| `gt_solution` | string | verified correct program (for SWE-smith: the full content of `target_file`) |58| `buggy_code` | string | program with injected bug(s) |59| `gt_diff` | string (JSON) | `{line_no: {type, original, modified}}` mapping โ€” the fix |60| `bug_count` | int | number of independent bug blocks (range: {1, 2, 3}) |61| `bug_type`, `bug_subtype` | string \| null | Orthogonal Defect Classification label (populated for `bug_count == 1`) |62| `gt_length` | int | line count of `gt_solution` |63| `editable_lines`, `deletable_lines`, `frozen_lines` | int \| null | handler-derived line counts |64| `is_buggy` | bool \| null | `true` for single-bug examples, null for composed multi-bug examples |65| `repo`, `image_name`, `target_file` | string \| null | SWE-smith repository, Docker image and file path (SWE-smith examples only) |66 67SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see [`dataset/swesmith`](https://github.com/Bill1235813/PDB/tree/main/dataset/swesmith) in the code repository.68 69## Loading70 71```python72from datasets import load_dataset73ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test")74example = ds[0]75print(example["buggy_code"])76print(example["gt_solution"])77```78 79`gt_diff` is a JSON-encoded string; decode with `json.loads(example["gt_diff"])`.80 81## License82 83MIT.84