Precise-Debugging-Benchmarking/PDB-Single-Full
PDB-Single-Full: Precise Debugging Benchmarking โ unfiltered single-line bug pool ๐ Paper ยท ๐ป Code ยท ๐ Project page ยท ๐ Leaderboard PDB-Single-Full is the unfiltered single-line bug pool of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +โฆ See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full.
084
1---2language:3- en4- code5license: mit6task_categories:7- text-generation8size_categories:9- 1K<n<10K10tags:11- code12- debugging13- benchmark14configs:15- config_name: default16 data_files:17 - split: test18 path: data/test-*.parquet19---20 21# PDB-Single-Full: Precise Debugging Benchmarking โ unfiltered single-line bug pool22 23๐ [Paper](https://arxiv.org/abs/2604.17338) ยท 24๐ป [Code](https://github.com/Bill1235813/PDB) ยท 25๐ [Project page](https://precise-debugging-benchmark.github.io/) ยท 26๐ [Leaderboard](https://precise-debugging-benchmark.github.io/leaderboard.html)27 28`PDB-Single-Full` is the **unfiltered single-line bug pool** of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (`gt_diff`) that encodes the minimal correct fix.29 30- **Source datasets:** [BigCodeBench](https://huggingface.co/datasets/bigcode/bigcodebench) + [LiveCodeBench](https://huggingface.co/datasets/livecodebench/execution)31- **Sibling datasets:** [PDB-Single](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single) ยท [PDB-Wild](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild) ยท [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)32 33## TL;DR34 35Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level **precision** (were unnecessary lines touched?) and bug-level **recall** (were all faults resolved?). Experiments on PDB-Single-Full show frontier models score above 76% on unit tests but at or below 45% on precision โ they over-edit.36 37## Statistics38 39- **Total examples:** 758940- **Per source dataset:**41 - `bigcodebench`: 369742 - `livecodebench`: 389243- **Bug count distribution:**44 - `bug_count = 1`: 237545 - `bug_count = 2`: 233046 - `bug_count = 3`: 189447 - `bug_count = 4`: 99048- **Source-model mix (bug generator):**49 - `gpt-5.1-codex`: 265650 - `gemini-2.5-pro`: 260851 - `claude-sonnet-4.5`: 232552 53## Schema54 55| field | type | notes |56|---|---|---|57| `task_id` | string | unique identifier, includes `_<idx>` suffix per bug variant |58| `source_dataset` | string | `bigcodebench` or `livecodebench` |59| `source_model` | string | generator model that produced the bug |60| `task_prompt` | string | natural-language problem statement |61| `gt_solution` | string | verified correct program |62| `buggy_code` | string | program with injected bug(s) |63| `gt_diff` | string (JSON) | `{line_no: {type, original, modified}}` mapping โ the fix |64| `bug_count` | int | number of independent bug blocks (range: {1, 2, 3, 4}) |65| `bug_type`, `bug_subtype` | string | Orthogonal Defect Classification label (populated for `bug_count == 1`; omitted for composed multi-bug entries) |66| `gt_length` | int | line count of `gt_solution` |67| `editable_lines`, `deletable_lines`, `frozen_lines` | int | handler-derived line counts |68| `is_buggy` | bool | always `true` in the released splits |69 70## Loading71 72```python73from datasets import load_dataset74ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Single-Full", split="test")75example = ds[0]76print(example["buggy_code"])77print(example["gt_solution"])78```79 80`gt_diff` is a JSON-encoded string; decode with `json.loads(example["gt_diff"])`.81 82## Debugging with a model83 84The companion code repo ships a turn-key driver:85 86```bash87git clone https://github.com/Bill1235813/PDB88cd PDB89uv sync90# set your key in keys/<provider>_key.txt, then:91bash scripts/simple_debug_eval.sh single openai/gpt-5.1-codex92```93 94This loops your model over both BCB and LCB subsets, writes debug outputs under `results/<bench>/debug_results/`, and computes Unit / Precision / Recall / F1 per task.95 96To score a saved debug-results file directly (without rerunning the model):97 98```bash99python src/evaluator.py \100 --dataset_name bigcodebench \101 --eval_model_name my-model \102 --input_file <model>_on_bigcodebench_pdb_single.json \103 --eval_set_name bigcodebench_pdb_single104```105 106## How PDB works107 1081. **Bug synthesis.** An LLM generator rewrites one line of `gt_solution` following the Orthogonal Defect Classification ([Chillarege et al., 1992](https://ieeexplore.ieee.org/document/177364)). Each candidate is unit-tested: it must fail the tests *and* every proper-subset partial fix must still fail (the **atomicity check**, preventing compound-independent bugs).1092. **Composition.** Valid single-bug variants are composed into `bug_count โ {1, 2, 3, 4}` programs with a stride constraint between blocks so that bug regions never touch or overlap.1103. **Evaluation.** For a model's patch, PDB reports:111 - **Unit score** โ does the patch pass hidden tests?112 - **Precision** โ fraction of edited lines that fall inside the GT edit regions (strict; default tolerance ฮต=0).113 - **Recall** โ fraction of GT edit blocks that the patch resolves.114 - **F1** over the above.115 116## Citation117 118```119@article{zhu2026pdb,120 title={Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?},121 author={Zhu, Wang Bill and Chai, Miaosen and Wang, Shangshang and Liu, Yejia and Bian, Song and Dong, Honghua and Neiswanger, Willie and Jia, Robin},122 journal={arXiv preprint arXiv:2604.17338},123 year={2026}124}125```126 127## License128 129MIT.130 