Team Ai
Datasetpublic

Precise-Debugging-Benchmarking/PDB-Single-Full

PDB-Single-Full: Precise Debugging Benchmarking โ€” unfiltered single-line bug pool ๐Ÿ“„ Paper  ยท  ๐Ÿ’ป Code  ยท  ๐ŸŒ Project page  ยท  ๐Ÿ† Leaderboard PDB-Single-Full is the unfiltered single-line bug pool of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench +โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full.

sourceHugging Facemitupdated 3d agoView on Hugging Face
0likes84downloads
README.md130 linesDownload Raw Back to root
1---2language:3- en4- code5license: mit6task_categories:7- text-generation8size_categories:9- 1K<n<10K10tags:11- code12- debugging13- benchmark14configs:15- config_name: default16  data_files:17  - split: test18    path: data/test-*.parquet19---20 21# PDB-Single-Full: Precise Debugging Benchmarking โ€” unfiltered single-line bug pool22 23๐Ÿ“„ [Paper](https://arxiv.org/abs/2604.17338) &nbsp;ยท&nbsp;24๐Ÿ’ป [Code](https://github.com/Bill1235813/PDB) &nbsp;ยท&nbsp;25๐ŸŒ [Project page](https://precise-debugging-benchmark.github.io/) &nbsp;ยท&nbsp;26๐Ÿ† [Leaderboard](https://precise-debugging-benchmark.github.io/leaderboard.html)27 28`PDB-Single-Full` is the **unfiltered single-line bug pool** of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (`gt_diff`) that encodes the minimal correct fix.29 30- **Source datasets:** [BigCodeBench](https://huggingface.co/datasets/bigcode/bigcodebench) + [LiveCodeBench](https://huggingface.co/datasets/livecodebench/execution)31- **Sibling datasets:** [PDB-Single](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single) ยท [PDB-Wild](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild) ยท [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)32 33## TL;DR34 35Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level **precision** (were unnecessary lines touched?) and bug-level **recall** (were all faults resolved?). Experiments on PDB-Single-Full show frontier models score above 76% on unit tests but at or below 45% on precision โ€” they over-edit.36 37## Statistics38 39- **Total examples:** 758940- **Per source dataset:**41  - `bigcodebench`: 369742  - `livecodebench`: 389243- **Bug count distribution:**44  - `bug_count = 1`: 237545  - `bug_count = 2`: 233046  - `bug_count = 3`: 189447  - `bug_count = 4`: 99048- **Source-model mix (bug generator):**49  - `gpt-5.1-codex`: 265650  - `gemini-2.5-pro`: 260851  - `claude-sonnet-4.5`: 232552 53## Schema54 55| field | type | notes |56|---|---|---|57| `task_id` | string | unique identifier, includes `_<idx>` suffix per bug variant |58| `source_dataset` | string | `bigcodebench` or `livecodebench` |59| `source_model` | string | generator model that produced the bug |60| `task_prompt` | string | natural-language problem statement |61| `gt_solution` | string | verified correct program |62| `buggy_code` | string | program with injected bug(s) |63| `gt_diff` | string (JSON) | `{line_no: {type, original, modified}}` mapping โ€” the fix |64| `bug_count` | int | number of independent bug blocks (range: {1, 2, 3, 4}) |65| `bug_type`, `bug_subtype` | string | Orthogonal Defect Classification label (populated for `bug_count == 1`; omitted for composed multi-bug entries) |66| `gt_length` | int | line count of `gt_solution` |67| `editable_lines`, `deletable_lines`, `frozen_lines` | int | handler-derived line counts |68| `is_buggy` | bool | always `true` in the released splits |69 70## Loading71 72```python73from datasets import load_dataset74ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Single-Full", split="test")75example = ds[0]76print(example["buggy_code"])77print(example["gt_solution"])78```79 80`gt_diff` is a JSON-encoded string; decode with `json.loads(example["gt_diff"])`.81 82## Debugging with a model83 84The companion code repo ships a turn-key driver:85 86```bash87git clone https://github.com/Bill1235813/PDB88cd PDB89uv sync90# set your key in keys/<provider>_key.txt, then:91bash scripts/simple_debug_eval.sh single openai/gpt-5.1-codex92```93 94This loops your model over both BCB and LCB subsets, writes debug outputs under `results/<bench>/debug_results/`, and computes Unit / Precision / Recall / F1 per task.95 96To score a saved debug-results file directly (without rerunning the model):97 98```bash99python src/evaluator.py \100  --dataset_name bigcodebench \101  --eval_model_name my-model \102  --input_file <model>_on_bigcodebench_pdb_single.json \103  --eval_set_name bigcodebench_pdb_single104```105 106## How PDB works107 1081. **Bug synthesis.** An LLM generator rewrites one line of `gt_solution` following the Orthogonal Defect Classification ([Chillarege et al., 1992](https://ieeexplore.ieee.org/document/177364)). Each candidate is unit-tested: it must fail the tests *and* every proper-subset partial fix must still fail (the **atomicity check**, preventing compound-independent bugs).1092. **Composition.** Valid single-bug variants are composed into `bug_count โˆˆ {1, 2, 3, 4}` programs with a stride constraint between blocks so that bug regions never touch or overlap.1103. **Evaluation.** For a model's patch, PDB reports:111   - **Unit score** โ€” does the patch pass hidden tests?112   - **Precision** โ€” fraction of edited lines that fall inside the GT edit regions (strict; default tolerance ฮต=0).113   - **Recall** โ€” fraction of GT edit blocks that the patch resolves.114   - **F1** over the above.115 116## Citation117 118```119@article{zhu2026pdb,120  title={Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?},121  author={Zhu, Wang Bill and Chai, Miaosen and Wang, Shangshang and Liu, Yejia and Bian, Song and Dong, Honghua and Neiswanger, Willie and Jia, Robin},122  journal={arXiv preprint arXiv:2604.17338},123  year={2026}124}125```126 127## License128 129MIT.130