Tejaswiniprabhakaran19/patchpilot-patchgen
PatchPilot patch-generation dataset Supervised fine-tuning chats for PatchPilot's patch generator (A2). Each chat is exactly the prompt PatchPilot's agent sends to its model, followed by the developers' real fix written in the agent's SEARCH/REPLACE edit format. Source Built by scripts/build_patchgen_data.py (seed 42) from the SWE-bench training split (princeton-nlp/SWE-bench, train) and the gold files' contents from princeton-nlp/SWE-bench_oracle. The same seeded… See the full description on the dataset page: https://huggingface.co/datasets/Tejaswiniprabhakaran19/patchpilot-patchgen.
PatchPilot patch-generation dataset
Supervised fine-tuning chats for PatchPilot's patch generator (A2). Each chat is exactly the prompt PatchPilot's agent sends to its model, followed by the developers' real fix written in the agent's SEARCH/REPLACE edit format.
Source
Built by scripts/build_patchgen_data.py (seed 42) from the SWE-bench training split (princeton-nlp/SWE-bench, train) and the gold files' contents from princeton-nlp/SWE-bench_oracle. The same seeded selection (at most 400 instances per repository) and the same repository-level train/validation/test splits as the localization dataset are used.
Leakage check
The training split's 35 repositories are disjoint from SWE-bench Lite's 12. scripts/check_leakage.py found zero overlap with all 300 SWE-bench Lite test instances on repository, instance id, normalised issue text and normalised gold patch (results/leakage_check.json).
Construction
- Filter (same shape as SWE-bench Lite's own selection): the gold patch edits exactly one existing non-test Python file, in at most three hunks. Larger multi-file fixes are skipped.
- Target: each hunk becomes a SEARCH/REPLACE block. If the original lines are not unique in the file, real neighbouring lines are added until they are. Applying the blocks is checked to give exactly the same file as
git applyof the gold patch; examples that fail are skipped. - Prompt: the agent's system prompt; a user turn with the issue text (max 3,000 characters), the tests added by the fix (max 1,500 characters) and the gold file with line numbers, whole if at most 150 lines, otherwise windows of 15 lines around each edit.
- Length: user + assistant at most 10,000 characters (about 2,800 Gemma tokens).
Size
From results/patchgen/dataset_stats.json:
Of the 8,867 selected training-split instances, 4,137 became chats. Skipped: 4,113 not SWE-bench-Lite-shaped (more than one file or more than three hunks), 235 too long, 197 with no existing non-test Python file edited, 101 absent from SWE-bench_oracle, 78 whose diff could not be turned into unique SEARCH/REPLACE blocks that reproduce git apply, and 6 whose edited file was missing from the oracle text. User + assistant length: median 5,207 characters, maximum 9,995. The splits use the same repositories as the localization dataset.
Format
{train,val,test}.jsonl.gz, one chat per line:
{"instance_id": "...", "repo": "owner/name", "split": "train", "files": ["pkg/mod.py"],
"messages": [{"role": "system", "content": "..."},
{"role": "user", "content": "## Bug report ..."},
{"role": "assistant", "content": "Fix:\n\npkg/mod.py\n<<<<<<< SEARCH\n..."}]}Limitations
- The context shows the gold file around the edit, so the model learns to fix given good localization; at run time it sees the agent's localized files, which may be wrong.
- The tests are shown as the test patch, not as the failing test output the agent sees at run time.
- Only small, single-file fixes; the developers' fix is one correct answer among possibly many.
Licence
Derived from SWE-bench (MIT). The underlying code belongs to the respective open-source projects under their own licences.
