lubem/patchpilot-agentic-ai
0
1# PatchPilot Statistical Analysis2 3This document records the reproducible statistical comparison between the full PatchPilot agent and the one-shot live Qwen baseline on PatchPilot-Bench v0.4 5## Inputs6 7- Full-agent runs CSV: `artifacts/evaluation/20260704-091749/runs.csv`8- One-shot baseline runs CSV: `artifacts/evaluation/20260704-135312/runs.csv`9- Paired tasks: 1210 11## Repair Success12 13| Condition | Successes | Repair Rate |14| --- | ---: | ---: |15| `full-agent-live-qwen` | 12/12 | 100.0% |16| `one-shot-live-qwen` | 8/12 | 66.7% |17 18## Paired Success Table19 20| Outcome | Count |21| --- | ---: |22| Both succeeded | 8 |23| Full agent only succeeded | 4 |24| One-shot baseline only succeeded | 0 |25| Both failed | 0 |26 27## Exact McNemar Test28 29Because both conditions ran on the same benchmark tasks, repair success is compared as paired binary outcomes.30 31| Test | Value |32| --- | ---: |33| Discordant pairs | 4 |34| Exact McNemar two-sided p-value | 0.1250 |35 36## Interpretation37 38The full agent repaired 12/12 tasks, while the one-shot baseline repaired 8/12 tasks. The paired effect size is 4 tasks repaired only by the full agent versus 0 tasks repaired only by the one-shot baseline.39 40Because the benchmark currently contains 12 tasks, p-values should be interpreted cautiously. The strongest evidence is the paired success difference and the fact that the full agent succeeds on tasks where a single-pass repair attempt fails verification.41 