lubem/patchpilot-agentic-ai
0
1# Accepted Project Scope2 3## Goal4 5Repair small Python defects through a bounded6Plan–Act–Observe–Reflect–Verify loop.7 8## Required capabilities9 10- inspect a repository and failing tests;11- maintain structured agent state and execution budgets;12- choose restricted debugging tools dynamically;13- apply minimal source-code patches;14- run targeted and full regression tests;15- reflect and replan after unsuccessful attempts;16- roll back invalid or unsafe changes;17- terminate with verified success, bounded failure, or escalation;18- save auditable execution traces.19 20## Evaluation21 22The full agent will be compared with:23 241. one-shot patch generation;252. a fixed one-pass repair workflow;263. a tool-using agent without reflection.27 28Primary metrics are complete repair rate and full regression-test pass rate.29Secondary metrics include invalid patches, tool calls, repair attempts,30latency, rollback frequency, patch size, and budget exhaustion.31 32## Scope controls33 34- Python repositories only;35- approximately 10–12 validated repair tasks;36- mutation-seeded defects reviewed before evaluation;37- one primary agent;38- no unrestricted shell access;39- no editing benchmark tests;40- no success claim without executable verification.41 