lubem/patchpilot-agentic-ai
PatchPilot: A Tool-Using Agent for Python Debugging and Repair
PatchPilot is a bounded Agentic AI system for autonomous Python debugging and repair. It treats software repair as an agent-environment loop: reproduce the failure, inspect relevant code, form a repair hypothesis, apply a controlled patch, verify with tests, and stop only when the current repository revision is executable and verified.
The project focuses on safe, auditable agent execution rather than unrestricted code generation. The language model proposes repair actions inside a constrained runtime; the runtime validates tool calls, restricts file access, applies patches safely, records traces, enforces budgets, and confirms success through pytest.
Key Features
- Single-agent Plan-Act-Observe-Reflect-Verify workflow
- Restricted repository tools for listing, reading, searching, testing, patching, diff inspection, and rollback
- Isolated benchmark workspaces for repair attempts
- Path restrictions to prevent edits outside allowed source files
- Patch validation and rollback support
- Loop guard against repeated identical tool actions
- Execution budgets for steps, tool calls, patch attempts, and runtime
- Auditable JSON traces for every repair run
- Ollama/Qwen local-model integration
- Final success allowed only after full-suite verification
Current Status
PatchPilot has a working end-to-end live repair workflow using a local Ollama model.
Verified live workflow:
run_tests -> search_code -> read_file -> apply_patch -> run_tests -> finishThe current local quality gate passes:
pytest: 107 passed
ruff: passed
mypy: passed
git diff --check: passedA successful live run repairs the seeded calculator benchmark and finishes only after the test suite passes.
Architecture
User / Benchmark Task
|
v
Agent State
|
v
Structured LLM Policy
|
v
Validated Tool Action
|
v
Restricted Tool Executor
|
+--> Repository tools
+--> Test runner
+--> Patch manager
+--> Rollback / diff tools
|
v
Tool Observation
|
v
Trace Recorder
|
v
Verified Finish / Escalation / FailurePatchPilot is intentionally designed so the model does not directly mutate the repository. The model proposes actions, while the runtime validates and executes them.
Model Backend
The live local pilot uses Ollama with Qwen2.5-Coder.
The default live script currently uses:
qwen2.5-coder:1.5bThis model was selected for reliable CPU-only execution in a low-memory local environment. The pipeline is model-swappable: stronger local or hosted code models can be connected through the same text-generation interface.
Repository Layout
benchmarks/ Seed repair tasks
docs/ Project documentation
scripts/ Demo and live-run scripts
src/patchpilot/agent/ Agent policy, executor, loop control, tracing
src/patchpilot/tools/ Repository, test, and patch tools
src/patchpilot/models/ Model backends
src/patchpilot/schemas/ Shared state and tool schemas
tests/ Unit and integration testsQuickstart
Create and activate a Python 3.12 virtual environment:
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"Run the local quality gate:
python -m pytest -q
python -m ruff check .
python -m mypy src
git diff --checkRunning the Live Ollama Demo
Install and start Ollama, then pull the local model:
ollama pull qwen2.5-coder:1.5bRun the live repair demo:
python scripts/run_live_qwen.pyA successful run prints the agent status, tool steps, changed files, workspace path, and trace path.
Expected successful flow:
STATUS=succeeded
STEP_1=run_tests|error
STEP_2=search_code|ok
STEP_3=read_file|ok
STEP_4=apply_patch|ok
STEP_5=run_tests|ok
STEP_6=finish|okInteractive Demo
Hosted demo: https://lubem-patchpilot-agentic-ai.hf.space/
Hugging Face Space: https://huggingface.co/spaces/lubem/patchpilot-agentic-ai
PatchPilot includes a Streamlit prototype that allows users to select a benchmark task, inspect the broken source and tests, run the agent locally, view the tool-use trace, inspect the patch diff, and confirm final pytest verification.
Install demo dependencies:
python -m pip install -r demo/requirements-demo.txtRun the demo:
streamlit run demo/streamlit_app.py --server.headless true --browser.gatherUsageStats falseOpen:
http://localhost:8501The live repair button requires Ollama and the selected model to be available locally.
Docker Demo
Build and run the demo container:
docker compose up --buildThen open:
http://localhost:8501The Docker demo packages the repository and Streamlit frontend. Live Ollama execution requires access to a running Ollama service from the container environment.
Safety Design
PatchPilot uses a restricted execution boundary:
- tool calls are schema-validated before execution;
- file access is limited to repository-relative paths;
- benchmark tests are protected from modification;
- patch attempts are budgeted;
- repeated identical actions are rejected;
- success requires full-suite verification on the current revision;
- every action and observation is recorded in an auditable trace.
Evaluation Plan
The completed evaluation compares the full PatchPilot agent against:
- one-shot patch generation;
- a no-retry/reduced-budget ablation.
A supplementary QuixBugs smoke test evaluates external public-benchmark behavior.
Primary metrics:
- full repair rate;
- full regression-test pass rate;
- invalid patch rate;
- average repair attempts;
- tool-call count;
- execution time;
- rollback frequency;
- budget exhaustion rate.
Project Scope
PatchPilot targets small Python repositories with executable pytest suites. It is designed for research and demonstration of bounded agentic software repair, not unrestricted production code modification.
License
No open-source license has been declared yet.
