Krishp1/Autonomous-Coding-Agent
0
1# π€ Autonomous Python Coding Agent2 3> **A production-grade, self-healing multi-agent pipeline that doesn't just generate Python code β it autonomously writes, validates, tests, secures, benchmarks, and reflects on its own output before shipping.**4 5[](https://python.org)6[](https://github.com/langchain-ai/langgraph)7[](https://groq.com)8[](https://chromadb.com)9[](https://streamlit.io)10[](LICENSE)11[](https://huggingface.co/spaces/krishpatel/autonomous-coding-agent)12 13---14 15## π Live Demo16 17**[βΆ Try it on Hugging Face Spaces](https://huggingface.co/spaces/krishpatel/autonomous-coding-agent)**18 19---20 21## πΈ Demo22 2324 25---26 27## π₯ What makes this different from just using ChatGPT?28 29| Feature | ChatGPT / Basic Agent | This Agent |30|---|---|---|31| Code generation | β
| β
|32| Syntax validation | β Run and hope | β
AST parse before running |33| Test cases | β Manual | β
Auto-generated by agent |34| Stress testing | β | β
500+ random inputs via Hypothesis |35| Memory | β Stateless | β
ChromaDB learns from past bugs |36| Security audit | β | β
Detects eval, exec, hardcoded keys |37| Performance check | β | β
Benchmarks 1000 runs, rejects slow code |38| Self-review | β | β
Agent scores own confidence 1-10 |39| Self-healing | β | β
Loops back and fixes failures automatically |40| Separate retry counters | β | β
Per-node counters prevent pipeline blockage |41 42---43 44## π Key Metrics45 46| Metric | Value |47|---|---|48| Pipeline nodes | 13 |49| Verification layers | 5 (AST β Tests β Hypothesis β Security β Complexity) |50| Max retries (debugger) | 3 |51| Max retries (security, complexity) | 2 each β independent counters |52| Hypothesis test cases | 500+ random inputs per run |53| Benchmark iterations | 1,000 runs |54| Performance threshold | < 5ms per call |55| Memory backend | ChromaDB vector similarity search |56| LLM | Llama 3.1 8B Instant via Groq |57| Avg pipeline runtime | ~20β40 seconds |58| Lines of code | ~600 across 5 files |59 60---61 62## ποΈ Architecture β 13-Node Pipeline63 64```65User Input (Python Task)66 β67 βΌ68 βββββββββββ69 β Planner β ββ Breaks task into blueprint70 ββββββ¬βββββ71 β72 βΌ73 βββββββββ74 β Coder β ββ Writes code using plan + ChromaDB memory75 ββββββ¬βββ76 β77 βΌ78 βββββββββββββββββ79 β AST Validator β ββ Syntax + hallucinated imports + type hints80 ββββββββ¬βββββββββ (no execution needed β milliseconds)81 β82 Pass β Fail βββΊ Debugger βββΊ back to AST83 βΌ84ββββββββββββββββββ85β Test Generator β ββ Auto-generates pytest-style test cases86βββββββββ¬βββββββββ87 β88 βΌ89 ββββββββββ90 β Tester β ββ Runs code + generated tests in sandbox91 βββββ¬βββββ92 β93 Pass β Fail βββΊ Debugger (max 3 retries)94 βΌ95ββββββββββββββ96β Hypothesis β ββ 500+ random inputs, property-based testing97βββββββ¬βββββββ (never blocks pipeline β informational only)98 β99 βΌ100βββββββββββββ101β Benchmark β ββ Runs 1000x, rejects if > 5ms/call102βββββββ¬ββββββ103 β104 βΌ105ββββββββββββ106β Security β ββ Detects eval/exec/hardcoded secrets107βββββββ¬βββββ (own retry counter β max 2)108 β109 βΌ110ββββββββββββββ111β Complexity β ββ Line count + nesting depth + LLM score/10112ββββββββ¬ββββββ (own retry counter β max 2)113 β114 βΌ115βββββββββββββββββββ116β Self Reflection β ββ Agent scores own confidence 1-10117ββββββββββ¬βββββββββ Rewrites if confidence < 7118 β119 βΌ120 ββββββββββββ121 β Reviewer β ββ Polishes + docstrings + type hints122 βββββββ¬βββββ123 β124 βΌ125 ββββββββββββ126 βExplainer β ββ Writes human-readable explanation127 βββββββ¬βββββ128 β129 βΌ130 OUTPUT131 Final Code + Explanation132```133 134---135 136## π Project Structure137 138```139autonomous-coding-agent/140βββ app.py β Streamlit UI141βββ main.py β Graph builder + entry point142βββ state.py β Shared TypedDict state (whiteboard)143βββ nodes.py β All 13 node functions + LLM + ChromaDB144βββ edges.py β All 7 conditional route functions145βββ requirements.txt β Dependencies146βββ README.md147```148 149---150 151## β‘ Run Locally152 153### Prerequisites154- Python 3.11+155- Groq API key β get free at [console.groq.com](https://console.groq.com)156 157### Step 1 β Clone the repo158```bash159git clone https://github.com/krishpatel/autonomous-coding-agent.git160cd autonomous-coding-agent161```162 163### Step 2 β Create virtual environment164```bash165python -m venv venv166 167# Mac/Linux168source venv/bin/activate169 170# Windows171venv\Scripts\activate172```173 174### Step 3 β Install dependencies175```bash176pip install -r requirements.txt177```178 179### Step 4 β Set your API key180```bash181# Mac/Linux182export GROQ_API_KEY=your_groq_api_key_here183 184# Windows185set GROQ_API_KEY=your_groq_api_key_here186```187 188Or create a `.env` file:189```bash190echo "GROQ_API_KEY=your_groq_api_key_here" > .env191```192 193### Step 5 β Run CLI (no UI)194```bash195python main.py196```197 198### Step 6 β Run Streamlit UI199```bash200streamlit run app.py201```202 203Open [http://localhost:8501](http://localhost:8501) in your browser.204 205---206 207## π³ Run with Docker (optional)208 209```dockerfile210# Dockerfile211FROM python:3.11-slim212WORKDIR /app213COPY requirements.txt .214RUN pip install -r requirements.txt215COPY . .216EXPOSE 8501217CMD ["streamlit", "run", "app.py", "--server.port=8501"]218```219 220```bash221# Build222docker build -t coding-agent .223 224# Run225docker run -e GROQ_API_KEY=your_key -p 8501:8501 coding-agent226```227 228---229 230## π Deploy to Hugging Face Spaces231 232```bash233# Install HF CLI234pip install huggingface_hub235 236# Login237huggingface-cli login238 239# Create space and push240huggingface-cli repo create autonomous-coding-agent --type space --space_sdk streamlit241git remote add hf https://huggingface.co/spaces/YOUR_USERNAME/autonomous-coding-agent242git push hf main243```244 245Then add your secret in HF Spaces Settings:246```247GROQ_API_KEY = your_key_here248```249 250---251 252## π οΈ Tech Stack253 254```255LangGraph β Stateful multi-agent graph orchestration256Groq API β LLM inference (Llama 3.1 8B Instant)257ChromaDB β Vector database for bug fix memory258Hypothesis β Property-based stress testing259Streamlit β Production UI260subprocess β Sandboxed isolated code execution261ast β Static code analysis without execution262hashlib β Deterministic ChromaDB IDs263importlib β Real-time import hallucination detection264```265 266---267 268## π‘ Key Engineering Decisions269 270### Why LangGraph over plain LangChain?271LangGraph handles **cyclic workflows** β when tests fail, the agent loops back through the debugger and restarts verification from AST. LangChain's linear chains can't do this cleanly.272 273### Why AST validation before running?274Running broken code wastes subprocess time. AST parsing catches syntax errors in **milliseconds** without execution β like a proofreader checking spelling before printing.275 276### Why Hypothesis for testing?277Hand-written tests only cover cases you think of. Hypothesis **auto-generates 500+ random inputs** and verifies properties that should always hold. Catches edge cases no human would write.278 279### Why separate retry counters per node?280One shared counter caused security failing 3 times to kill the entire pipeline before the debugger got its attempts. Separate counters for security and complexity mean each node fails independently without blocking others.281 282### Why hashlib instead of Python's hash()?283Python's `hash()` is **randomized every session** for security. Same error β different ChromaDB ID β agent can never retrieve past fixes. `hashlib.md5` is deterministic across all sessions.284 285### Why combined Reviewer + Explainer?286Two separate LLM calls for polishing and explaining wasted ~8 seconds. One combined call with structured output (`FINAL_CODE:` / `EXPLANATION:`) saves an entire API round trip.287 288---289 290## π Real Bugs Found and Fixed291 292**Bug 1 β False Positive in Tester**293`returncode == 0` doesn't mean the function was called. A file that only defines functions exits successfully but prints nothing. Fixed by checking `stdout` is not empty after successful run.294 295**Bug 2 β ChromaDB Hash Randomization**296Python's `hash()` is session-randomized. Same bug β different ID every run β memory retrieval never works. Fixed with `hashlib.md5().hexdigest()[:8]` for deterministic cross-session IDs.297 298**Bug 3 β Python 3.11 F-string Backslash**299Python 3.11 doesn't allow backslashes inside f-string expressions. Benchmark node embedded code inside f-strings. Fixed using string concatenation instead.300 301**Bug 4 β Shared Retry Counter**302One `retries` counter shared across all nodes caused security/complexity failures to consume the debugger's retry budget. Fixed by adding `security_retries` and `complexity_retries` as independent counters.303 304---305 306## π Environment Variables307 308| Variable | Required | Description |309|---|---|---|310| `GROQ_API_KEY` | β
Yes | Get free at console.groq.com |311| `GITHUB_TOKEN` | β No | Only needed for AutoReview AI project |312 313---314 315## π Resume Line316 317> **Autonomous Python Coding Agent** | LangGraph Β· Groq Β· ChromaDB Β· Streamlit318> Built a 13-node self-healing pipeline with 5-layer verification β AST validation, auto-generated tests, Hypothesis property testing (500+ random inputs), security audit, and self-reflection confidence scoring. ChromaDB vector memory enables cross-session bug fix learning. Deployed on Hugging Face Spaces.319 320---321 322## π¨βπ» Author323 324**Krish Patel** β AI Engineer 325[GitHub](https://github.com/krishpatel) Β· [LinkedIn](https://linkedin.com/in/krishpatel) Β· [Live Demo](https://huggingface.co/spaces/krishpatel/autonomous-coding-agent)326 327---328 329*Built as part of AI Engineer internship portfolio β Bangalore, 2026*330 