Team Ai
Apppublic

Krishp1/Autonomous-Coding-Agent

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes
README-3.md330 linesDownload Raw Back to root
1# πŸ€– Autonomous Python Coding Agent2 3> **A production-grade, self-healing multi-agent pipeline that doesn't just generate Python code β€” it autonomously writes, validates, tests, secures, benchmarks, and reflects on its own output before shipping.**4 5[![Python](https://img.shields.io/badge/Python-3.11-blue?style=flat-square&logo=python)](https://python.org)6[![LangGraph](https://img.shields.io/badge/LangGraph-0.2.0-green?style=flat-square)](https://github.com/langchain-ai/langgraph)7[![Groq](https://img.shields.io/badge/Groq-Llama%203.1-orange?style=flat-square)](https://groq.com)8[![ChromaDB](https://img.shields.io/badge/ChromaDB-0.5.0-purple?style=flat-square)](https://chromadb.com)9[![Streamlit](https://img.shields.io/badge/Streamlit-1.35-red?style=flat-square)](https://streamlit.io)10[![License](https://img.shields.io/badge/License-MIT-lightgrey?style=flat-square)](LICENSE)11[![Live Demo](https://img.shields.io/badge/πŸ€—%20Live%20Demo-HuggingFace-yellow?style=flat-square)](https://huggingface.co/spaces/krishpatel/autonomous-coding-agent)12 13---14 15## πŸš€ Live Demo16 17**[β–Ά Try it on Hugging Face Spaces](https://huggingface.co/spaces/krishpatel/autonomous-coding-agent)**18 19---20 21## πŸ“Έ Demo22 23![Agent Demo](demo.gif)24 25---26 27## πŸ”₯ What makes this different from just using ChatGPT?28 29| Feature | ChatGPT / Basic Agent | This Agent |30|---|---|---|31| Code generation | βœ… | βœ… |32| Syntax validation | ❌ Run and hope | βœ… AST parse before running |33| Test cases | ❌ Manual | βœ… Auto-generated by agent |34| Stress testing | ❌ | βœ… 500+ random inputs via Hypothesis |35| Memory | ❌ Stateless | βœ… ChromaDB learns from past bugs |36| Security audit | ❌ | βœ… Detects eval, exec, hardcoded keys |37| Performance check | ❌ | βœ… Benchmarks 1000 runs, rejects slow code |38| Self-review | ❌ | βœ… Agent scores own confidence 1-10 |39| Self-healing | ❌ | βœ… Loops back and fixes failures automatically |40| Separate retry counters | ❌ | βœ… Per-node counters prevent pipeline blockage |41 42---43 44## πŸ“Š Key Metrics45 46| Metric | Value |47|---|---|48| Pipeline nodes | 13 |49| Verification layers | 5 (AST β†’ Tests β†’ Hypothesis β†’ Security β†’ Complexity) |50| Max retries (debugger) | 3 |51| Max retries (security, complexity) | 2 each β€” independent counters |52| Hypothesis test cases | 500+ random inputs per run |53| Benchmark iterations | 1,000 runs |54| Performance threshold | < 5ms per call |55| Memory backend | ChromaDB vector similarity search |56| LLM | Llama 3.1 8B Instant via Groq |57| Avg pipeline runtime | ~20–40 seconds |58| Lines of code | ~600 across 5 files |59 60---61 62## πŸ—οΈ Architecture β€” 13-Node Pipeline63 64```65User Input (Python Task)66         β”‚67         β–Ό68    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”69    β”‚ Planner β”‚ ── Breaks task into blueprint70    β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜71         β”‚72         β–Ό73    β”Œβ”€β”€β”€β”€β”€β”€β”€β”74    β”‚ Coder β”‚ ── Writes code using plan + ChromaDB memory75    β””β”€β”€β”€β”€β”¬β”€β”€β”˜76         β”‚77         β–Ό78    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”79    β”‚ AST Validator β”‚ ── Syntax + hallucinated imports + type hints80    β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜    (no execution needed β€” milliseconds)81           β”‚82      Pass β”‚   Fail ──► Debugger ──► back to AST83           β–Ό84β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”85β”‚ Test Generator β”‚ ── Auto-generates pytest-style test cases86β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜87        β”‚88        β–Ό89    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”90    β”‚ Tester β”‚ ── Runs code + generated tests in sandbox91    β””β”€β”€β”€β”¬β”€β”€β”€β”€β”˜92        β”‚93   Pass β”‚   Fail ──► Debugger (max 3 retries)94        β–Ό95β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”96β”‚ Hypothesis β”‚ ── 500+ random inputs, property-based testing97β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜    (never blocks pipeline β€” informational only)98      β”‚99      β–Ό100β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”101β”‚ Benchmark β”‚ ── Runs 1000x, rejects if > 5ms/call102β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜103      β”‚104      β–Ό105β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”106β”‚ Security β”‚ ── Detects eval/exec/hardcoded secrets107β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜    (own retry counter β€” max 2)108      β”‚109      β–Ό110β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”111β”‚ Complexity β”‚ ── Line count + nesting depth + LLM score/10112β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜    (own retry counter β€” max 2)113       β”‚114       β–Ό115β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”116β”‚ Self Reflection β”‚ ── Agent scores own confidence 1-10117β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜    Rewrites if confidence < 7118         β”‚119         β–Ό120    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”121    β”‚ Reviewer β”‚ ── Polishes + docstrings + type hints122    β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜123          β”‚124          β–Ό125    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”126    β”‚Explainer β”‚ ── Writes human-readable explanation127    β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜128          β”‚129          β–Ό130       OUTPUT131  Final Code + Explanation132```133 134---135 136## πŸ“ Project Structure137 138```139autonomous-coding-agent/140β”œβ”€β”€ app.py              ← Streamlit UI141β”œβ”€β”€ main.py             ← Graph builder + entry point142β”œβ”€β”€ state.py            ← Shared TypedDict state (whiteboard)143β”œβ”€β”€ nodes.py            ← All 13 node functions + LLM + ChromaDB144β”œβ”€β”€ edges.py            ← All 7 conditional route functions145β”œβ”€β”€ requirements.txt    ← Dependencies146└── README.md147```148 149---150 151## ⚑ Run Locally152 153### Prerequisites154- Python 3.11+155- Groq API key β€” get free at [console.groq.com](https://console.groq.com)156 157### Step 1 β€” Clone the repo158```bash159git clone https://github.com/krishpatel/autonomous-coding-agent.git160cd autonomous-coding-agent161```162 163### Step 2 β€” Create virtual environment164```bash165python -m venv venv166 167# Mac/Linux168source venv/bin/activate169 170# Windows171venv\Scripts\activate172```173 174### Step 3 β€” Install dependencies175```bash176pip install -r requirements.txt177```178 179### Step 4 β€” Set your API key180```bash181# Mac/Linux182export GROQ_API_KEY=your_groq_api_key_here183 184# Windows185set GROQ_API_KEY=your_groq_api_key_here186```187 188Or create a `.env` file:189```bash190echo "GROQ_API_KEY=your_groq_api_key_here" > .env191```192 193### Step 5 β€” Run CLI (no UI)194```bash195python main.py196```197 198### Step 6 β€” Run Streamlit UI199```bash200streamlit run app.py201```202 203Open [http://localhost:8501](http://localhost:8501) in your browser.204 205---206 207## 🐳 Run with Docker (optional)208 209```dockerfile210# Dockerfile211FROM python:3.11-slim212WORKDIR /app213COPY requirements.txt .214RUN pip install -r requirements.txt215COPY . .216EXPOSE 8501217CMD ["streamlit", "run", "app.py", "--server.port=8501"]218```219 220```bash221# Build222docker build -t coding-agent .223 224# Run225docker run -e GROQ_API_KEY=your_key -p 8501:8501 coding-agent226```227 228---229 230## 🌐 Deploy to Hugging Face Spaces231 232```bash233# Install HF CLI234pip install huggingface_hub235 236# Login237huggingface-cli login238 239# Create space and push240huggingface-cli repo create autonomous-coding-agent --type space --space_sdk streamlit241git remote add hf https://huggingface.co/spaces/YOUR_USERNAME/autonomous-coding-agent242git push hf main243```244 245Then add your secret in HF Spaces Settings:246```247GROQ_API_KEY = your_key_here248```249 250---251 252## πŸ› οΈ Tech Stack253 254```255LangGraph    β€” Stateful multi-agent graph orchestration256Groq API     β€” LLM inference (Llama 3.1 8B Instant)257ChromaDB     β€” Vector database for bug fix memory258Hypothesis   β€” Property-based stress testing259Streamlit    β€” Production UI260subprocess   β€” Sandboxed isolated code execution261ast          β€” Static code analysis without execution262hashlib      β€” Deterministic ChromaDB IDs263importlib    β€” Real-time import hallucination detection264```265 266---267 268## πŸ’‘ Key Engineering Decisions269 270### Why LangGraph over plain LangChain?271LangGraph handles **cyclic workflows** β€” when tests fail, the agent loops back through the debugger and restarts verification from AST. LangChain's linear chains can't do this cleanly.272 273### Why AST validation before running?274Running broken code wastes subprocess time. AST parsing catches syntax errors in **milliseconds** without execution β€” like a proofreader checking spelling before printing.275 276### Why Hypothesis for testing?277Hand-written tests only cover cases you think of. Hypothesis **auto-generates 500+ random inputs** and verifies properties that should always hold. Catches edge cases no human would write.278 279### Why separate retry counters per node?280One shared counter caused security failing 3 times to kill the entire pipeline before the debugger got its attempts. Separate counters for security and complexity mean each node fails independently without blocking others.281 282### Why hashlib instead of Python's hash()?283Python's `hash()` is **randomized every session** for security. Same error β†’ different ChromaDB ID β†’ agent can never retrieve past fixes. `hashlib.md5` is deterministic across all sessions.284 285### Why combined Reviewer + Explainer?286Two separate LLM calls for polishing and explaining wasted ~8 seconds. One combined call with structured output (`FINAL_CODE:` / `EXPLANATION:`) saves an entire API round trip.287 288---289 290## πŸ› Real Bugs Found and Fixed291 292**Bug 1 β€” False Positive in Tester**293`returncode == 0` doesn't mean the function was called. A file that only defines functions exits successfully but prints nothing. Fixed by checking `stdout` is not empty after successful run.294 295**Bug 2 β€” ChromaDB Hash Randomization**296Python's `hash()` is session-randomized. Same bug β†’ different ID every run β†’ memory retrieval never works. Fixed with `hashlib.md5().hexdigest()[:8]` for deterministic cross-session IDs.297 298**Bug 3 β€” Python 3.11 F-string Backslash**299Python 3.11 doesn't allow backslashes inside f-string expressions. Benchmark node embedded code inside f-strings. Fixed using string concatenation instead.300 301**Bug 4 β€” Shared Retry Counter**302One `retries` counter shared across all nodes caused security/complexity failures to consume the debugger's retry budget. Fixed by adding `security_retries` and `complexity_retries` as independent counters.303 304---305 306## πŸ”‘ Environment Variables307 308| Variable | Required | Description |309|---|---|---|310| `GROQ_API_KEY` | βœ… Yes | Get free at console.groq.com |311| `GITHUB_TOKEN` | ❌ No | Only needed for AutoReview AI project |312 313---314 315## πŸ“ Resume Line316 317> **Autonomous Python Coding Agent** | LangGraph Β· Groq Β· ChromaDB Β· Streamlit318> Built a 13-node self-healing pipeline with 5-layer verification β€” AST validation, auto-generated tests, Hypothesis property testing (500+ random inputs), security audit, and self-reflection confidence scoring. ChromaDB vector memory enables cross-session bug fix learning. Deployed on Hugging Face Spaces.319 320---321 322## πŸ‘¨β€πŸ’» Author323 324**Krish Patel** β€” AI Engineer  325[GitHub](https://github.com/krishpatel) Β· [LinkedIn](https://linkedin.com/in/krishpatel) Β· [Live Demo](https://huggingface.co/spaces/krishpatel/autonomous-coding-agent)326 327---328 329*Built as part of AI Engineer internship portfolio β€” Bangalore, 2026*330