aady161103/reflection-debug-agent
<div align="center">
๐ฌ Code Therapy โ Reflection-Guided Debugging Agent
An OpenEnv RL environment where AI agents learn to debug code _and_ explain their reasoning.
   
Most debugging environments reward only code correctness. This one rewards **how the agent thinks** โ not just what it outputs.
</div>
๐ก The Core Idea
Traditional code-debugging RL environments define reward as a binary: tests pass or they don't. This discards all signal about the agent's reasoning process.
Code Therapy introduces process supervision for debugging. At every step, the agent must produce:
The reward function then combines both dimensions:
reward = 0.6 ร code_correctness + 0.4 ร reflection_qualityReflection quality is scored by an LLM-as-a-Judge (Qwen2.5-72B-Instruct) using a structured rubric โ making this a process-supervised debugging environment.
๐งฉ Tasks โ 3 Difficulty Levels
Each task presents real-world buggy Python code with automated test suites. Rewards are in [0.0, 1.0].
Task 1 ยท api_json_fix ยท Easy
Bug: API handler fails to catchJSONDecodeErroron malformed payloads and uses JavaScript-style dot notation (data.username) on Python dicts.
- 3 unit tests ยท Focuses on exception handling and dict access syntax.
Task 2 ยท csv_processor_fix ยท Medium
Bug: CSV sales processor crashes withValueErroron empty amount fields (float("")) and has missingstore_idgrouping logic.
- 3 unit tests (including edge cases) ยท Focuses on real-world data sanitization.
Task 3 ยท retry_decorator_fix ยท Hard
Bug: A @retry_on_exception decorator silently swallows exceptions when retries exhaust and loops an incorrect number of times.- 4 unit tests tracking exact retry counts and exception propagation ยท Focuses on closures, decorators, and control flow.
๐ฆ OpenEnv Specification
Action Space โ DebugAction
class DebugAction(BaseModel):
edits: List[CodeEdit] # search-and-replace patches
hypothesis: str # why the bug exists
action_description: str # what was changed and why
expected_result: str # predicted outcome after fixObservation Space โ DebugObservation
class DebugObservation(BaseModel):
buggy_code: str # current source code
test_output: str # stdout/stderr from test run
tests_passed: int # passing test count
tests_total: int # total test count
reflection_prompt: str # structured HโAโR prompt
step_number: int # current step (1-based)
done: bool # episode complete?
reward: Optional[float] # step reward (null on reset)
reward_breakdown: Optional[dict] # detailed scoring componentsState โ DebugState
class DebugState(BaseModel):
episode_id: str
step_count: int
task_name: str
max_steps: int # 8
best_score: floatAPI Endpoints
๐ Reward Function โ Dual-Axis Scoring
total_reward = 0.6 ร code_correctness + 0.4 ร reflection_qualityCode Correctness (weight: 0.6)
code_score = tests_passed / tests_totalReflection Quality (weight: 0.4) โ LLM-as-a-Judge
An LLM evaluates the agent's structured reflection on four axes:
๐๏ธ Project Structure
reflection-debug-agent/
โโโ openenv.yaml # OpenEnv manifest (spec v1)
โโโ pyproject.toml # Python package config + dependencies
โโโ inference.py # Baseline inference script (root)
โโโ Dockerfile # Multi-stage build (Node + Python)
โโโ docker-compose.yml # Local development
โ
โโโ backend/
โ โโโ main.py # FastAPI server โ /reset, /step, /state
โ โโโ agent.py # LLM agent wrapper
โ โโโ requirements.txt # Python dependencies
โ โโโ models/ # Pydantic typed models
โ โ โโโ action.py # DebugAction + CodeEdit
โ โ โโโ observation.py # DebugObservation
โ โ โโโ state.py # DebugState
โ โ โโโ reward.py # RewardBreakdown
โ โโโ tasks/ # 3 difficulty-graded tasks
โ โ โโโ base_task.py # Abstract task interface
โ โ โโโ task_easy.py # api_json_fix
โ โ โโโ task_medium.py # csv_processor_fix
โ โ โโโ task_hard.py # retry_decorator_fix
โ โ โโโ custom_task.py # User-provided code support
โ โโโ engine/
โ โโโ environment.py # OpenEnv environment loop
โ โโโ reflection_scorer.py # LLM-as-a-Judge scorer
โ
โโโ server/
โ โโโ app.py # Multi-mode deployment entry point
โ
โโโ frontend/ # React + Vite dashboard
โโโ src/
โโโ components/ # UI components
โโโ store/ # Zustand state management
โโโ services/ # API client๐ Quick Start
Environment Variables
# Required
export API_BASE_URL="https://router.huggingface.co/v1"
export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
export HF_TOKEN="your_huggingface_token"
# Optional
export ENV_URL="http://localhost:7860" # defaults to localhostOption 1 โ Docker (Recommended)
docker build -t reflection-debug-agent .
docker run -p 7860:7860 \
-e HF_TOKEN=$HF_TOKEN \
-e API_BASE_URL=$API_BASE_URL \
-e MODEL_NAME=$MODEL_NAME \
reflection-debug-agentOption 2 โ Docker Compose
# Create .env file with your variables first
docker compose up --buildOption 3 โ Local Development
# Backend
cd backend && pip install -r requirements.txt && cd ..
uvicorn backend.main:app --host 0.0.0.0 --port 7860 --reload
# Frontend (separate terminal)
cd frontend && npm install && npm run devRunning Inference
# Start the environment server first, then:
python inference.pyExpected stdout format:
[START] task=api_json_fix env=reflection_debug_agent model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=fix(Incorrect dict access ...) reward=0.45 done=false error=null
[STEP] step=2 action=fix(Missing JSONDecodeError...) reward=0.78 done=false error=null
[STEP] step=3 action=fix(Final cleanup of error ...) reward=1.00 done=true error=null
[END] success=true steps=3 rewards=0.45,0.78,1.00โ Pre-Submission Checklist
- [x] HF Space deploys โ automated ping returns 200 and responds to
POST /reset - [x] OpenEnv spec compliance โ
openenv.yaml+ typed Pydantic models +/step/reset/state - [x] Dockerfile builds โ multi-stage (Node.js frontend โ Python backend)
- [x] Baseline inference reproduces โ
inference.pyat root with[START]/[STEP]/[END]stdout - [x] 3+ tasks with graders โ easy/medium/hard, all scores in
[0.0, 1.0] - [x] OpenAI client โ all LLM calls via
openai.OpenAI(base_url=API_BASE_URL, api_key=HF_TOKEN) - [x] pyproject.toml โ with
openenv-core>=0.2.0dependency and[project.scripts]entry - [x] uv.lock โ dependency lock file present
๐ Links
๐ License
MIT
