Team Ai
Apppublic

aady161103/reflection-debug-agent

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes
App README

<div align="center">

๐Ÿ”ฌ Code Therapy โ€” Reflection-Guided Debugging Agent

An OpenEnv RL environment where AI agents learn to debug code _and_ explain their reasoning.

![Live Demo](https://huggingface.co/spaces/aady161103/reflection-debug-agent) ![OpenEnv](https://github.com/meta-pytorch/OpenEnv) ![Python](https://python.org) ![Docker](Dockerfile)


Most debugging environments reward only code correctness. This one rewards **how the agent thinks** โ€” not just what it outputs.

</div>


๐Ÿ’ก The Core Idea

Traditional code-debugging RL environments define reward as a binary: tests pass or they don't. This discards all signal about the agent's reasoning process.

Code Therapy introduces process supervision for debugging. At every step, the agent must produce:

ComponentWhat the agent writesWhy it matters
Hypothesis"The bug exists because `float("")` throws `ValueError` on empty CSV cells"Forces root-cause analysis, not random edits
Action"Wrap the conversion in a try/except and skip malformed rows"Demands intentional, justified fixes
Expected Result"Tests 1-3 should now pass since empty rows are gracefully skipped"Requires predictive reasoning about outcomes

The reward function then combines both dimensions:

reward = 0.6 ร— code_correctness + 0.4 ร— reflection_quality

Reflection quality is scored by an LLM-as-a-Judge (Qwen2.5-72B-Instruct) using a structured rubric โ€” making this a process-supervised debugging environment.


๐Ÿงฉ Tasks โ€” 3 Difficulty Levels

Each task presents real-world buggy Python code with automated test suites. Rewards are in [0.0, 1.0].

Task 1 ยท api_json_fix ยท Easy

Bug: API handler fails to catch JSONDecodeError on malformed payloads and uses JavaScript-style dot notation (data.username) on Python dicts.
  • โ€”3 unit tests ยท Focuses on exception handling and dict access syntax.

Task 2 ยท csv_processor_fix ยท Medium

Bug: CSV sales processor crashes with ValueError on empty amount fields (float("")) and has missing store_id grouping logic.
  • โ€”3 unit tests (including edge cases) ยท Focuses on real-world data sanitization.

Task 3 ยท retry_decorator_fix ยท Hard

Bug: A @retry_on_exception decorator silently swallows exceptions when retries exhaust and loops an incorrect number of times.
  • โ€”4 unit tests tracking exact retry counts and exception propagation ยท Focuses on closures, decorators, and control flow.

๐Ÿ“ฆ OpenEnv Specification

Action Space โ€” DebugAction

python
class DebugAction(BaseModel):
    edits: List[CodeEdit]          # search-and-replace patches
    hypothesis: str                # why the bug exists
    action_description: str        # what was changed and why
    expected_result: str           # predicted outcome after fix

Observation Space โ€” DebugObservation

python
class DebugObservation(BaseModel):
    buggy_code: str                # current source code
    test_output: str               # stdout/stderr from test run
    tests_passed: int              # passing test count
    tests_total: int               # total test count
    reflection_prompt: str         # structured Hโ†’Aโ†’R prompt
    step_number: int               # current step (1-based)
    done: bool                     # episode complete?
    reward: Optional[float]        # step reward (null on reset)
    reward_breakdown: Optional[dict]  # detailed scoring components

State โ€” DebugState

python
class DebugState(BaseModel):
    episode_id: str
    step_count: int
    task_name: str
    max_steps: int                 # 8
    best_score: float

API Endpoints

MethodEndpointDescription
POST/resetReset env for a new task, returns initial DebugObservation
POST/stepSubmit a DebugAction, returns observation + reward + done
GET/stateQuery current environment state
GET/healthHealth check

๐Ÿ“Š Reward Function โ€” Dual-Axis Scoring

total_reward = 0.6 ร— code_correctness + 0.4 ร— reflection_quality

Code Correctness (weight: 0.6)

code_score = tests_passed / tests_total

Reflection Quality (weight: 0.4) โ€” LLM-as-a-Judge

An LLM evaluates the agent's structured reflection on four axes:

CriterionWeightWhat's evaluated
Bug Identification25%Did the agent correctly identify a plausible root cause?
Fix Relevance25%Does the proposed edit directly address the identified bug?
Reasoning Consistency25%Does the expected result logically follow from the action?
Improvement Signal25%Did test results actually improve after the fix?

๐Ÿ—๏ธ Project Structure

reflection-debug-agent/
โ”œโ”€โ”€ openenv.yaml                # OpenEnv manifest (spec v1)
โ”œโ”€โ”€ pyproject.toml              # Python package config + dependencies
โ”œโ”€โ”€ inference.py                # Baseline inference script (root)
โ”œโ”€โ”€ Dockerfile                  # Multi-stage build (Node + Python)
โ”œโ”€โ”€ docker-compose.yml          # Local development
โ”‚
โ”œโ”€โ”€ backend/
โ”‚   โ”œโ”€โ”€ main.py                 # FastAPI server โ€” /reset, /step, /state
โ”‚   โ”œโ”€โ”€ agent.py                # LLM agent wrapper
โ”‚   โ”œโ”€โ”€ requirements.txt        # Python dependencies
โ”‚   โ”œโ”€โ”€ models/                 # Pydantic typed models
โ”‚   โ”‚   โ”œโ”€โ”€ action.py           #   DebugAction + CodeEdit
โ”‚   โ”‚   โ”œโ”€โ”€ observation.py      #   DebugObservation
โ”‚   โ”‚   โ”œโ”€โ”€ state.py            #   DebugState
โ”‚   โ”‚   โ””โ”€โ”€ reward.py           #   RewardBreakdown
โ”‚   โ”œโ”€โ”€ tasks/                  # 3 difficulty-graded tasks
โ”‚   โ”‚   โ”œโ”€โ”€ base_task.py        #   Abstract task interface
โ”‚   โ”‚   โ”œโ”€โ”€ task_easy.py        #   api_json_fix
โ”‚   โ”‚   โ”œโ”€โ”€ task_medium.py      #   csv_processor_fix
โ”‚   โ”‚   โ”œโ”€โ”€ task_hard.py        #   retry_decorator_fix
โ”‚   โ”‚   โ””โ”€โ”€ custom_task.py      #   User-provided code support
โ”‚   โ””โ”€โ”€ engine/
โ”‚       โ”œโ”€โ”€ environment.py      #   OpenEnv environment loop
โ”‚       โ””โ”€โ”€ reflection_scorer.py #  LLM-as-a-Judge scorer
โ”‚
โ”œโ”€โ”€ server/
โ”‚   โ””โ”€โ”€ app.py                  # Multi-mode deployment entry point
โ”‚
โ””โ”€โ”€ frontend/                   # React + Vite dashboard
    โ””โ”€โ”€ src/
        โ”œโ”€โ”€ components/         # UI components
        โ”œโ”€โ”€ store/              # Zustand state management
        โ””โ”€โ”€ services/           # API client

๐Ÿš€ Quick Start

Environment Variables

bash
# Required
export API_BASE_URL="https://router.huggingface.co/v1"
export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
export HF_TOKEN="your_huggingface_token"

# Optional
export ENV_URL="http://localhost:7860"   # defaults to localhost

Option 1 โ€” Docker (Recommended)

bash
docker build -t reflection-debug-agent .
docker run -p 7860:7860 \
  -e HF_TOKEN=$HF_TOKEN \
  -e API_BASE_URL=$API_BASE_URL \
  -e MODEL_NAME=$MODEL_NAME \
  reflection-debug-agent

Option 2 โ€” Docker Compose

bash
# Create .env file with your variables first
docker compose up --build

Option 3 โ€” Local Development

bash
# Backend
cd backend && pip install -r requirements.txt && cd ..
uvicorn backend.main:app --host 0.0.0.0 --port 7860 --reload

# Frontend (separate terminal)
cd frontend && npm install && npm run dev

Running Inference

bash
# Start the environment server first, then:
python inference.py

Expected stdout format:

[START] task=api_json_fix env=reflection_debug_agent model=Qwen/Qwen2.5-72B-Instruct
[STEP]  step=1 action=fix(Incorrect dict access ...) reward=0.45 done=false error=null
[STEP]  step=2 action=fix(Missing JSONDecodeError...) reward=0.78 done=false error=null
[STEP]  step=3 action=fix(Final cleanup of error ...) reward=1.00 done=true error=null
[END]   success=true steps=3 rewards=0.45,0.78,1.00

โœ… Pre-Submission Checklist

  • โ€”[x] HF Space deploys โ€” automated ping returns 200 and responds to POST /reset
  • โ€”[x] OpenEnv spec compliance โ€” openenv.yaml + typed Pydantic models + /step /reset /state
  • โ€”[x] Dockerfile builds โ€” multi-stage (Node.js frontend โ†’ Python backend)
  • โ€”[x] Baseline inference reproduces โ€” inference.py at root with [START]/[STEP]/[END] stdout
  • โ€”[x] 3+ tasks with graders โ€” easy/medium/hard, all scores in [0.0, 1.0]
  • โ€”[x] OpenAI client โ€” all LLM calls via openai.OpenAI(base_url=API_BASE_URL, api_key=HF_TOKEN)
  • โ€”[x] pyproject.toml โ€” with openenv-core>=0.2.0 dependency and [project.scripts] entry
  • โ€”[x] uv.lock โ€” dependency lock file present

๐Ÿ”— Links


๐Ÿ“„ License

MIT