Team Ai
Apppublic

rakshithakho/bugfixerenv

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

๐Ÿ› BugFixerEnv

An OpenEnv-compatible reinforcement-learning environment where an agent reads a buggy Python function and submits a corrected version. Reward = fraction of test cases passed (0.0 โ€“ 1.0).

Table of Contents


Environment Description

BugFixerEnv simulates real-world code debugging. At each episode:

  1. 1.The environment presents a buggy Python function along with a plain-English description of what it should do.
  2. 2.The agent submits fixed_code containing the corrected function.
  3. 3.The environment executes the submission against hidden test cases and returns a reward equal to the fraction passed.
  4. 4.The episode ends when the agent achieves a perfect score (1.0) or exhausts max_steps (5) attempts.

This mirrors the daily workflow of software engineers who must identify and fix bugs in unfamiliar code โ€” a high-value, real-world task.


Observation Space

json
{
  "task_id":     "easy | medium | hard",
  "description": "What the function should do (natural language)",
  "buggy_code":  "The broken Python function",
  "steps_taken": 0,
  "max_steps":   5,
  "last_reward": 0.0,
  "done":        false
}

Action Space

json
{
  "fixed_code": "<corrected Python source code as a string>"
}

The submitted string must contain a valid Python function with the same name as the original.


Tasks

IDDescriptionBugsBaseline Score
easysum_to_n(n) โ€” sum integers 1โ€ฆn1 (off-by-one in range)1.0
mediummax_subarray_sum(nums) โ€” Kadane's algorithm1 (- instead of +)1.0
hardis_valid_parentheses(s) โ€” bracket matching2 (wrong return values)1.0

Difficulty progression

  • โ€”Easy โ€” Single character change. All 4 test cases fail with buggy code.
  • โ€”Medium โ€” Subtle algorithmic error. 2 of 5 tests pass with bug (partial reward = 0.40).
  • โ€”Hard โ€” Two independent bugs. Fixing one yields partial reward (0.29โ€“0.71), both needed for 1.0.

Reward Function

reward = tests_passed / total_tests   # range: 0.0 โ€“ 1.0

Partial progress is always rewarded. A submission passing 3 of 5 tests receives 0.60, not 0. This creates a meaningful gradient for learning agents.


API Reference

MethodPathDescription
GET/Health check
POST/resetStart a new episode
POST/stepSubmit a fix attempt
GET/statePeek at current state

Setup & Running

bash
git clone https://github.com/rakshithakh/bugfixerenv
cd bugfixerenv
pip install -r requirements.txt
uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Browse docs at http://localhost:8000/docs.


Docker

bash
docker build -t bugfixerenv .
docker run -p 7860:7860 bugfixerenv

Baseline Inference

bash
export API_BASE_URL=https://rakshithakho-bugfixerenv.hf.space
export MODEL_NAME=meta-llama/Llama-3.3-70B-Instruct
export HF_TOKEN=hf_xxx

python inference.py

Baseline Scores

TaskScore
easy1.00
medium1.00
hard1.00
average1.00

Reproducible Scores

All test cases are fully deterministic. temperature=0.0 ensures reproducible LLM outputs. Same input always produces same score.


Project Structure

bugfixerenv/
โ”œโ”€โ”€ main.py          # FastAPI app
โ”œโ”€โ”€ env.py           # Core environment logic
โ”œโ”€โ”€ grader.py        # Scoring / test execution
โ”œโ”€โ”€ tasks.py         # Task definitions
โ”œโ”€โ”€ inference.py     # Baseline agent
โ”œโ”€โ”€ openenv.yaml     # OpenEnv spec
โ”œโ”€โ”€ pyproject.toml   # Project metadata
โ”œโ”€โ”€ uv.lock          # Locked dependencies
โ”œโ”€โ”€ requirements.txt
โ”œโ”€โ”€ Dockerfile
โ”œโ”€โ”€ README.md
โ””โ”€โ”€ server/
    โ””โ”€โ”€ app.py       # Server entry point