rakshithakho/bugfixerenv
0
๐ BugFixerEnv
An OpenEnv-compatible reinforcement-learning environment where an agent reads a buggy Python function and submits a corrected version. Reward = fraction of test cases passed (0.0 โ 1.0).
Table of Contents
- Environment Description
- Observation Space
- Action Space
- Tasks
- Reward Function
- API Reference
- Setup & Running
- Docker
- Baseline Inference
- Reproducible Scores
Environment Description
BugFixerEnv simulates real-world code debugging. At each episode:
- The environment presents a buggy Python function along with a plain-English description of what it should do.
- The agent submits
fixed_codecontaining the corrected function. - The environment executes the submission against hidden test cases and returns a reward equal to the fraction passed.
- The episode ends when the agent achieves a perfect score (1.0) or exhausts
max_steps(5) attempts.
This mirrors the daily workflow of software engineers who must identify and fix bugs in unfamiliar code โ a high-value, real-world task.
Observation Space
{
"task_id": "easy | medium | hard",
"description": "What the function should do (natural language)",
"buggy_code": "The broken Python function",
"steps_taken": 0,
"max_steps": 5,
"last_reward": 0.0,
"done": false
}Action Space
{
"fixed_code": "<corrected Python source code as a string>"
}The submitted string must contain a valid Python function with the same name as the original.
Tasks
Difficulty progression
- Easy โ Single character change. All 4 test cases fail with buggy code.
- Medium โ Subtle algorithmic error. 2 of 5 tests pass with bug (partial reward = 0.40).
- Hard โ Two independent bugs. Fixing one yields partial reward (0.29โ0.71), both needed for 1.0.
Reward Function
reward = tests_passed / total_tests # range: 0.0 โ 1.0Partial progress is always rewarded. A submission passing 3 of 5 tests receives 0.60, not 0. This creates a meaningful gradient for learning agents.
API Reference
Setup & Running
git clone https://github.com/rakshithakh/bugfixerenv
cd bugfixerenv
pip install -r requirements.txt
uvicorn main:app --host 0.0.0.0 --port 8000 --reloadBrowse docs at http://localhost:8000/docs.
Docker
docker build -t bugfixerenv .
docker run -p 7860:7860 bugfixerenvBaseline Inference
export API_BASE_URL=https://rakshithakho-bugfixerenv.hf.space
export MODEL_NAME=meta-llama/Llama-3.3-70B-Instruct
export HF_TOKEN=hf_xxx
python inference.pyBaseline Scores
Reproducible Scores
All test cases are fully deterministic. temperature=0.0 ensures reproducible LLM outputs. Same input always produces same score.
Project Structure
bugfixerenv/
โโโ main.py # FastAPI app
โโโ env.py # Core environment logic
โโโ grader.py # Scoring / test execution
โโโ tasks.py # Task definitions
โโโ inference.py # Baseline agent
โโโ openenv.yaml # OpenEnv spec
โโโ pyproject.toml # Project metadata
โโโ uv.lock # Locked dependencies
โโโ requirements.txt
โโโ Dockerfile
โโโ README.md
โโโ server/
โโโ app.py # Server entry point