PRANAV05092003/autonomous-code-refactoring-env
๐ ACRE โ Autonomous Code Refactoring Environment
OpenEnv-powered AI system for real-world code optimization, refactoring, and evaluation.
๐ฅ Overview
ACRE is an OpenEnv-compliant environment designed to simulate real-world software engineering workflows such as code cleanup, optimization, and refactoring using AI agents.
It enables agents to iteratively improve code through structured actions while receiving dense, step-wise reward feedback.
Environment Overview and Motivation
ACRE models a realistic developer workflow where an agent incrementally improves Python code quality under a fixed action budget. The environment is designed for OpenEnv Round 1 requirements: typed APIs, deterministic grading, multi-difficulty tasks, and reproducible inference behavior.
๐ก Why This Matters
Modern software systems require automated code optimization and intelligent tooling.
ACRE enables:
- ๐ค AI coding assistants
- ๐ Automated code review systems
- โก Reinforcement learning-based optimization agents
- ๐ง Learning real developer workflows
๐ How It Works
Code โ Action โ Refactor โ Reward โ Repeat
- Load messy code
- Apply transformation
- Evaluate using grader
- Compute reward
- Iterate until optimal
๐ง Key Features
- โ Autonomous code refactoring
- โก Step-wise reward feedback
- ๐งช OpenEnv compliant interface
- ๐ Deterministic grading system
- ๐ Reproducible inference pipeline
- ๐ณ Fully containerized (Docker + Hugging Face Spaces)
๐ Tasks
Each task uses AST-based transformations and deterministic grading.
Task Descriptions with Expected Difficulty Levels
- Easy (
rename_variables): rename generic names likex,tmp,iinto descriptive identifiers. - Medium (
remove_dead_code): remove unreachable branches and unused assignments while preserving behavior. - Hard (
full_refactor): combine renaming, dead-code elimination, loop simplification, condition cleanup, and helper inlining.
๐ฏ Reward System
Rewards are computed at every step:
- โ Valid executable code โ positive reward
- ๐ Reduced complexity โ reward
- โก Improved performance โ reward
- โ Errors or invalid code โ penalty
- ๐ No progress โ penalty
Normalization:
(raw_reward + 32) / 52 โ [0, 1]
๐ Example Execution
[START] task=rename_variables
[STEP] action=0
[END] task=rename_variables score=1.00
[START] task=remove_dead_code
[STEP] action=1
[END] task=remove_dead_code score=0.25
[START] task=full_refactor
[STEP] action=3
[END] task=full_refactor score=0.71
Final Score: 0.65๐๏ธ Architecture
server/app.pyโ FastAPI entry point used by OpenEnv + Dockerserver.pyโ legacy local runner / UI helperopenenv_interface.pyโ OpenEnv wrapperacre/env/โ Core environment logicacre/tasks/โ Task definitionsacre/utils/โ Metrics and helpersinference.pyโ Evaluation pipeline
โ๏ธ OpenEnv Interface
observation = env.reset()
observation, reward, done, info = env.step(action)
state = env.state()Uses Pydantic models:
ObservationModelActionModelRewardModel
Definitions of Action and Observation Spaces
- Observation space: Box(4) with fields
code_length,complexity_score,runtime_s,error_flag. - Action space: Discrete(5) with actions
rename_variable,remove_dead_code,simplify_loop,optimize_condition,inline_function.
๐ HTTP API
๐ Run Locally
Setup and Usage Instructions
pip install -r requirements.txt
uvicorn server.app:app --host 0.0.0.0 --port 7860๐ณ Docker / Hugging Face Spaces
docker build -t acre .
docker run -p 7860:7860 \
-e API_BASE_URL=https://api.openai.com/v1 \
-e MODEL_NAME=gpt-4o-mini \
-e API_KEY=your_key \
-e ENV_URL=http://localhost:7860 \
acre๐งช Inference
Set environment variables:
export API_BASE_URL=https://api.openai.com/v1
export MODEL_NAME=gpt-4o-mini
export API_KEY=your_key
export ENV_URL=http://localhost:7860Run:
python inference.pyExpected output:
Easy: 1.00
Medium: 0.25
Hard: 0.71
Final: 0.65๐ OpenEnv Compliance
- โ
step()implemented - โ
reset()implemented - โ
state()implemented - โ reward shaping
- โ deterministic grading
- โ structured logs
๐งช Validation
python validate.py --url http://localhost:7860Or:
openenv validate๐ Live Demo
๐ Running on Hugging Face Spaces
๐ Baseline Performance
Baseline Performance Scores
๐ Use Cases
- AI-powered code optimization
- Automated refactoring tools
- Reinforcement learning environments
- Developer productivity systems
๐ License
MIT License
