harshit078/code-review-agent1
Code Review Agent — OpenEnv Environment
A reinforcement learning environment where an AI agent learns to review Python code across three difficulty levels.
Real-World Problem
Millions of code reviews happen daily at tech companies. Junior developers miss syntax errors, security vulnerabilities, and style issues. This environment trains an AI agent to perform accurate, structured code reviews — reducing review time and improving code quality.
Environment Overview
How It Works
The agent interacts with the environment through 3 API endpoints:
POST /reset— loads a random code sample and returns it as observationPOST /step— accepts the agent's text review and returns a reward scoreGET /state— returns current task info and history
Tasks
Task 1 — Easy: Syntax Error Detection
Agent must identify the exact error type (SyntaxError, IndentationError, etc.) and the line number where it occurs.
Reward tiers:
- 1.0 — correct error type + correct line number
- 0.6 — correct error type, wrong line number
- 0.3 — related keywords mentioned
- 0.0 — missed completely
Task 2 — Medium: Security Vulnerability Detection
Agent must identify security vulnerabilities — SQL injection, hardcoded credentials, eval injection, path traversal, command injection, and more.
Reward tiers:
- 1.0 — 4+ vulnerability keywords matched
- 0.75 — 3 keywords matched
- 0.5 — 2 keywords matched
- 0.25 — 1 keyword matched
- 0.0 — missed completely
Task 3 — Hard: Full Code Quality Review
Agent must cover all 4 review categories. Each category is worth 0.25.
Observation Space
{
"code": "def login(user, pwd):\n ...",
"task_type": "easy",
"difficulty": "easy",
"language": "python",
"instructions": "Review this code and identify all syntax errors..."
}Action Space
Natural language text — the agent's code review response.
Reward Design
Partial scoring is used throughout — agents are never penalized with binary right/wrong. Every partially correct answer earns credit, which enables meaningful gradient signal for RL training.
Running Locally
# Install dependencies
pip install -r requirements.txt
# Start the environment server
uvicorn app:app --reload --port 8000
# Run the baseline agent
python inference.pyRunning with Docker
docker build -t code-review-agent .
docker run -p 7860:7860 \
-e API_BASE_URL=https://api.openai.com/v1 \
-e MODEL_NAME=gpt-4o-mini \
-e HF_TOKEN=your_key_here \
code-review-agentEnvironment Variables
Project Structure
code-review-agent/
├── code_samples/
│ ├── easy_samples.json # 10 syntax error samples
│ ├── medium_samples.json # 10 security vulnerability samples
│ └── hard_samples.json # 5 multi-issue samples
├── graders/
│ ├── easy_grader.py # Keyword + line number scoring
│ ├── medium_grader.py # Vulnerability keyword scoring
│ └── hard_grader.py # 4-category partial scoring
├── environment.py # Core OpenEnv class
├── app.py # FastAPI server
├── inference.py # Baseline AI agent
├── openenv.yaml # OpenEnv spec config
├── Dockerfile # HF Spaces deployment
└── requirements.txt