kaysav18/code-review-assistant
Code Review Assistant — OpenEnv Environment
Submission-ready OpenEnv environment that challenges LLM agents to act as senior software engineers performing structured, multi-task code reviews.
1. Environment Description
The Code Review Assistant Environment presents an LLM agent with real Python code snippets and asks it to produce a structured review consisting of:
The environment evaluates responses deterministically using a weighted grading function and provides step-level reward feedback so that the agent can improve its review on subsequent attempts within the same task.
2. Real-World Motivation
Code review is one of the highest-value—and most expensive—activities in modern software development. Senior engineers must simultaneously detect functional bugs, security vulnerabilities, performance anti-patterns, and style issues, then communicate fixes clearly.
This environment measures whether an LLM can:
- Identify specific, named issues rather than vague generalities.
- Correctly calibrate severity (not everything is "high").
- Produce actionable, keyword-rich suggestions that a developer can act on.
3. Action Space
{
"issues": ["<issue 1>", "<issue 2>", "..."],
"severity": "low | medium | high",
"suggestion": "<detailed fix recommendation>"
}4. Observation Space
{
"code_snippet": "<source code string>",
"language": "<python | javascript | sql>",
"history": [{"step": 0, "action": {...}, "reward": 0.72}, ...],
"task_type": "<bug_detection | optimization | full_review>",
"task_index": 0,
"task_description":"<what the agent should look for>",
"max_steps": 3,
"step": 0
}The history field contains all previous (action, reward) pairs within the current task, enabling the agent to self-correct across attempts.
5. Task Descriptions
Task 0 — Easy: Bug Detection
The agent reviews an index-validation function containing a classic off-by-one error: index <= len(lst) should be index < len(lst). A call with index == len(lst) returns True despite being out of bounds.
Expected issues: off-by-one, wrong boundary operator, index out of bounds. Suggestion keywords: index < len, strict less than, off-by-one, fix.
Task 1 — Medium: Optimization + Logic
The agent reviews a duplicate-detection function with O(n²) nested loops, range(len()) anti-patterns, and redundant not in membership checks.
Expected issues: quadratic complexity, nested loop, use range(len()), redundant check. Suggestion keywords: set, Counter, O(n), collections, efficient.
Task 2 — Hard: Full Code Review
The agent reviews a user-authentication module with 7 distinct issues:
- MD5 is cryptographically insecure → use
bcrypt/argon2 - SQL injection (string interpolation in queries) → parameterized queries
- Connection resource leak (never closed) → context managers (
with) - No duplicate-username check
- No input validation
- No exception handling
- Timing attack vulnerability →
hmac.compare_digest/secrets
6. Reward Design
Reward at each step is computed deterministically:
reward = 0.50 × issue_overlap_score
+ 0.20 × severity_score
+ 0.30 × suggestion_keyword_score
- penaltiesPenalties (subtracted after weighting):
All rewards are clamped to [0.0, 1.0].
Episode score = mean of per-task best-step rewards (best-of-3 per task).
7. Setup Instructions
Prerequisites
- Python ≥ 3.9
- pip
Local setup
git clone <repo-url>
cd openenv
# Create virtual environment
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# Install dependencies
pip install -r requirements.txtDocker setup
# Build
docker build -t code-review-env .
# Run (override env vars at runtime)
docker run --rm \
-e API_BASE_URL="https://api.openai.com/v1" \
-e MODEL_NAME="gpt-4o" \
-e HF_TOKEN="sk-..." \
code-review-env8. How to Run Inference
Environment variables
Run locally
export API_BASE_URL="https://api.openai.com/v1"
export MODEL_NAME="gpt-4o"
export HF_TOKEN="sk-YOUR_KEY_HERE"
python inference.pyExpected log output
[START]
Model : gpt-4o
API Base : https://api.openai.com/v1
Tasks : 3
[STEP] #1
Task index : 0
Task step : 1
Severity : high
Issues found : ['off-by-one error', 'index boundary check is wrong']
Suggestion : Change `index <= len(lst)` to `index < len(lst)` to fix the off-by-one error...
Step reward : 0.8150
Grading : issue=0.800 | severity=1.000 | suggestion=0.667 | penalty=0.000
Info message : Task 'easy_bug_detection' — step 1/3. Step reward: 0.8150.
...
[END]
Total steps : 9
All step rewards: [0.8150, 0.7300, ...]
Final score : 0.7867
RESULT_JSON: {"final_score": 0.7867, "total_steps": 9, ...}9. Example Baseline Scores
Scores above 0.70 are considered competitive. Scores above 0.85 indicate expert-level code review performance.
Project Structure
openenv/
├── env/
│ ├── __init__.py # Public API
│ ├── models.py # Pydantic models (Action, Observation, StepReward)
│ ├── environment.py # Core OpenEnv environment class
│ ├── tasks.py # Task definitions (Easy / Medium / Hard)
│ └── grader.py # Deterministic grading logic
├── inference.py # Inference entry point
├── openenv.yaml # OpenEnv config
├── Dockerfile # Container build
├── requirements.txt # Python dependencies
└── README.md # This fileLicense
MIT — see LICENSE.
