AshutoshRajS/code-review-assistant
Code Review Assistant — OpenEnv Environment
An OpenEnv environment where AI agents learn to perform structured, professional code reviews — one of the most common and high-value tasks in software engineering.
Why Code Review?
Code review is a universal real-world task: every software team does it daily, yet it requires synthesizing multiple skills — bug detection, security intuition, performance awareness, and clear communication. It's a rich domain for agentic evaluation because:
- Ground truth is deterministic (known bugs at known lines)
- Partial credit is meaningful (catching 2/4 bugs is better than 0/4)
- Difficulty scales naturally (style nits → critical security flaws)
- Output is structured (decision + line comments + summary)
Environment Overview
The agent receives a pull request (title, description, code files with line numbers) and must produce a code review consisting of:
Observation Space
{
"pull_request": {
"title": "string",
"description": "string",
"author": "string",
"files": [{"filename": "...", "language": "...", "content": "...", "line_count": 42}],
"diff_stats": {"additions": 54, "deletions": 0}
},
"step": 1,
"max_steps": 3,
"task_name": "easy-bug-detection",
"instructions": "Review this module...",
"previous_feedback": "Step 1 score: 0.612. Caught 2/3 bugs."
}Action Space
{
"decision": "request_changes",
"summary": "This PR has three critical bugs that must be fixed before merging.",
"comments": [
{
"line": 5,
"category": "bug",
"severity": "critical",
"message": "ZeroDivisionError when input list is empty.",
"suggestion": "Add: if not numbers: return 0.0"
}
]
}Category values: bug, security, performance, style, maintainability, logic, documentation Severity values: critical, major, minor, suggestion
Tasks
Task 1 — easy-bug-detection 🟢
Difficulty: Easy | File: utils.py (25 lines)
Review a Python utility module. Find obvious runtime bugs:
- Line 5:
ZeroDivisionErrorwhen input is empty list - Line 8:
IndexErrorwhen input is empty list - Line 15: Off-by-one error in
range()causingIndexError
Expected decision: request_changes Success threshold: 0.50
Task 2 — medium-security-review 🟡
Difficulty: Medium | File: app.py (54 lines)
Review a Flask authentication API. Find security vulnerabilities:
- Line 8: Hardcoded secret key
- Line 21: SQL injection via f-string query
- Line 24: MD5 used for token generation (broken crypto)
- Line 37: Password exposed in API response + SQL injection
- Line 44: No authentication on
update_passwordendpoint - Line 53:
debug=Truein production
Expected decision: request_changes Success threshold: 0.50
Task 3 — hard-async-pipeline-review 🔴
Difficulty: Hard | File: pipeline.py (66 lines)
Review an async data pipeline. Find subtle issues:
- Line 20:
lru_cacheapplied toasynccoroutine (breaks caching) - Line 15:
aiohttp.ClientSessioncreated outside event loop - Line 28: Silent
exceptswallowing all errors - Line 44:
KeyErroron missingid/valuefields - Line 56: Sequential batch processing instead of
asyncio.gather - Line 63: Unsafe
__del__callingrun_until_complete
Expected decision: request_changes Success threshold: 0.50
Reward Function
Each step returns:
reward = grader_score + 0.1 × max(0, grader_score - previous_score)The grader_score is a weighted combination:
Scores are clamped to [0.0, 1.0].
Baseline Scores
Setup & Usage
Prerequisites
- Python 3.10+
- Docker
- A Hugging Face account and API token
Local Development
# Clone the repo
git clone https://huggingface.co/spaces/<your-username>/code-review-assistant
cd code-review-assistant
# Install server deps
pip install -r server/requirements.txt
# Start the server
cd server
uvicorn main:app --host 0.0.0.0 --port 7860 --reloadDocker
# Build
docker build -t code-review-assistant .
# Run
docker run -p 7860:7860 code-review-assistantRunning the Inference Script
# Install inference deps
pip install -r requirements-inference.txt
# Set environment variables
export HF_TOKEN="your_hf_token"
export API_BASE_URL="https://router.huggingface.co/v1"
export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
export ENV_BASE_URL="https://ashutoshrajs-code-review-assistant.hf.space"
# Run all 3 tasks
python inference.py
# Run a single task
export CODE_REVIEW_TASK="easy-bug-detection"
python inference.pyExample Output
============================================================
Running task: easy-bug-detection
============================================================
[START] task=easy-bug-detection env=code-review-assistant model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=decision=request_changes comments=4 summary='Three critical bugs will cause runtime errors on empty input...' reward=0.74 done=true error=null
[END] success=true steps=1 score=0.740 rewards=0.74API Endpoints
Quick curl test
# Reset
curl -s -X POST http://localhost:7860/reset \
-H "Content-Type: application/json" \
-d '{"task_name": "easy-bug-detection"}' | python -m json.tool
# Step
curl -s -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"decision": "request_changes",
"summary": "Critical bugs: empty list crashes on lines 5, 8, and 15.",
"comments": [
{"line": 5, "category": "bug", "severity": "critical",
"message": "ZeroDivisionError if numbers is empty",
"suggestion": "if not numbers: return 0.0"},
{"line": 8, "category": "bug", "severity": "critical",
"message": "IndexError if lst is empty",
"suggestion": "if not lst: return None"},
{"line": 15, "category": "bug", "severity": "major",
"message": "Off-by-one: range(len(s), 0, -1) is out of range",
"suggestion": "range(len(s)-1, -1, -1)"}
]
}
}' | python -m json.toolProject Structure
code-review-assistant/
├── Dockerfile # Container for HF Spaces
├── openenv.yaml # OpenEnv spec compliance
├── inference.py # Baseline agent script
├── requirements-inference.txt
├── README.md
└── server/
├── main.py # FastAPI server
├── env.py # Core environment logic
├── models.py # Pydantic data models
├── tasks.py # Task definitions + code snippets
├── graders.py # Scoring/grading logic
└── requirements.txtLicense
MIT
