RitwikGupta-0501/Code-Review
๐ OpenEnv: Code Review
A real-world reinforcement learning environment where AI agents review pull requests, identify bugs, audit security vulnerabilities, and produce actionable code review output.
   ![Tests: 25 passing]()
What Is This?
OpenEnv: Code Review simulates a real-world software engineering task: reviewing pull requests.
An AI agent is presented with realistic pull request diffs across Python, JavaScript, Go, and TypeScript. The agent must:
- Identify bugs โ off-by-one errors, race conditions, memory leaks, logic errors
- Audit security โ SQL injection, command injection, hardcoded secrets, broken auth, insecure deserialization
- Route correctly โ assign the PR to the right reviewer role (backend, security, senior)
- Summarise findings โ write actionable
REQUEST_CHANGESorAPPROVEdecisions
This is a task that software engineers do every day. Training agents to do it well has immediate real-world value.
Quick Start
Option A: Docker (recommended)
git clone https://github.com/your-repo/openenv-code-review
cd openenv-code-review
docker build -t openenv-code-review .
docker run -p 7860:7860 openenv-code-review
# Server is live at http://localhost:7860Option B: Local Python
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 7860Option C: HuggingFace Space
Visit the live deployment at: https://huggingface.co/spaces/YOUR_USERNAME/openenv-code-review
Running an Agent
import requests
BASE = "http://localhost:7860"
# 1. Start a bug detection episode
obs = requests.post(f"{BASE}/reset?task=bug_detection").json()
print(obs["message"])
print(obs["current_pr"]["title"])
# 2. Flag a bug
result = requests.post(f"{BASE}/step", json={
"action_type": "flag_bug",
"bug_category": "off_by_one",
"line_number": 22,
"severity": "medium"
}).json()
print(f"Reward: {result['reward']['value']}") # +0.28
print(f"Explanation: {result['reward']['explanation']}")
# 3. Flag a security issue
result = requests.post(f"{BASE}/step", json={
"action_type": "flag_security",
"security_category": "sql_injection",
"line_number": 17,
"severity": "high"
}).json()
# 4. Assign reviewer and close out the PR
requests.post(f"{BASE}/step", json={
"action_type": "assign_reviewer",
"reviewer_role": "backend"
})
requests.post(f"{BASE}/step", json={
"action_type": "request_changes",
"comment": "Found an off-by-one in retry logic and SQL injection in get_user. Both need fixing before merge."
})
# 5. Get your final score
score = requests.get(f"{BASE}/grader").json()
print(f"Score: {score['score']}") # 0.0 โ 1.0
print(f"Breakdown: {score['breakdown']}")API Reference
Observation Space
Every call to /reset and /step returns an Observation:
{
"task": "bug_detection",
"step_number": 3,
"current_pr": {
"pr_id": "PR-101",
"title": "Add user pagination to /api/users endpoint",
"description": "...",
"author": "junior_dev_1",
"files": [
{
"filename": "api/users.py",
"language": "python",
"patch": "--- a/api/users.py\n+++ b/api/users.py\n..."
}
],
"total_additions": 45,
"total_deletions": 4
},
"current_file_index": 0,
"prs_remaining": 2,
"flags_raised": [...],
"cumulative_reward": 0.56,
"elapsed_steps": 3,
"message": "Bug flagged: off_by_one at line 22",
"done": false
}Action Space
Every call to /step takes an Action:
{
"action_type": "flag_bug | flag_security | add_comment | assign_reviewer | request_changes | approve | skip",
"bug_category": "logic_error | null_pointer | off_by_one | race_condition | memory_leak | infinite_loop | type_error | unhandled_exception",
"security_category": "sql_injection | xss | hardcoded_secret | path_traversal | insecure_deserialization | broken_auth | sensitive_data_exposure | command_injection",
"line_number": 22,
"severity": "critical | high | medium | low | info",
"reviewer_role": "backend | frontend | security | devops | senior",
"comment": "Found SQL injection at line 17 โ use parameterized queries.",
"escalation_reason": "optional string"
}Required fields per action type:
Tasks
Task 1 โ Bug Detection bug_detection ๐ข Easy
Objective: Review 2 PRs (Python + JavaScript) and identify all bugs.
PRs in this task:
PR-101โ Python pagination endpoint with an off-by-one in retry logic and SQL injection in a route handlerPR-102โ JavaScript shopping cart with an off-by-one loop bug and raw payment card data exposure
What a good agent does: FLAGBUG for the off-by-one errors, FLAGSECURITY for the injection/data issues, ASSIGNREVIEWER correctly, REQUESTCHANGES with a clear summary.
Max steps: 30 | Expected baseline score: ~0.71
Task 2 โ Security Audit security_audit ๐ก Medium
Objective: Security-focused review of 2 PRs (Python + Go). Find all vulnerabilities.
PRs in this task:
PR-201โ Python file export feature with path traversal, SQL injection, command injection, SSTI, and two separate sets of hardcoded credentialsPR-202โ Go JWT authentication with MD5-signed tokens, unverified token signature, SQL injection in login, MD5 password hashing, and no token expiry check
What a good agent does: Must recognise multiple vulnerability classes in the same file, correctly classify each (not just "this looks bad"), and produce a security-focused REQUEST_CHANGES summary.
Max steps: 40 | Expected baseline score: ~0.58
Task 3 โ Full Review full_review ๐ด Hard
Objective: Comprehensive review of 2 complex PRs (Python + TypeScript). Bugs, security issues, race conditions, and design flaws are interleaved.
PRs in this task:
PR-301โ Python async job queue with insecure pickle deserialization (RCE risk), race conditions on the result cache, a thread leak inschedule_recurring, worker list never cleared onstop(), and a cache stampede vulnerabilityPR-302โ TypeScript multi-tenant SaaS middleware where a module-level mutable variable causes cross-tenant data leakage in async Node.js, missing tenant filter on UPDATE/DELETE queries, SQL injection via filter key interpolation, and secrets returned in full config responses
What a good agent does: Distinguish between bug and security categories for subtle issues (e.g. the race condition is a bug AND a security issue), flag all 9+ issues, assign SENIOR reviewer, write a structured multi-section summary.
Max steps: 60 | Expected baseline score: ~0.39
Reward Function
Reward is shaped across the full trajectory โ agents get signal on every step, not just at the end.
Line Number Tolerance
Agents are given ยฑ3 line tolerance when matching flags to ground truth. A bug at line 22 is credited if the agent flags lines 19โ25. This accounts for agents reasoning about a code block rather than the exact line.
Grader
The grader runs deterministically at episode end and returns a score in [0.0, 1.0] with a full breakdown:
{
"score": 0.74,
"breakdown": {
"bug_recall": 0.875,
"vuln_recall": 0.72,
"precision": 0.9,
"reviewer_correct": 1.0,
"final_action": 1.0,
"comment_quality": 0.6,
"efficiency_bonus": 0.047,
"weighted_total": 0.7401
},
"task": "bug_detection"
}Score interpretation:
Baseline Scores
Run against gpt-4o-mini with the included baseline.py:
Running the Baseline Yourself
export OPENAI_API_KEY=sk-...
python -m app.baseline
# Or via the API:
curl http://localhost:7860/baselineProject Structure
openenv-code-review/
โ
โโโ app/
โ โโโ __init__.py โ package marker
โ โโโ main.py โ FastAPI app, all HTTP endpoints
โ โโโ env.py โ Core env: reset() / step() / state() / grade()
โ โโโ models.py โ Pydantic v2 typed models (Observation, Action, Reward)
โ โโโ corpus.py โ 6 annotated PRs with hidden ground truth
โ โโโ tasks.py โ Task definitions + deterministic grader
โ โโโ baseline.py โ GPT-4o-mini baseline inference script
โ
โโโ tests/
โ โโโ conftest.py โ pytest configuration
โ โโโ test_env.py โ 25 unit tests (reset, step, state, grader, rewards)
โ
โโโ openenv.yaml โ OpenEnv spec metadata
โโโ Dockerfile โ Container for HuggingFace Spaces
โโโ requirements.txt โ Python dependencies
โโโ README.md โ This fileDeploying to HuggingFace Spaces
- Create a new Space at huggingface.co/new-space
- SDK: Docker
- Visibility: Public
- Push this repo to the Space:
git remote add hf https://huggingface.co/spaces/YOUR_USERNAME/openenv-code-review
git push hf main- (Optional) Add your
OPENAI_API_KEYas a Space secret for live baseline: - Space Settings โ Variables and Secrets โ New Secret
- Verify deployment:
curl https://YOUR_USERNAME-openenv-code-review.hf.space/tasksRunning Tests
pip install -r requirements.txt pytest
python -m pytest tests/ -vExpected output: 25 passed
OpenEnv Spec Compliance
License
MIT โ see LICENSE
