Team Ai
Apppublic

RitwikGupta-0501/Code-Review

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

๐Ÿ” OpenEnv: Code Review

A real-world reinforcement learning environment where AI agents review pull requests, identify bugs, audit security vulnerabilities, and produce actionable code review output.

![openenv](https://huggingface.co) ![Python 3.11+](https://python.org) ![License: MIT](LICENSE) ![Tests: 25 passing]()


What Is This?

OpenEnv: Code Review simulates a real-world software engineering task: reviewing pull requests.

An AI agent is presented with realistic pull request diffs across Python, JavaScript, Go, and TypeScript. The agent must:

  • โ€”Identify bugs โ€” off-by-one errors, race conditions, memory leaks, logic errors
  • โ€”Audit security โ€” SQL injection, command injection, hardcoded secrets, broken auth, insecure deserialization
  • โ€”Route correctly โ€” assign the PR to the right reviewer role (backend, security, senior)
  • โ€”Summarise findings โ€” write actionable REQUEST_CHANGES or APPROVE decisions

This is a task that software engineers do every day. Training agents to do it well has immediate real-world value.


Quick Start

Option A: Docker (recommended)

bash
git clone https://github.com/your-repo/openenv-code-review
cd openenv-code-review

docker build -t openenv-code-review .
docker run -p 7860:7860 openenv-code-review

# Server is live at http://localhost:7860

Option B: Local Python

bash
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 7860

Option C: HuggingFace Space

Visit the live deployment at: https://huggingface.co/spaces/YOUR_USERNAME/openenv-code-review


Running an Agent

python
import requests

BASE = "http://localhost:7860"

# 1. Start a bug detection episode
obs = requests.post(f"{BASE}/reset?task=bug_detection").json()
print(obs["message"])
print(obs["current_pr"]["title"])

# 2. Flag a bug
result = requests.post(f"{BASE}/step", json={
    "action_type": "flag_bug",
    "bug_category": "off_by_one",
    "line_number": 22,
    "severity": "medium"
}).json()

print(f"Reward: {result['reward']['value']}")        # +0.28
print(f"Explanation: {result['reward']['explanation']}")

# 3. Flag a security issue
result = requests.post(f"{BASE}/step", json={
    "action_type": "flag_security",
    "security_category": "sql_injection",
    "line_number": 17,
    "severity": "high"
}).json()

# 4. Assign reviewer and close out the PR
requests.post(f"{BASE}/step", json={
    "action_type": "assign_reviewer",
    "reviewer_role": "backend"
})

requests.post(f"{BASE}/step", json={
    "action_type": "request_changes",
    "comment": "Found an off-by-one in retry logic and SQL injection in get_user. Both need fixing before merge."
})

# 5. Get your final score
score = requests.get(f"{BASE}/grader").json()
print(f"Score: {score['score']}")           # 0.0 โ€“ 1.0
print(f"Breakdown: {score['breakdown']}")

API Reference

MethodEndpointDescription
POST/reset?task=<task_id>Start a new episode
POST/stepSubmit an action, receive observation + reward
GET/stateCurrent full environment state
GET/tasksList all tasks + full action JSON schema
GET/graderScore the current episode (0.0โ€“1.0)
GET/baselineRun baseline inference (requires OPENAI_API_KEY)
GET/docsInteractive Swagger UI

Observation Space

Every call to /reset and /step returns an Observation:

json
{
  "task": "bug_detection",
  "step_number": 3,
  "current_pr": {
    "pr_id": "PR-101",
    "title": "Add user pagination to /api/users endpoint",
    "description": "...",
    "author": "junior_dev_1",
    "files": [
      {
        "filename": "api/users.py",
        "language": "python",
        "patch": "--- a/api/users.py\n+++ b/api/users.py\n..."
      }
    ],
    "total_additions": 45,
    "total_deletions": 4
  },
  "current_file_index": 0,
  "prs_remaining": 2,
  "flags_raised": [...],
  "cumulative_reward": 0.56,
  "elapsed_steps": 3,
  "message": "Bug flagged: off_by_one at line 22",
  "done": false
}

Action Space

Every call to /step takes an Action:

json
{
  "action_type": "flag_bug | flag_security | add_comment | assign_reviewer | request_changes | approve | skip",

  "bug_category": "logic_error | null_pointer | off_by_one | race_condition | memory_leak | infinite_loop | type_error | unhandled_exception",
  "security_category": "sql_injection | xss | hardcoded_secret | path_traversal | insecure_deserialization | broken_auth | sensitive_data_exposure | command_injection",
  "line_number": 22,
  "severity": "critical | high | medium | low | info",
  "reviewer_role": "backend | frontend | security | devops | senior",
  "comment": "Found SQL injection at line 17 โ€” use parameterized queries.",
  "escalation_reason": "optional string"
}

Required fields per action type:

action_typeRequired fields
flag_bugbug_category, line_number, severity
flag_securitysecurity_category, line_number, severity
add_commentcomment
assign_reviewerreviewer_role
request_changescomment
approve(none required, comment optional)
skip(none required)

Tasks

Task 1 โ€” Bug Detection bug_detection ๐ŸŸข Easy

Objective: Review 2 PRs (Python + JavaScript) and identify all bugs.

PRs in this task:

  • โ€”PR-101 โ€” Python pagination endpoint with an off-by-one in retry logic and SQL injection in a route handler
  • โ€”PR-102 โ€” JavaScript shopping cart with an off-by-one loop bug and raw payment card data exposure

What a good agent does: FLAGBUG for the off-by-one errors, FLAGSECURITY for the injection/data issues, ASSIGNREVIEWER correctly, REQUESTCHANGES with a clear summary.

Max steps: 30 | Expected baseline score: ~0.71


Task 2 โ€” Security Audit security_audit ๐ŸŸก Medium

Objective: Security-focused review of 2 PRs (Python + Go). Find all vulnerabilities.

PRs in this task:

  • โ€”PR-201 โ€” Python file export feature with path traversal, SQL injection, command injection, SSTI, and two separate sets of hardcoded credentials
  • โ€”PR-202 โ€” Go JWT authentication with MD5-signed tokens, unverified token signature, SQL injection in login, MD5 password hashing, and no token expiry check

What a good agent does: Must recognise multiple vulnerability classes in the same file, correctly classify each (not just "this looks bad"), and produce a security-focused REQUEST_CHANGES summary.

Max steps: 40 | Expected baseline score: ~0.58


Task 3 โ€” Full Review full_review ๐Ÿ”ด Hard

Objective: Comprehensive review of 2 complex PRs (Python + TypeScript). Bugs, security issues, race conditions, and design flaws are interleaved.

PRs in this task:

  • โ€”PR-301 โ€” Python async job queue with insecure pickle deserialization (RCE risk), race conditions on the result cache, a thread leak in schedule_recurring, worker list never cleared on stop(), and a cache stampede vulnerability
  • โ€”PR-302 โ€” TypeScript multi-tenant SaaS middleware where a module-level mutable variable causes cross-tenant data leakage in async Node.js, missing tenant filter on UPDATE/DELETE queries, SQL injection via filter key interpolation, and secrets returned in full config responses

What a good agent does: Distinguish between bug and security categories for subtle issues (e.g. the race condition is a bug AND a security issue), flag all 9+ issues, assign SENIOR reviewer, write a structured multi-section summary.

Max steps: 60 | Expected baseline score: ~0.39


Reward Function

Reward is shaped across the full trajectory โ€” agents get signal on every step, not just at the end.

EventReward
Correctly flag a critical issue+0.50
Correctly flag a high severity issue+0.40
Correctly flag a medium severity issue+0.30
Correctly flag a low severity issue+0.15
Correct category label (vs just right line)Full credit
Wrong category but right line50% credit
False positive (flagging clean code)โˆ’0.20
Correct reviewer assigned+0.10
Wrong reviewer assignedโˆ’0.05
REQUEST_CHANGES when PR has issues+0.15
APPROVE when PR has issuesโˆ’0.15
Meaningful ADD_COMMENT (>30 chars)+0.05
SKIP actionโˆ’0.05
Each step takenโˆ’0.02 (efficiency pressure)

Line Number Tolerance

Agents are given ยฑ3 line tolerance when matching flags to ground truth. A bug at line 22 is credited if the agent flags lines 19โ€“25. This accounts for agents reasoning about a code block rather than the exact line.


Grader

The grader runs deterministically at episode end and returns a score in [0.0, 1.0] with a full breakdown:

json
{
  "score": 0.74,
  "breakdown": {
    "bug_recall": 0.875,
    "vuln_recall": 0.72,
    "precision": 0.9,
    "reviewer_correct": 1.0,
    "final_action": 1.0,
    "comment_quality": 0.6,
    "efficiency_bonus": 0.047,
    "weighted_total": 0.7401
  },
  "task": "bug_detection"
}

Score interpretation:

RangeMeaning
0.0 โ€“ 0.3Poor โ€” missed most issues
0.3 โ€“ 0.6Partial โ€” found some, missed key ones
0.6 โ€“ 0.8Good โ€” found most issues with reasonable precision
0.8 โ€“ 1.0Excellent โ€” comprehensive review with correct labels

Baseline Scores

Run against gpt-4o-mini with the included baseline.py:

TaskScoreDifficulty
bug_detection0.71Easy
security_audit0.58Medium
full_review0.39Hard
Average0.56โ€”

Running the Baseline Yourself

bash
export OPENAI_API_KEY=sk-...
python -m app.baseline

# Or via the API:
curl http://localhost:7860/baseline

Project Structure

openenv-code-review/
โ”‚
โ”œโ”€โ”€ app/
โ”‚   โ”œโ”€โ”€ __init__.py     โ€” package marker
โ”‚   โ”œโ”€โ”€ main.py         โ€” FastAPI app, all HTTP endpoints
โ”‚   โ”œโ”€โ”€ env.py          โ€” Core env: reset() / step() / state() / grade()
โ”‚   โ”œโ”€โ”€ models.py       โ€” Pydantic v2 typed models (Observation, Action, Reward)
โ”‚   โ”œโ”€โ”€ corpus.py       โ€” 6 annotated PRs with hidden ground truth
โ”‚   โ”œโ”€โ”€ tasks.py        โ€” Task definitions + deterministic grader
โ”‚   โ””โ”€โ”€ baseline.py     โ€” GPT-4o-mini baseline inference script
โ”‚
โ”œโ”€โ”€ tests/
โ”‚   โ”œโ”€โ”€ conftest.py     โ€” pytest configuration
โ”‚   โ””โ”€โ”€ test_env.py     โ€” 25 unit tests (reset, step, state, grader, rewards)
โ”‚
โ”œโ”€โ”€ openenv.yaml        โ€” OpenEnv spec metadata
โ”œโ”€โ”€ Dockerfile          โ€” Container for HuggingFace Spaces
โ”œโ”€โ”€ requirements.txt    โ€” Python dependencies
โ””โ”€โ”€ README.md           โ€” This file

Deploying to HuggingFace Spaces

  1. 1.Create a new Space at huggingface.co/new-space
  2. 2.SDK: Docker
  3. 3.Visibility: Public
  1. 1.Push this repo to the Space:
bash
git remote add hf https://huggingface.co/spaces/YOUR_USERNAME/openenv-code-review
git push hf main
  1. 1.(Optional) Add your OPENAI_API_KEY as a Space secret for live baseline:
  2. 2.Space Settings โ†’ Variables and Secrets โ†’ New Secret
  1. 1.Verify deployment:
bash
curl https://YOUR_USERNAME-openenv-code-review.hf.space/tasks

Running Tests

bash
pip install -r requirements.txt pytest
python -m pytest tests/ -v

Expected output: 25 passed


OpenEnv Spec Compliance

RequirementStatus
Typed Observation modelโœ… Pydantic v2
Typed Action modelโœ… Pydantic v2
Typed Reward modelโœ… with component breakdown
POST /resetโœ…
POST /stepโœ… returns observation, reward, done, info
GET /stateโœ…
GET /tasksโœ… with action schema
GET /graderโœ… deterministic 0.0โ€“1.0
GET /baselineโœ…
openenv.yamlโœ…
3+ tasks with difficulty rangeโœ… easy โ†’ medium โ†’ hard
Shaped reward functionโœ… per-step signal
Dockerfileโœ… builds and runs
HuggingFace Spaceโœ… Docker SDK

License

MIT โ€” see LICENSE