Team Ai
Apppublic

kaysav18/code-review-assistant

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

Code Review Assistant — OpenEnv Environment

Submission-ready OpenEnv environment that challenges LLM agents to act as senior software engineers performing structured, multi-task code reviews.

1. Environment Description

The Code Review Assistant Environment presents an LLM agent with real Python code snippets and asks it to produce a structured review consisting of:

FieldDescription
issuesList of specific bugs / bad practices found
severityOverall severity: low, medium, or high
suggestionActionable natural-language fix recommendation

The environment evaluates responses deterministically using a weighted grading function and provides step-level reward feedback so that the agent can improve its review on subsequent attempts within the same task.


2. Real-World Motivation

Code review is one of the highest-value—and most expensive—activities in modern software development. Senior engineers must simultaneously detect functional bugs, security vulnerabilities, performance anti-patterns, and style issues, then communicate fixes clearly.

This environment measures whether an LLM can:

  • —Identify specific, named issues rather than vague generalities.
  • —Correctly calibrate severity (not everything is "high").
  • —Produce actionable, keyword-rich suggestions that a developer can act on.

3. Action Space

json
{
  "issues": ["<issue 1>", "<issue 2>", "..."],
  "severity": "low | medium | high",
  "suggestion": "<detailed fix recommendation>"
}
FieldTypeConstraints
issueslist[str]≥ 0 strings; each non-empty
severityenumexactly low, medium, or high
suggestionstrShould be ≥ 20 words for full credit

4. Observation Space

json
{
  "code_snippet":    "<source code string>",
  "language":        "<python | javascript | sql>",
  "history":         [{"step": 0, "action": {...}, "reward": 0.72}, ...],
  "task_type":       "<bug_detection | optimization | full_review>",
  "task_index":      0,
  "task_description":"<what the agent should look for>",
  "max_steps":       3,
  "step":            0
}

The history field contains all previous (action, reward) pairs within the current task, enabling the agent to self-correct across attempts.


5. Task Descriptions

Task 0 — Easy: Bug Detection

PropertyValue
IDeasy_bug_detection
DifficultyEasy
Max Steps3
Expected Severityhigh

The agent reviews an index-validation function containing a classic off-by-one error: index <= len(lst) should be index < len(lst). A call with index == len(lst) returns True despite being out of bounds.

Expected issues: off-by-one, wrong boundary operator, index out of bounds. Suggestion keywords: index < len, strict less than, off-by-one, fix.


Task 1 — Medium: Optimization + Logic

PropertyValue
IDmedium_optimization_logic
DifficultyMedium
Max Steps3
Expected Severitymedium

The agent reviews a duplicate-detection function with O(n²) nested loops, range(len()) anti-patterns, and redundant not in membership checks.

Expected issues: quadratic complexity, nested loop, use range(len()), redundant check. Suggestion keywords: set, Counter, O(n), collections, efficient.


Task 2 — Hard: Full Code Review

PropertyValue
IDhard_full_review
DifficultyHard
Max Steps3
Expected Severityhigh

The agent reviews a user-authentication module with 7 distinct issues:

  1. 1.MD5 is cryptographically insecure → use bcrypt / argon2
  2. 2.SQL injection (string interpolation in queries) → parameterized queries
  3. 3.Connection resource leak (never closed) → context managers (with)
  4. 4.No duplicate-username check
  5. 5.No input validation
  6. 6.No exception handling
  7. 7.Timing attack vulnerability → hmac.compare_digest / secrets

6. Reward Design

Reward at each step is computed deterministically:

reward = 0.50 × issue_overlap_score
       + 0.20 × severity_score
       + 0.30 × suggestion_keyword_score
       - penalties
ComponentWeightMethod
Issue overlap50%Substring match: expected keywords found in reported issues
Severity20%Exact match (binary)
Suggestion30%Keyword hit ratio: required keywords found in suggestion text

Penalties (subtracted after weighting):

ViolationPenalty
Empty issues list−0.20
Empty / too-short suggestion−0.15
Repeated identical action−0.10

All rewards are clamped to [0.0, 1.0].

Episode score = mean of per-task best-step rewards (best-of-3 per task).


7. Setup Instructions

Prerequisites

  • —Python ≥ 3.9
  • —pip

Local setup

bash
git clone <repo-url>
cd openenv

# Create virtual environment
python -m venv .venv
source .venv/bin/activate      # Windows: .venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

Docker setup

bash
# Build
docker build -t code-review-env .

# Run (override env vars at runtime)
docker run --rm \
  -e API_BASE_URL="https://api.openai.com/v1" \
  -e MODEL_NAME="gpt-4o" \
  -e HF_TOKEN="sk-..." \
  code-review-env

8. How to Run Inference

Environment variables

VariableDescriptionDefault
API_BASE_URLOpenAI-compatible base URLhttps://api.openai.com/v1
MODEL_NAMEModel identifiergpt-4o
HF_TOKENAPI bearer token(required)

Run locally

bash
export API_BASE_URL="https://api.openai.com/v1"
export MODEL_NAME="gpt-4o"
export HF_TOKEN="sk-YOUR_KEY_HERE"

python inference.py

Expected log output

[START]
  Model      : gpt-4o
  API Base   : https://api.openai.com/v1
  Tasks      : 3

[STEP] #1
  Task index   : 0
  Task step    : 1
  Severity     : high
  Issues found : ['off-by-one error', 'index boundary check is wrong']
  Suggestion   : Change `index <= len(lst)` to `index < len(lst)` to fix the off-by-one error...
  Step reward  : 0.8150
  Grading      : issue=0.800 | severity=1.000 | suggestion=0.667 | penalty=0.000
  Info message : Task 'easy_bug_detection' — step 1/3. Step reward: 0.8150.
...

[END]
  Total steps     : 9
  All step rewards: [0.8150, 0.7300, ...]
  Final score     : 0.7867

RESULT_JSON: {"final_score": 0.7867, "total_steps": 9, ...}

9. Example Baseline Scores

ModelEasy ScoreMedium ScoreHard Score**Final Score**
GPT-4o0.890.780.740.80
GPT-3.5-turbo0.720.550.430.57
Llama-3-8B0.650.480.380.50
Random baseline0.080.060.050.06
Scores above 0.70 are considered competitive. Scores above 0.85 indicate expert-level code review performance.

Project Structure

openenv/
├── env/
│   ├── __init__.py        # Public API
│   ├── models.py          # Pydantic models (Action, Observation, StepReward)
│   ├── environment.py     # Core OpenEnv environment class
│   ├── tasks.py           # Task definitions (Easy / Medium / Hard)
│   └── grader.py          # Deterministic grading logic
├── inference.py           # Inference entry point
├── openenv.yaml           # OpenEnv config
├── Dockerfile             # Container build
├── requirements.txt       # Python dependencies
└── README.md              # This file

License

MIT — see LICENSE.