Team Ai
Apppublic

AshutoshRajS/code-review-assistant

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

Code Review Assistant — OpenEnv Environment

An OpenEnv environment where AI agents learn to perform structured, professional code reviews — one of the most common and high-value tasks in software engineering.

Why Code Review?

Code review is a universal real-world task: every software team does it daily, yet it requires synthesizing multiple skills — bug detection, security intuition, performance awareness, and clear communication. It's a rich domain for agentic evaluation because:

  • —Ground truth is deterministic (known bugs at known lines)
  • —Partial credit is meaningful (catching 2/4 bugs is better than 0/4)
  • —Difficulty scales naturally (style nits → critical security flaws)
  • —Output is structured (decision + line comments + summary)

Environment Overview

The agent receives a pull request (title, description, code files with line numbers) and must produce a code review consisting of:

FieldTypeDescription
decisionenumapprove, request_changes, or comment
commentslistInline comments: {line, category, severity, message, suggestion}
summarystring1–3 sentence overall assessment

Observation Space

json
{
  "pull_request": {
    "title": "string",
    "description": "string",
    "author": "string",
    "files": [{"filename": "...", "language": "...", "content": "...", "line_count": 42}],
    "diff_stats": {"additions": 54, "deletions": 0}
  },
  "step": 1,
  "max_steps": 3,
  "task_name": "easy-bug-detection",
  "instructions": "Review this module...",
  "previous_feedback": "Step 1 score: 0.612. Caught 2/3 bugs."
}

Action Space

json
{
  "decision": "request_changes",
  "summary": "This PR has three critical bugs that must be fixed before merging.",
  "comments": [
    {
      "line": 5,
      "category": "bug",
      "severity": "critical",
      "message": "ZeroDivisionError when input list is empty.",
      "suggestion": "Add: if not numbers: return 0.0"
    }
  ]
}

Category values: bug, security, performance, style, maintainability, logic, documentation Severity values: critical, major, minor, suggestion


Tasks

Task 1 — easy-bug-detection 🟢

Difficulty: Easy | File: utils.py (25 lines)

Review a Python utility module. Find obvious runtime bugs:

  • —Line 5: ZeroDivisionError when input is empty list
  • —Line 8: IndexError when input is empty list
  • —Line 15: Off-by-one error in range() causing IndexError

Expected decision: request_changes Success threshold: 0.50


Task 2 — medium-security-review 🟡

Difficulty: Medium | File: app.py (54 lines)

Review a Flask authentication API. Find security vulnerabilities:

  • —Line 8: Hardcoded secret key
  • —Line 21: SQL injection via f-string query
  • —Line 24: MD5 used for token generation (broken crypto)
  • —Line 37: Password exposed in API response + SQL injection
  • —Line 44: No authentication on update_password endpoint
  • —Line 53: debug=True in production

Expected decision: request_changes Success threshold: 0.50


Task 3 — hard-async-pipeline-review 🔴

Difficulty: Hard | File: pipeline.py (66 lines)

Review an async data pipeline. Find subtle issues:

  • —Line 20: lru_cache applied to async coroutine (breaks caching)
  • —Line 15: aiohttp.ClientSession created outside event loop
  • —Line 28: Silent except swallowing all errors
  • —Line 44: KeyError on missing id/value fields
  • —Line 56: Sequential batch processing instead of asyncio.gather
  • —Line 63: Unsafe __del__ calling run_until_complete

Expected decision: request_changes Success threshold: 0.50


Reward Function

Each step returns:

reward = grader_score + 0.1 × max(0, grader_score - previous_score)

The grader_score is a weighted combination:

ComponentWeightDescription
Issue line coverage30–40%Did the agent find the real bugs?
Decision accuracy15–25%Was the overall decision correct?
Category correctness20–25%Were issues categorized right?
Severity accuracy10–15%Were severities appropriate?
Summary quality10%Does the summary mention key concerns?
False positive penalty−5% eachSpurious critical comments on clean code

Scores are clamped to [0.0, 1.0].


Baseline Scores

TaskModelScore
easy-bug-detectionQwen2.5-72B-Instruct~0.72
medium-security-reviewQwen2.5-72B-Instruct~0.65
hard-async-pipeline-reviewQwen2.5-72B-Instruct~0.48

Setup & Usage

Prerequisites

  • —Python 3.10+
  • —Docker
  • —A Hugging Face account and API token

Local Development

bash
# Clone the repo
git clone https://huggingface.co/spaces/<your-username>/code-review-assistant
cd code-review-assistant

# Install server deps
pip install -r server/requirements.txt

# Start the server
cd server
uvicorn main:app --host 0.0.0.0 --port 7860 --reload

Docker

bash
# Build
docker build -t code-review-assistant .

# Run
docker run -p 7860:7860 code-review-assistant

Running the Inference Script

bash
# Install inference deps
pip install -r requirements-inference.txt

# Set environment variables
export HF_TOKEN="your_hf_token"
export API_BASE_URL="https://router.huggingface.co/v1"
export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
export ENV_BASE_URL="https://ashutoshrajs-code-review-assistant.hf.space"

# Run all 3 tasks
python inference.py

# Run a single task
export CODE_REVIEW_TASK="easy-bug-detection"
python inference.py

Example Output

============================================================
Running task: easy-bug-detection
============================================================
[START] task=easy-bug-detection env=code-review-assistant model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=decision=request_changes comments=4 summary='Three critical bugs will cause runtime errors on empty input...' reward=0.74 done=true error=null
[END] success=true steps=1 score=0.740 rewards=0.74

API Endpoints

MethodPathDescription
GET/healthHealth check
GET/tasksList available tasks
POST/resetStart new episode ({"task_name": "..."})
POST/stepSubmit review action
GET/stateCurrent episode state

Quick curl test

bash
# Reset
curl -s -X POST http://localhost:7860/reset \
  -H "Content-Type: application/json" \
  -d '{"task_name": "easy-bug-detection"}' | python -m json.tool

# Step
curl -s -X POST http://localhost:7860/step \
  -H "Content-Type: application/json" \
  -d '{
    "action": {
      "decision": "request_changes",
      "summary": "Critical bugs: empty list crashes on lines 5, 8, and 15.",
      "comments": [
        {"line": 5, "category": "bug", "severity": "critical",
         "message": "ZeroDivisionError if numbers is empty",
         "suggestion": "if not numbers: return 0.0"},
        {"line": 8, "category": "bug", "severity": "critical",
         "message": "IndexError if lst is empty",
         "suggestion": "if not lst: return None"},
        {"line": 15, "category": "bug", "severity": "major",
         "message": "Off-by-one: range(len(s), 0, -1) is out of range",
         "suggestion": "range(len(s)-1, -1, -1)"}
      ]
    }
  }' | python -m json.tool

Project Structure

code-review-assistant/
├── Dockerfile               # Container for HF Spaces
├── openenv.yaml             # OpenEnv spec compliance
├── inference.py             # Baseline agent script
├── requirements-inference.txt
├── README.md
└── server/
    ├── main.py              # FastAPI server
    ├── env.py               # Core environment logic
    ├── models.py            # Pydantic data models
    ├── tasks.py             # Task definitions + code snippets
    ├── graders.py           # Scoring/grading logic
    └── requirements.txt

License

MIT