Team Ai
Apppublic

harshit078/code-review-agent1

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

Code Review Agent — OpenEnv Environment

A reinforcement learning environment where an AI agent learns to review Python code across three difficulty levels.

Real-World Problem

Millions of code reviews happen daily at tech companies. Junior developers miss syntax errors, security vulnerabilities, and style issues. This environment trains an AI agent to perform accurate, structured code reviews — reducing review time and improving code quality.

Environment Overview

PropertyValue
Environment TypeText-based RL
Tasks3 (easy, medium, hard)
LanguagePython
Reward Range0.0 – 1.0
DeploymentHugging Face Spaces (Docker)

How It Works

The agent interacts with the environment through 3 API endpoints:

  • —POST /reset — loads a random code sample and returns it as observation
  • —POST /step — accepts the agent's text review and returns a reward score
  • —GET /state — returns current task info and history

Tasks

Task 1 — Easy: Syntax Error Detection

Agent must identify the exact error type (SyntaxError, IndentationError, etc.) and the line number where it occurs.

Reward tiers:

  • —1.0 — correct error type + correct line number
  • —0.6 — correct error type, wrong line number
  • —0.3 — related keywords mentioned
  • —0.0 — missed completely

Task 2 — Medium: Security Vulnerability Detection

Agent must identify security vulnerabilities — SQL injection, hardcoded credentials, eval injection, path traversal, command injection, and more.

Reward tiers:

  • —1.0 — 4+ vulnerability keywords matched
  • —0.75 — 3 keywords matched
  • —0.5 — 2 keywords matched
  • —0.25 — 1 keyword matched
  • —0.0 — missed completely

Task 3 — Hard: Full Code Quality Review

Agent must cover all 4 review categories. Each category is worth 0.25.

CategoryWorthWhat Agent Must Cover
Security0.25Vulnerabilities, unsafe functions
Bugs0.25Logic errors, missing error handling
Style0.25Variable names, docstrings, type hints
Suggestions0.25Specific fixes with examples

Observation Space

json
{
  "code": "def login(user, pwd):\n    ...",
  "task_type": "easy",
  "difficulty": "easy",
  "language": "python",
  "instructions": "Review this code and identify all syntax errors..."
}

Action Space

Natural language text — the agent's code review response.

Reward Design

Partial scoring is used throughout — agents are never penalized with binary right/wrong. Every partially correct answer earns credit, which enables meaningful gradient signal for RL training.

Running Locally

bash
# Install dependencies
pip install -r requirements.txt

# Start the environment server
uvicorn app:app --reload --port 8000

# Run the baseline agent
python inference.py

Running with Docker

bash
docker build -t code-review-agent .
docker run -p 7860:7860 \
  -e API_BASE_URL=https://api.openai.com/v1 \
  -e MODEL_NAME=gpt-4o-mini \
  -e HF_TOKEN=your_key_here \
  code-review-agent

Environment Variables

VariableDescription
API_BASE_URLLLM API base URL
MODEL_NAMEModel to use for inference
HF_TOKENAPI key / HF token
ENV_URLEnvironment server URL (default: localhost:8000)

Project Structure

code-review-agent/
├── code_samples/
│   ├── easy_samples.json       # 10 syntax error samples
│   ├── medium_samples.json     # 10 security vulnerability samples
│   └── hard_samples.json       # 5 multi-issue samples
├── graders/
│   ├── easy_grader.py          # Keyword + line number scoring
│   ├── medium_grader.py        # Vulnerability keyword scoring
│   └── hard_grader.py          # 4-category partial scoring
├── environment.py              # Core OpenEnv class
├── app.py                      # FastAPI server
├── inference.py                # Baseline AI agent
├── openenv.yaml                # OpenEnv spec config
├── Dockerfile                  # HF Spaces deployment
└── requirements.txt