Team Ai
Apppublic

anomalouslady/code-security-reviewer

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

CodeSecure — OpenEnv Environment

An AI agent environment for automated code security review. Agents analyze real Python code snippets and identify security vulnerabilities. Built for the OpenEnv Hackathon (Scaler × Hugging Face).


What the Agent Does

The agent receives a Python code snippet and must identify security vulnerabilities, explain their risk, and suggest fixes. Three tasks of increasing difficulty:

TaskDifficultyVulnerabilities
sql-injectionEasySQL injection via f-string query
hardcoded-secretsMediumHardcoded API keys, passwords, tokens
multi-vulnHardSQL injection + secrets + pickle deserialization + path traversal

API Reference

POST /reset

Start a new episode.

Request:

json
{ "task": "sql-injection" }

Valid tasks: sql-injection, hardcoded-secrets, multi-vuln

Response:

json
{
  "session_id": "uuid",
  "task": "sql-injection",
  "difficulty": "easy",
  "code_snippet": "...",
  "instructions": "...",
  "step": 0,
  "max_steps": 3,
  "done": false
}

POST /step

Submit the agent's analysis.

Request:

json
{
  "session_id": "uuid",
  "action": "This code is vulnerable to SQL injection because..."
}

Response:

json
{
  "state": { ... },
  "reward": 0.75,
  "done": false,
  "error": null,
  "feedback": "Matched 7/9 key indicators."
}

GET /state?session_id=<uuid>

Retrieve current session state.

GET /health

Health check. Returns {"status": "ok"}.


Observation Space

Type: Text The agent receives a Python code snippet alongside task instructions.

Action Space

Type: Text The agent returns a natural language security analysis.

Reward

  • —Range: [0.0, 1.0]
  • —Partial credit for each correctly identified vulnerability indicator
  • —Bonus credit for recommending correct fixes (parameterized queries, env vars, etc.)
  • —Episode ends when reward ≥ 0.8 or max steps (3) reached

Running Locally

bash
pip install -r requirements.txt
uvicorn main:app --host 0.0.0.0 --port 8000

Then in another terminal:

bash
export HF_TOKEN=your_token
export API_BASE_URL=https://router.huggingface.co/v1
export MODEL_NAME=Qwen/Qwen2.5-72B-Instruct
export BASE_URL=http://localhost:8000
python inference.py

Running with Docker

bash
docker build -t code-security-env .
docker run -p 7860:7860 \
  -e HF_TOKEN=your_token \
  -e API_BASE_URL=https://router.huggingface.co/v1 \
  -e MODEL_NAME=Qwen/Qwen2.5-72B-Instruct \
  code-security-env

Baseline Performance Scores

These scores were produced by running inference.py with Qwen/Qwen2.5-72B-Instruct via the HuggingFace router:

TaskDifficultyBaseline ScoreSuccess
sql-injectionEasy0.83✅
hardcoded-secretsMedium0.80✅
multi-vulnHard1.00✅

Scores are deterministic — graders are keyword-based with no randomness.


Reward Design

Partial progress: Each correctly identified vulnerability keyword adds to the reward incrementally — the agent gets credit even for partial analysis.

Penalties:

  • —Empty action → reward = 0.0, episode counts the step
  • —Repeated identical actions do not gain additional reward
  • —Exceeding max_steps (3) ends the episode with whatever score was last achieved

Project Structure

.
├── main.py          # FastAPI server (reset/step/state endpoints)
├── tasks.py         # Task definitions and grading logic
├── inference.py     # Baseline inference script
├── openenv.yaml     # OpenEnv spec
├── Dockerfile       # Container config
├── requirements.txt
└── README.md