anomalouslady/code-security-reviewer
CodeSecure — OpenEnv Environment
An AI agent environment for automated code security review. Agents analyze real Python code snippets and identify security vulnerabilities. Built for the OpenEnv Hackathon (Scaler × Hugging Face).
What the Agent Does
The agent receives a Python code snippet and must identify security vulnerabilities, explain their risk, and suggest fixes. Three tasks of increasing difficulty:
API Reference
POST /reset
Start a new episode.
Request:
{ "task": "sql-injection" }Valid tasks: sql-injection, hardcoded-secrets, multi-vuln
Response:
{
"session_id": "uuid",
"task": "sql-injection",
"difficulty": "easy",
"code_snippet": "...",
"instructions": "...",
"step": 0,
"max_steps": 3,
"done": false
}POST /step
Submit the agent's analysis.
Request:
{
"session_id": "uuid",
"action": "This code is vulnerable to SQL injection because..."
}Response:
{
"state": { ... },
"reward": 0.75,
"done": false,
"error": null,
"feedback": "Matched 7/9 key indicators."
}GET /state?session_id=<uuid>
Retrieve current session state.
GET /health
Health check. Returns {"status": "ok"}.
Observation Space
Type: Text The agent receives a Python code snippet alongside task instructions.
Action Space
Type: Text The agent returns a natural language security analysis.
Reward
- Range:
[0.0, 1.0] - Partial credit for each correctly identified vulnerability indicator
- Bonus credit for recommending correct fixes (parameterized queries, env vars, etc.)
- Episode ends when reward ≥ 0.8 or max steps (3) reached
Running Locally
pip install -r requirements.txt
uvicorn main:app --host 0.0.0.0 --port 8000Then in another terminal:
export HF_TOKEN=your_token
export API_BASE_URL=https://router.huggingface.co/v1
export MODEL_NAME=Qwen/Qwen2.5-72B-Instruct
export BASE_URL=http://localhost:8000
python inference.pyRunning with Docker
docker build -t code-security-env .
docker run -p 7860:7860 \
-e HF_TOKEN=your_token \
-e API_BASE_URL=https://router.huggingface.co/v1 \
-e MODEL_NAME=Qwen/Qwen2.5-72B-Instruct \
code-security-envBaseline Performance Scores
These scores were produced by running inference.py with Qwen/Qwen2.5-72B-Instruct via the HuggingFace router:
Scores are deterministic — graders are keyword-based with no randomness.
Reward Design
Partial progress: Each correctly identified vulnerability keyword adds to the reward incrementally — the agent gets credit even for partial analysis.
Penalties:
- Empty action → reward = 0.0, episode counts the step
- Repeated identical actions do not gain additional reward
- Exceeding max_steps (3) ends the episode with whatever score was last achieved
Project Structure
.
├── main.py # FastAPI server (reset/step/state endpoints)
├── tasks.py # Task definitions and grading logic
├── inference.py # Baseline inference script
├── openenv.yaml # OpenEnv spec
├── Dockerfile # Container config
├── requirements.txt
└── README.md