veerskywalker/RL_Security_Vul_Agent
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
set up
$env:HFTOKEN="yourhftokenhere" $env:APIBASEURL="https://router.huggingface.co/v1" $env:MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
note : we tested and and used Qwen/Qwen2.5-72B-Instruct
Security Vulnerability Detection Environment
An OpenEnv environment where an AI agent reviews real source-code files and identifies planted security vulnerabilities. The agent receives code, submits a structured vulnerability report, and is scored on accuracy — mimicking the real-world task of security code review.
Why This Environment?
Security code review is one of the most high-stakes tasks in software engineering. Human reviewers miss vulnerabilities due to fatigue, cognitive load, or simply the sheer volume of code. Training AI agents to reliably detect vulnerabilities — and evaluating how well they do it — has immediate practical value.
This environment fills a real gap: a reproducible, gradable benchmark for security-focused code analysis agents. Unlike CTF challenges (which are often contrived), every task here mirrors a vulnerability class that appears regularly in production codebases.
Environment Description
The agent is shown one or more Python source files containing a planted security vulnerability. The agent must:
- Identify the vulnerability type (e.g.
sql_injection,command_injection) - Name the affected files
- Point to the line numbers where the vulnerability exists
- Write a description of how the vulnerability can be exploited
- Assign a severity level
The grader scores the submission against a hidden answer key. Partial credit is awarded for partially correct answers. The agent has up to 3 attempts per task, but later attempts receive a score penalty to incentivise getting it right first.
Action Space
The agent submits a SecurityVulnAction JSON object:
Example action:
{
"vulnerability_type": "sql_injection",
"affected_files": ["user_auth.py"],
"line_numbers": [7, 16],
"description": "User input is interpolated directly into the SQL query string using an f-string. An attacker can inject SQL syntax like ' OR '1'='1 to bypass authentication or dump the database.",
"severity": "critical"
}Observation Space
After each reset() or step(), the agent receives a SecurityVulnObservation:
Tasks
Task 1 — sql_injection_basic (Easy)
Description: A single Python authentication module with two SQL queries built using unsafe string interpolation. The agent must identify the SQL injection vulnerability and point to the affected lines.
Why it's easy: The vulnerability is in one file, the pattern (f-string / string concatenation in SQL) is one of the most well-documented vulnerability classes, and the fix is well-known.
Expected agent score: 0.80 – 1.00
Task 2 — auth_bypass_jwt (Medium)
Description: A two-file Flask application with an admin panel. The JWT middleware accepts the none algorithm, and the route handler trusts the role claim from the token without server-side verification. An attacker can forge a token with alg: none and role: admin.
Why it's medium: The vulnerability spans two files. The agent must understand how JWT works, recognise the none algorithm attack, and connect the middleware flaw to the route-level role check.
Expected agent score: 0.50 – 0.85
Task 3 — vuln_chain_ssrf_cmdi (Hard)
Description: A three-file Python microservice. The vulnerability is a chain: user-controlled input flows into a subprocess.run(..., shell=True) call (command injection), and a separate function forwards HTTP requests to any caller-supplied URL without validation (SSRF). Combined, these allow an attacker to achieve remote code execution and probe the internal network.
Why it's hard: The vulnerability only exists as a chain across three files. The agent must trace data flow across module boundaries, identify both vulnerability classes, and explain how they combine into a single attack. Even frontier models frequently miss the cross-file interaction.
Expected agent score: 0.20 – 0.60
Reward Function
The grader scores each submission on four components:
Attempt penalty — later attempts receive a multiplier:
Attempt 1 → score × 1.00
Attempt 2 → score × 0.85
Attempt 3 → score × 0.70This gives a continuous signal over the full trajectory rather than a binary end-of-episode reward.
Baseline Scores
Baseline model: Qwen/Qwen2.5-72B-Instruct via HuggingFace Router
Setup & Usage
Option 1 — Use the live HuggingFace Space
The environment is deployed at:
https://VeerTheSani-security-vuln-env.hf.spaceTest it immediately:
curl -X POST https://VeerTheSani-security-vuln-env.hf.space/resetInteractive API docs:
https://VeerTheSani-security-vuln-env.hf.space/docsOption 2 — Run locally with Docker
# Build
docker build -t security-vuln-env .
# Run
docker run -p 7860:7860 security-vuln-envServer will be available at http://localhost:7860
Option 3 — Run locally with uv
# Install dependencies
uv sync
# Start server
uv run serverRunning the Inference Script
# Install dependencies
pip install openai openenv-core
# Set credentials
export HF_TOKEN=your_huggingface_token
export API_BASE_URL=https://router.huggingface.co/v1
export MODEL_NAME=Qwen/Qwen2.5-72B-Instruct
export IMAGE_NAME=security-vuln-env:latest
# Run
python inference.pyExpected output:
[START] task=sql_injection_basic env=security_vuln model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=type=sql_injection files=['user_auth.py'] lines=[7, 16] reward=0.85 done=true error=null
[END] success=true steps=1 score=0.850 rewards=0.85
[START] task=auth_bypass_jwt env=security_vuln model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=type=auth_bypass files=['middleware.py', 'admin_routes.py'] lines=[15, 10] reward=0.62 done=true error=null
[END] success=true steps=1 score=0.620 rewards=0.62
[START] task=vuln_chain_ssrf_cmdi env=security_vuln model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=type=command_injection files=['report_api.py'] lines=[10] reward=0.31 done=false error=null
[END] success=false steps=3 score=0.310 rewards=0.31,0.22,0.19
---
## Environment Variables
| Variable | Required | Description |
|---|---|---|
| `HF_TOKEN` | Yes | HuggingFace API token for LLM calls |
| `API_BASE_URL` | No | LLM endpoint (default: HF router) |
| `MODEL_NAME` | No | Model identifier (default: Qwen2.5-72B) |
| `IMAGE_NAME` | No | Docker image name for local runs |
| `SECURITY_VULN_TASK` | No | Pin to a specific task ID for testing |
---
## OpenEnv Validation
pip install openenv-core openenv validate
All three checks should pass: