Team Ai
Apppublic

veerskywalker/RL_Security_Vul_Agent

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference

set up

$env:HFTOKEN="yourhftokenhere" $env:APIBASEURL="https://router.huggingface.co/v1" $env:MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"

note : we tested and and used Qwen/Qwen2.5-72B-Instruct

Security Vulnerability Detection Environment

An OpenEnv environment where an AI agent reviews real source-code files and identifies planted security vulnerabilities. The agent receives code, submits a structured vulnerability report, and is scored on accuracy — mimicking the real-world task of security code review.


Why This Environment?

Security code review is one of the most high-stakes tasks in software engineering. Human reviewers miss vulnerabilities due to fatigue, cognitive load, or simply the sheer volume of code. Training AI agents to reliably detect vulnerabilities — and evaluating how well they do it — has immediate practical value.

This environment fills a real gap: a reproducible, gradable benchmark for security-focused code analysis agents. Unlike CTF challenges (which are often contrived), every task here mirrors a vulnerability class that appears regularly in production codebases.


Environment Description

The agent is shown one or more Python source files containing a planted security vulnerability. The agent must:

  1. 1.Identify the vulnerability type (e.g. sql_injection, command_injection)
  2. 2.Name the affected files
  3. 3.Point to the line numbers where the vulnerability exists
  4. 4.Write a description of how the vulnerability can be exploited
  5. 5.Assign a severity level

The grader scores the submission against a hidden answer key. Partial credit is awarded for partially correct answers. The agent has up to 3 attempts per task, but later attempts receive a score penalty to incentivise getting it right first.


Action Space

The agent submits a SecurityVulnAction JSON object:

FieldTypeDescription
vulnerability_typestringOne of: sql_injection, xss, command_injection, ssrf, auth_bypass, path_traversal, insecure_deserialization, xxe, idor
affected_fileslist[string]Filenames that contain or contribute to the vulnerability
line_numberslist[int]Line numbers where the vulnerable code exists (±3 tolerance)
descriptionstringHow the vulnerability works and how it can be exploited
severitystringOne of: low, medium, high, critical

Example action:

json
{
  "vulnerability_type": "sql_injection",
  "affected_files": ["user_auth.py"],
  "line_numbers": [7, 16],
  "description": "User input is interpolated directly into the SQL query string using an f-string. An attacker can inject SQL syntax like ' OR '1'='1 to bypass authentication or dump the database.",
  "severity": "critical"
}

Observation Space

After each reset() or step(), the agent receives a SecurityVulnObservation:

FieldTypeDescription
task_idstringIdentifier of the current task
task_descriptionstringWhat the agent is being asked to find
code_filesdict[str, str]Map of filename → source code to analyse
feedbackstringGrader feedback from the last submission
attempts_remainingintHow many attempts the agent has left
doneboolWhether the episode is complete
rewardfloatScore for the last submission (0.0–1.0)

Tasks

Task 1 — sql_injection_basic (Easy)

Description: A single Python authentication module with two SQL queries built using unsafe string interpolation. The agent must identify the SQL injection vulnerability and point to the affected lines.

Why it's easy: The vulnerability is in one file, the pattern (f-string / string concatenation in SQL) is one of the most well-documented vulnerability classes, and the fix is well-known.

Expected agent score: 0.80 – 1.00


Task 2 — auth_bypass_jwt (Medium)

Description: A two-file Flask application with an admin panel. The JWT middleware accepts the none algorithm, and the route handler trusts the role claim from the token without server-side verification. An attacker can forge a token with alg: none and role: admin.

Why it's medium: The vulnerability spans two files. The agent must understand how JWT works, recognise the none algorithm attack, and connect the middleware flaw to the route-level role check.

Expected agent score: 0.50 – 0.85


Task 3 — vuln_chain_ssrf_cmdi (Hard)

Description: A three-file Python microservice. The vulnerability is a chain: user-controlled input flows into a subprocess.run(..., shell=True) call (command injection), and a separate function forwards HTTP requests to any caller-supplied URL without validation (SSRF). Combined, these allow an attacker to achieve remote code execution and probe the internal network.

Why it's hard: The vulnerability only exists as a chain across three files. The agent must trace data flow across module boundaries, identify both vulnerability classes, and explain how they combine into a single attack. Even frontier models frequently miss the cross-file interaction.

Expected agent score: 0.20 – 0.60


Reward Function

The grader scores each submission on four components:

ComponentWeightDescription
Vulnerability type40%Exact match = full credit. Same family = partial credit.
Affected files30%Proportional to overlap with expected files
Line numbers20%Credit for each line within ±3 of expected
Description keywords10%Proportion of key terms present in description

Attempt penalty — later attempts receive a multiplier:

Attempt 1 → score × 1.00
Attempt 2 → score × 0.85
Attempt 3 → score × 0.70

This gives a continuous signal over the full trajectory rather than a binary end-of-episode reward.


Baseline Scores

Baseline model: Qwen/Qwen2.5-72B-Instruct via HuggingFace Router

TaskDifficultyBaseline Score
sql_injection_basicEasy0.94
auth_bypass_jwtMedium0.83
vuln_chain_ssrf_cmdiHard0.68
Overall0.82

Setup & Usage

Option 1 — Use the live HuggingFace Space

The environment is deployed at:

https://VeerTheSani-security-vuln-env.hf.space

Test it immediately:

bash
curl -X POST https://VeerTheSani-security-vuln-env.hf.space/reset

Interactive API docs:

https://VeerTheSani-security-vuln-env.hf.space/docs

Option 2 — Run locally with Docker

bash
# Build
docker build -t security-vuln-env .

# Run
docker run -p 7860:7860 security-vuln-env

Server will be available at http://localhost:7860


Option 3 — Run locally with uv

bash
# Install dependencies
uv sync

# Start server
uv run server

Running the Inference Script

bash
# Install dependencies
pip install openai openenv-core

# Set credentials
export HF_TOKEN=your_huggingface_token
export API_BASE_URL=https://router.huggingface.co/v1
export MODEL_NAME=Qwen/Qwen2.5-72B-Instruct
export IMAGE_NAME=security-vuln-env:latest

# Run
python inference.py

Expected output:

[START] task=sql_injection_basic env=security_vuln model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=type=sql_injection files=['user_auth.py'] lines=[7, 16] reward=0.85 done=true error=null
[END] success=true steps=1 score=0.850 rewards=0.85

[START] task=auth_bypass_jwt env=security_vuln model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=type=auth_bypass files=['middleware.py', 'admin_routes.py'] lines=[15, 10] reward=0.62 done=true error=null
[END] success=true steps=1 score=0.620 rewards=0.62

[START] task=vuln_chain_ssrf_cmdi env=security_vuln model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=type=command_injection files=['report_api.py'] lines=[10] reward=0.31 done=false error=null
[END] success=false steps=3 score=0.310 rewards=0.31,0.22,0.19


---

## Environment Variables

| Variable | Required | Description |
|---|---|---|
| `HF_TOKEN` | Yes | HuggingFace API token for LLM calls |
| `API_BASE_URL` | No | LLM endpoint (default: HF router) |
| `MODEL_NAME` | No | Model identifier (default: Qwen2.5-72B) |
| `IMAGE_NAME` | No | Docker image name for local runs |
| `SECURITY_VULN_TASK` | No | Pin to a specific task ID for testing |

---

## OpenEnv Validation

pip install openenv-core openenv validate


All three checks should pass: