Rudransh12121/incident_debugging_env
๐ก๏ธ OpenEnv: Incident Debugging Environment
A production-grade RL environment for training and evaluating AI agents on real-world SRE incident triage.
  
๐ฏ What This Is
This is an OpenEnv-compliant reinforcement learning environment that simulates real-world production incident debugging โ a task performed daily by Site Reliability Engineers at companies like Google, Meta, and Netflix.
An AI agent is dropped into a simulated production outage with noisy system logs. It must iteratively analyze telemetry, filter noise from signal, identify the root cause, and propose a fix โ just like a real SRE during a 3 AM pager alert.
This is not a game or a toy benchmark. It models the exact cognitive workflow that human SREs perform: observe logs โ form hypothesis โ refine diagnosis โ propose remediation.
๐ง Why This Matters
๐ Live Environment
Quick verification:
curl -X POST https://rudransh12121-incident-debugging-env.hf.space/reset?level=medium๐ OpenEnv Spec Compliance
This environment implements the full OpenEnv specification:
๐๏ธ Architecture
scaler/
โโโ server/
โ โโโ app.py # FastAPI server, dashboard, grader endpoints
โ โโโ environment.py # Core RL environment (reset/step/state)
โ โโโ log_generator.py # Procedural telemetry engine with noise interleaving
โ โโโ scorer.py # Hybrid F1 + semantic scoring with shaped rewards
โ โโโ parser.py # Structured text parser with synonym normalization
โ โโโ tasks.py # Gold-standard scenario definitions
โ โโโ models.py # Pydantic typed models (Observation, Action, State)
โโโ inference.py # Baseline agent โ runs all 3 tasks via OpenAI API
โโโ openenv.yaml # OpenEnv spec (v1) with task/grader config
โโโ outputs/ # Sample baseline run outputs
โโโ docs/ # Technical documentation
โโโ Dockerfile # Production container
โโโ tests/ # Validation scriptsData Flow
Agent โ POST /reset?task_id=easy
โ
LogGenerator produces procedural logs (seeded)
Environment initializes State(step=0, best_score=0)
โ
Returns: Observation(logs=[...], context="...")
โ
Agent โ POST /step { "raw_text": "ROOT_CAUSE: ..." }
โ
Parser extracts: ROOT_CAUSE, FACTORS, FIX, SEVERITY
Scorer computes: F1 + semantic similarity vs gold labels
Reward = max(0, improvement - penalties)
โ
Returns: { observation, reward, done, info: { score, components } }
โ
Agent refines diagnosis based on feedback...
(repeat until done=true or max_steps reached)๐ฎ Action Space
Agents submit structured diagnostic text:
Action(raw_text: str)Required format:
ROOT_CAUSE: <primary root cause of the incident>
FACTORS: <contributing factors, comma separated>
FIX: <recommended remediation actions>
SEVERITY: <incident severity level>๐๏ธ Observation Space
Progressive disclosure โ agents see more data as they explore:
Observation(
logs: List[str], # System telemetry (revealed progressively)
context: str, # Incident context description
step: int, # Current step in episode
last_reward: float, # Reward from previous action
feedback: str # Natural language diagnostic hint
)Key mechanic: Early steps reveal logs one at a time. Later steps switch to reasoning hints, forcing agents to work from memory and inference rather than raw data.
๐ State
State(
step: int, # Current step
best_score: float, # Highest score achieved so far
cumulative_reward: float, # Total reward accumulated
done: bool # Episode termination flag
)โ๏ธ Reward Function
Scoring Components
Each action is scored across 4 diagnostic dimensions:
Hybrid Scoring
Each component uses the maximum of two scoring methods:
score = max(token_f1, semantic_similarity * 0.85)This ensures agents aren't penalized for paraphrasing (e.g., "disk ran out of space" vs "disk space full").
Shaped Rewards (Not Binary!)
reward = max(0, score_improvement - penalties)Rewards are incremental โ agents only earn reward for discovering NEW information:
Score Calibration
All scores are clamped to the open interval (0.01, 0.99) at two layers (scorer + endpoint) to satisfy platform constraints.
๐ Task Difficulty Levels
Why Difficulty Increases
- Episode length: Easy (5 steps) โ Hard (15 steps) โ harder tasks need more exploration
- Log volume: More logs means more data to process
- Noise placement: Easy has noise clearly at the end. Medium inserts noise randomly. Hard shuffles everything โ the agent must identify signal from a jumbled timeline
- Causal depth: Hard tasks separate the cause (memory leak) from the symptom (kernel OOM panic) by multiple inference steps
๐ง Setup Instructions
Run Locally
pip install -r requirements.txt
uvicorn server.app:app --host 0.0.0.0 --port 7860Docker
docker build -t incident-env .
docker run -p 7860:7860 incident-envRun Baseline Inference
# Mock mode (no API key needed):
python inference.py
# With LLM:
OPENAI_API_KEY=<your_key> python inference.pyValidate OpenEnv Contract
openenv validate --url http://localhost:7860 --verboseDeploy to Hugging Face
huggingface-cli login
openenv push๐งช Example Interaction
Step 1 โ Reset:
curl -X POST "http://localhost:7860/reset?level=medium&seed=42"Step 2 โ Submit diagnosis:
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{"raw_text": "ROOT_CAUSE: database deadlock migration cascade\nFACTORS: lock wait timeout, migration sql\nFIX: restart and optimize\nSEVERITY: total outage"}'Response:
{
"observation": {
"logs": [],
"step": 1,
"feedback": "DIAGNOSTIC ACCURACY IMPROVED.",
"last_reward": 0.42
},
"reward": 0.42,
"done": false,
"info": {
"score": 0.78,
"best_score": 0.78,
"step": 1,
"max_steps": 10,
"components": {"rc": 0.92, "f": 0.67, "fix": 0.50, "sev": 1.0}
}
}