Team Ai
Apppublic

Rudransh12121/incident_debugging_env

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

๐Ÿ›ก๏ธ OpenEnv: Incident Debugging Environment

A production-grade RL environment for training and evaluating AI agents on real-world SRE incident triage.

![OpenEnv](https://github.com/openenv/spec) ![Deployment](https://huggingface.co/spaces/Rudransh12121/incidentdebuggingenv) ![Docker](https://hub.docker.com)


๐ŸŽฏ What This Is

This is an OpenEnv-compliant reinforcement learning environment that simulates real-world production incident debugging โ€” a task performed daily by Site Reliability Engineers at companies like Google, Meta, and Netflix.

An AI agent is dropped into a simulated production outage with noisy system logs. It must iteratively analyze telemetry, filter noise from signal, identify the root cause, and propose a fix โ€” just like a real SRE during a 3 AM pager alert.

This is not a game or a toy benchmark. It models the exact cognitive workflow that human SREs perform: observe logs โ†’ form hypothesis โ†’ refine diagnosis โ†’ propose remediation.

๐Ÿง  Why This Matters

ProblemHow We Solve It
Static benchmarks can be memorizedProcedurally generated logs with seeded randomness ensure infinite unique scenarios
Binary pass/fail doesn't capture partial understandingMulti-component weighted scoring rewards incremental diagnostic progress
Real incidents have noiseDifficulty-scaled noise interleaving (appended โ†’ randomly inserted โ†’ fully shuffled)
Agents game metrics with keyword stuffingSpam penalties, noise penalties, and repetition penalties discourage gaming
LLMs paraphrase correct answers differentlyHybrid token-F1 + semantic similarity scoring handles varied phrasing

๐Ÿš€ Live Environment

ResourceLink
Hugging Face SpaceRudransh12121/incident_debugging_env
Interactive Dashboardrudransh12121-incident-debugging-env.hf.space
API Docs (Swagger)/docs

Quick verification:

bash
curl -X POST https://rudransh12121-incident-debugging-env.hf.space/reset?level=medium

๐Ÿ“ OpenEnv Spec Compliance

This environment implements the full OpenEnv specification:

RequirementStatusImplementation
Typed Observation modelโœ…Pydantic model with logs, context, step, last_reward, feedback
Typed Action modelโœ…Pydantic model with raw_text
Typed State modelโœ…Pydantic model with step, best_score, cumulative_reward, done
POST /reset โ†’ observationโœ…Accepts task_id, level, seed params
POST /step โ†’ obs, reward, done, infoโœ…Returns shaped reward + component breakdown
GET /state โ†’ current stateโœ…Returns internal episode state
openenv.yamlโœ…spec_version: 1, 3 tasks, per-task graders
3+ tasks with gradersโœ…/grade/easy, /grade/medium, /grade/hard (GET+POST)
Scores in (0, 1)โœ…Double-clamped: scorer + endpoint level
openenv validateโœ…Passes all checks

๐Ÿ—๏ธ Architecture

text
scaler/
โ”œโ”€โ”€ server/
โ”‚   โ”œโ”€โ”€ app.py              # FastAPI server, dashboard, grader endpoints
โ”‚   โ”œโ”€โ”€ environment.py      # Core RL environment (reset/step/state)
โ”‚   โ”œโ”€โ”€ log_generator.py    # Procedural telemetry engine with noise interleaving
โ”‚   โ”œโ”€โ”€ scorer.py           # Hybrid F1 + semantic scoring with shaped rewards
โ”‚   โ”œโ”€โ”€ parser.py           # Structured text parser with synonym normalization
โ”‚   โ”œโ”€โ”€ tasks.py            # Gold-standard scenario definitions
โ”‚   โ””โ”€โ”€ models.py           # Pydantic typed models (Observation, Action, State)
โ”œโ”€โ”€ inference.py            # Baseline agent โ€” runs all 3 tasks via OpenAI API
โ”œโ”€โ”€ openenv.yaml            # OpenEnv spec (v1) with task/grader config
โ”œโ”€โ”€ outputs/                # Sample baseline run outputs
โ”œโ”€โ”€ docs/                   # Technical documentation
โ”œโ”€โ”€ Dockerfile              # Production container
โ””โ”€โ”€ tests/                  # Validation scripts

Data Flow

Agent โ†’ POST /reset?task_id=easy
                โ†“
        LogGenerator produces procedural logs (seeded)
        Environment initializes State(step=0, best_score=0)
                โ†“
        Returns: Observation(logs=[...], context="...")
                โ†“
Agent โ†’ POST /step { "raw_text": "ROOT_CAUSE: ..." }
                โ†“
        Parser extracts: ROOT_CAUSE, FACTORS, FIX, SEVERITY
        Scorer computes: F1 + semantic similarity vs gold labels
        Reward = max(0, improvement - penalties)
                โ†“
        Returns: { observation, reward, done, info: { score, components } }
                โ†“
        Agent refines diagnosis based on feedback...
        (repeat until done=true or max_steps reached)

๐ŸŽฎ Action Space

Agents submit structured diagnostic text:

python
Action(raw_text: str)

Required format:

text
ROOT_CAUSE: <primary root cause of the incident>
FACTORS: <contributing factors, comma separated>
FIX: <recommended remediation actions>
SEVERITY: <incident severity level>

๐Ÿ‘๏ธ Observation Space

Progressive disclosure โ€” agents see more data as they explore:

python
Observation(
    logs: List[str],       # System telemetry (revealed progressively)
    context: str,          # Incident context description
    step: int,             # Current step in episode
    last_reward: float,    # Reward from previous action
    feedback: str          # Natural language diagnostic hint
)

Key mechanic: Early steps reveal logs one at a time. Later steps switch to reasoning hints, forcing agents to work from memory and inference rather than raw data.

๐Ÿ“Š State

python
State(
    step: int,                # Current step
    best_score: float,        # Highest score achieved so far
    cumulative_reward: float,  # Total reward accumulated
    done: bool                # Episode termination flag
)

โš–๏ธ Reward Function

Scoring Components

Each action is scored across 4 diagnostic dimensions:

ComponentWeightMethodWhy This Weight
Root Cause40%Token F1 + Semantic SimilarityMost critical โ€” identifying the cause is the primary task
Contributing Factors20%Token F1 + Semantic SimilaritySupporting evidence validates the diagnosis
Fix Category30%Keyword Category CoverageCorrect remediation is operationally critical
Severity10%Exact MatchClassification accuracy matters but is secondary

Hybrid Scoring

Each component uses the maximum of two scoring methods:

python
score = max(token_f1, semantic_similarity * 0.85)

This ensures agents aren't penalized for paraphrasing (e.g., "disk ran out of space" vs "disk space full").

Shaped Rewards (Not Binary!)

python
reward = max(0, score_improvement - penalties)

Rewards are incremental โ€” agents only earn reward for discovering NEW information:

SignalTypePurpose
delta > 0Positive rewardDiagnostic improvement
delta > 0.1 at step โ‰ค 3Early convergence bonus (+0.03)Fast solvers rewarded
Same answer, different phrasingExploration bonus (0.01)Encourages hypothesis variation
delta = 0, same actionZero rewardNo credit for repetition
Excessive tokensSpam penalty (-0.05)Prevents log-dumping
Irrelevant tokensNoise penalty (-0.02/token)Penalizes hallucination
Repeated action (>95% similar)Repetition penalty (-0.10)Prevents infinite loops

Score Calibration

All scores are clamped to the open interval (0.01, 0.99) at two layers (scorer + endpoint) to satisfy platform constraints.


๐Ÿ“‹ Task Difficulty Levels

TaskScenarioMax StepsLog VolumeNoise StrategyGold Root Cause
EasyDisk Saturation54-6 logsAppended at endCleanup script disk saturation
MediumDatabase Deadlock106-10 logsRandomly interleavedDatabase deadlock migration cascade
HardMemory Leak / OOM1510-15 logsFully shuffledMemory leak cache large payload OOM

Why Difficulty Increases

  1. 1.Episode length: Easy (5 steps) โ†’ Hard (15 steps) โ€” harder tasks need more exploration
  2. 2.Log volume: More logs means more data to process
  3. 3.Noise placement: Easy has noise clearly at the end. Medium inserts noise randomly. Hard shuffles everything โ€” the agent must identify signal from a jumbled timeline
  4. 4.Causal depth: Hard tasks separate the cause (memory leak) from the symptom (kernel OOM panic) by multiple inference steps

๐Ÿ”ง Setup Instructions

Run Locally

bash
pip install -r requirements.txt
uvicorn server.app:app --host 0.0.0.0 --port 7860

Docker

bash
docker build -t incident-env .
docker run -p 7860:7860 incident-env

Run Baseline Inference

bash
# Mock mode (no API key needed):
python inference.py

# With LLM:
OPENAI_API_KEY=<your_key> python inference.py

Validate OpenEnv Contract

bash
openenv validate --url http://localhost:7860 --verbose

Deploy to Hugging Face

bash
huggingface-cli login
openenv push

๐Ÿงช Example Interaction

Step 1 โ€” Reset:

bash
curl -X POST "http://localhost:7860/reset?level=medium&seed=42"

Step 2 โ€” Submit diagnosis:

bash
curl -X POST http://localhost:7860/step \
  -H "Content-Type: application/json" \
  -d '{"raw_text": "ROOT_CAUSE: database deadlock migration cascade\nFACTORS: lock wait timeout, migration sql\nFIX: restart and optimize\nSEVERITY: total outage"}'

Response:

json
{
  "observation": {
    "logs": [],
    "step": 1,
    "feedback": "DIAGNOSTIC ACCURACY IMPROVED.",
    "last_reward": 0.42
  },
  "reward": 0.42,
  "done": false,
  "info": {
    "score": 0.78,
    "best_score": 0.78,
    "step": 1,
    "max_steps": 10,
    "components": {"rc": 0.92, "f": 0.67, "fix": 0.50, "sev": 1.0}
  }
}

๐Ÿ“š Documentation

GuideDescription
OverviewProblem statement and real-world utility
ArchitectureSystem design, data flow, and state management
Task LibraryDetailed breakdown of all 3 incident scenarios
Scoring LogicReward math, penalty strategies, and calibration
API ReferenceEndpoint schemas, examples, and grader routes
Design DecisionsRationale behind key architectural choices
EvaluationBenchmarking methodology and expected score ranges

๐Ÿ† Key Differentiators

FeatureDescription
Hybrid ScoringCombines token-level F1 with semantic similarity โ€” handles LLM paraphrasing gracefully
Difficulty-Scaled NoiseEasy: noise at end. Medium: randomly interleaved. Hard: fully shuffled timeline
Anti-Gaming PenaltiesSpam, noise, and repetition penalties prevent keyword stuffing and infinite loops
Progressive DisclosureEarly steps reveal logs; later steps provide reasoning hints โ€” mimics real SRE workflows
Deterministic ReproductionSeeded log generation enables exact reproducibility for research benchmarking
Per-Task Episode Length5/10/15 steps for easy/medium/hard โ€” difficulty controls both complexity AND time budget
Defense-in-Depth ScoringScores clamped at both scorer and endpoint level โ€” zero risk of out-of-range failures

โš™๏ธ Environment Variables

VariableDefaultDescription
OPENAI_API_KEYโ€”API key for LLM inference (or HF_TOKEN)
API_BASE_URLhttps://router.huggingface.co/v1LLM API base URL
MODEL_NAMEQwen/Qwen2.5-72B-InstructModel for baseline inference
ENV_URLhttp://127.0.0.1:7860Environment server URL