tek-wizard/devops-incident-responder
2
1# Training Safer Incident Copilots, Not Just Smarter Chatbots2 3At 2:13 AM, login failures spike, dashboards turn red, and every minute of delay costs real money and trust.4 5🚀 Try the Environment: `https://tek-wizard-devops-incident-responder.hf.space`6 7📊 Training Notebook: [devops_grpo_training.ipynb](devops_grpo_training.ipynb)8 9💻 Source Code: `https://github.com/tek-wizard/devops-incident-responder`10 11A language model can already *talk* about incidents. It can summarize logs, explain what a rollback is, and suggest generic debugging steps.12 13That is not the same as behaving well inside a live incident.14 15Real incident response is a long-horizon decision problem under pressure:16- signals are noisy17- the true cause is only partially visible18- several actions look plausible19- the wrong action order makes the outage worse20- success is not “sounding smart,” but restoring service safely21 22That is the gap our environment targets.23 24## The Problem25 26Today’s LLMs are strong at retrieval and explanation, but weak at **stateful operational reasoning**.27 28In production settings, that weakness is predictable. Untrained models jump to mitigations too early, recommend risky actions without enough evidence, lose track of what has already been checked, and often stop at plausible advice instead of verified recovery.29 30We built **DevOps Incident Responder** to train and evaluate a different capability:31 32**Can an LLM learn to act like a safer on-call copilot inside a realistic outage loop?**33 34Not a search engine.35Not a documentation bot.36Not a one-turn benchmark.37 38A sequential agent operating in a partially observable production world.39 40## Why This Matters41 42The near-term future of AI in operations is not fully autonomous production control.43 44It is **high-quality incident copilots**:45- systems that help engineers investigate faster46- systems that recommend safer next actions47- systems that preserve causal context across a multi-step incident48- systems that know when *not* to act49 50That is valuable for:51- SRE teams handling high-pressure outages52- platform teams building safe automation53- researchers studying long-horizon world models54- organizations that want AI assistance without blindly handing over production55 56If we want trustworthy operational AI, we need environments that reward **correct behavior over fluent behavior**.57 58## The Environment59 60DevOps Incident Responder simulates a small production microservice stack with realistic service interactions:61- `auth-api`62- `billing-api`63- `worker`64- `main-db`65- `cache`66- `queue`67 68The agent observes:69- service health70- latency and error metrics71- logs72- deploy history73- dependency signals74- discovered facts75- recent incident timeline76 77The agent acts through operational commands such as:78- `query_logs`79- `get_metrics`80- `get_recent_deploys`81- `inspect_config`82- `check_dependency`83- `rollback_deploy`84- `restart_service`85- `scale_service`86- `clear_connections`87- `disable_feature_flag`88- `drain_queue`89- `failover_db`90- `run_smoke_test`91- `post_status_update`92 93This is not a toy grid world. It is a compact model of the kind of reasoning human on-call engineers perform:94- investigate before mitigation95- infer hidden causes from partial evidence96- choose actions in the right order97- avoid unsafe interventions98- verify actual customer recovery99 100Under the hood, the environment is implemented as a seeded Python world model with typed OpenEnv observations, actions, rewards, and graders. Each command mutates the simulated production state through deterministic transitions, delayed effects, dependency changes, and scenario-specific failure logic. That lets us train and evaluate trajectory quality without spinning up real infrastructure, while still preserving the causal structure of real incidents.101 102## The Tasks103 104Our scenarios are designed around failure modes that feel operationally real:105- a bad auth deploy causing login failures106- a worker memory leak driving queue backlog and downstream latency107- database pool exhaustion requiring ordered recovery108- a feature flag storm amplifying cache and DB load109- a billing producer regression creating poison messages110- a risky failover scenario where acting too early is unsafe111 112These scenarios deliberately test capabilities that current LLM agents still struggle with:113- causal diagnosis114- long-horizon sequencing115- partial observability116- safety under uncertainty117 118## Why This Is Ambitious119 120The hackathon explicitly asks for environments that are ambitious, original, and actually useful for training.121 122This environment is ambitious in the right way:123 1241. It is not a static QA benchmark.125The agent must interact with a changing world.126 1272. It is not a one-turn reward hack.128Success depends on trajectories, not clever wording.129 1303. It combines two hard themes at once.131It is both **World Modeling** and **Long-Horizon Planning**.132 1334. It teaches a frontier-relevant capability.134Safe operational reasoning under uncertainty is underexplored and high impact.135 1365. It is difficult to game.137Plausible-looking but unsafe or out-of-order actions are penalized.138 139This is exactly the kind of environment where “good sounding” language is not enough.140 141## Reward Design142 143We designed the reward to teach behavior that matters in real incidents:144- positive reward for meaningful investigation145- higher reward for correct mitigations146- strong reward for verified recovery147- penalties for unsafe, irrelevant, or repeated actions148- penalties for wasting steps149 150The key principle is simple:151 152**The agent should not get high reward unless it actually behaves like a good incident responder.**153 154That means the reward is tied to:155- evidence gathering156- causal correctness157- safety158- recovery confirmation159- operational efficiency160 161This matters because many benchmarks are easy to exploit with fluent but shallow outputs. We wanted a benchmark where the model must *earn* success through behavior.162 163## What We Want the Agent to Learn164 165We are not trying to train a model that memorizes runbooks.166 167We are trying to train a model that improves on these axes:168- better first-action selection169- fewer unsafe actions170- stronger investigation coverage171- better action ordering172- higher solve rate on seeded incidents173 174That gives us a much more meaningful picture of progress than “did reward go up a bit?”175 176## Training Evidence177 178The core question for this environment is not whether the reward curve moves, but whether the agent’s behavior improves in a way an on-call engineer would actually trust.179 180In our evaluation, we compare:181- an untrained or zero-shot model182- simple baselines183- a trained model on the same seeded incidents184 185The most important signals are:186- average validation score187- solve rate188- first-action quality189- investigation coverage190- unsafe actions per episode191 192We also look at qualitative trajectory changes. A weak model often jumps straight to a risky or generic action like `restart_service` or picks the right command with the wrong target. A stronger model learns to investigate first, narrow the likely root cause, and only then apply the ordered mitigation that actually resolves the incident.193 194 195## Why OpenEnv Is the Right Fit196 197OpenEnv is valuable here because this problem is fundamentally interactive.198 199A static dataset cannot capture:200- counterfactual action choices201- environment transitions after mitigation202- delayed consequences of bad sequencing203- trajectory-level grading204 205This benchmark needs:206- resettable seeded incidents207- stepwise world transitions208- reproducible evaluation209- typed observations, actions, rewards, and graders210 211That is exactly the point of an environment.212 213## The Core Insight214 215There are already AI tools that can answer DevOps questions.216 217What is missing is a good environment for training and measuring whether an LLM can move from:218 219**“Here are some generic troubleshooting ideas”**220 221to222 223**“Given this evolving incident state, here is the safest and most useful next action.”**224 225That is the jump from chatbot intelligence to agent intelligence.226 227## What Makes This Useful Beyond The Hackathon228 229Even outside this competition, this environment can serve as:230- a benchmark for incident-response copilots231- a training simulator for junior SRE workflows232- a safe offline sandbox for evaluating operational AI233- a research environment for long-horizon causal reasoning234 235In other words, this is not just a demo.236 237It is a practical testbed for one of the most important open questions in applied AI:238 239**Can language agents make better decisions in high-stakes, partially observable real-world workflows?**240 241## Closing242 243We believe the most important AI systems of the next few years will not win by being more verbose.244 245They will win by being:246- safer247- more stateful248- more causally grounded249- better at acting under uncertainty250 251DevOps Incident Responder is our attempt to build exactly the kind of environment needed to train that behavior.252 253Because in a real incident, the goal is not to sound helpful.254 255The goal is to restore service without making the outage worse.256 