openenvforge/3ambench
3amBench: can your agent write alerts that page the right human at 3 a.m., and only then? 3amBench (package alertforge) is an RL environment and benchmark for a job SRE teams do every week: owning Prometheus alerting rules and Alertmanager routing as code. Each task drops the agent into a realistic monitoring/ repo for a fictional company with a handful of change requests: onboard a service onto multi-window burn-rate SLO alerts, fix the alert that paged 30 times last night… See the full description on the dataset page: https://huggingface.co/datasets/openenvforge/3ambench.
3amBench: can your agent write alerts that page the right human at 3 a.m., and only then?
3amBench (package alertforge) is an RL environment and benchmark for a job SRE teams do every week: owning Prometheus alerting rules and Alertmanager routing as code. Each task drops the agent into a realistic monitoring/ repo for a fictional company with a handful of change requests: onboard a service onto multi-window burn-rate SLO alerts, fix the alert that paged 30 times last night, find out why nobody was paged when a service was down for 47 minutes, route a newly formed team, add inhibition so an outage produces one page instead of forty.
Grading is behavioral. Nothing is judged by an LLM, and nothing is compared as text. Hidden, seed-varied outage replays run through Prometheus' own promtool test rules, and routing runs through Alertmanager's own amtool. Every alert must fire during the outage, stay silent during normal traffic, short spikes, recovery, low traffic and near-threshold traffic, fire on time (not before for elapses), fire once per service (not per pod), carry the right labels, and reach the right receivers. A task has 59-189 atomic checks, so the reward is dense.
Quick start
# oracle (expect mean 1.0) and no-op (expect 0.0)
uvx harbor run --repo https://huggingface.co/datasets/openenvforge/3ambench@v0.2.0 -d 3ambench@0.2.0 -a oracle -n 4
uvx harbor run --repo https://huggingface.co/datasets/openenvforge/3ambench@v0.2.0 -d 3ambench@0.2.0 -a nop -n 4
# your model on the 5-task smoke set
uvx harbor run --repo https://huggingface.co/datasets/openenvforge/3ambench@v0.2.0 -d 3ambench-mini@0.2.0 -a terminus-2 -m <model>
# expert tier: the 15 leaderboard tasks (all 45: -d 3ambench-expert@0.2.0)
uvx harbor run --repo https://huggingface.co/datasets/openenvforge/3ambench@v0.2.0 -d 3ambench-expert-lb@0.2.0 -a terminus-2 -m <model>The core tasks in 3ambench@0.2.0 are the 3ambench@0.1.1 tasks with their fictional company domains moved to reserved .example names (and the version stamp); checks, scenarios and grader are unchanged, and every recorded v0.1.1 run regrades to the same reward, so v0.1.1 numbers stay comparable.
Tasks set [agent] network_mode = "no-network", because this repo ships the solutions (see Anti-hacking). Harbor enforces that on Linux Docker hosts and sandboxed providers. Docker Desktop on macOS cannot enforce it (its VM kernel lacks CONFIG_NFT_FIB_INET), and Harbor refuses the task. For local smoke runs on a Mac, use scripts/harbor_local.sh, which runs a temp copy without the network override. Those runs are not leaderboard-safe.
OpenEnv (multi-turn, per-step reward), against a local server started as in Try it. The client needs pip install "openenv-core==0.3.0" and a clone of this repo on PYTHONPATH=openenv:
from alertforge_env.client import AlertForgeEnv
with AlertForgeEnv(base_url="http://localhost:8000").sync() as env:
obs = env.reset(seed=7, split="train", workflow="missed-page-postmortem", tier="medium")
obs = env.step({"tool": "read_file", "path": "README.md"})
obs = env.step({"tool": "write_file", "path": "rules/service-health.yml", "content": "..."})
print(obs.reward, obs.done) # parity mode: rewards sum to the Harbor rewardFresh tasks (training, private held-out sets): python generator/generate.py --master-seed <secret> --out <dir> (needs promtool and amtool; without --out it overwrites this repo's tasks).
Try it
- Replay runs in the browser: 3amBench Replay steps through recorded runs (tool calls, file diffs, the reward curve, which checks pass after each step, and for expert runs which tickets were solved). Its data comes from
scripts/record_reference_runs.py,scripts/record_agent_runs.shandscripts/export_runs.py; the page itself is inspace/. - Run the environment locally and open the web interface at http://localhost:8000/web:
docker run -p 8000:8000 -e ENABLE_WEB_INTERFACE=true ghcr.io/devesh-maheshwari/3ambench-env:0.2.0 The image is built for linux/amd64 and linux/arm64. To build it yourself from a clone: docker build -f openenv/alertforge_env/server/Dockerfile -t 3ambench-env . (then run 3ambench-env).
Leaderboard (core, 6 tasks: af-001, af-009, af-013, af-017, af-021, af-029)
<!-- LEADERBOARD:core --> | Agent | Model | Runs | Mean reward | Solved | |---|---|---|---|---| | claude-code | claude-opus-5-5 | 12 | 1.000 | 12/12 | | codex | gpt-6-astra | 12 | 1.000 | 12/12 | | terminus-2 | gpt-oss-120b | 13 | 0.436 | 3/13 |
37 runs on 6 tasks. <!-- /LEADERBOARD -->
The runs were recorded on the v0.1.1 tasks; their saved workspaces, read with the renamed domains, regrade to the same rewards on the 0.2.0 tasks. Local Harbor runs on Docker Desktop (agent network enabled, so not sealed). Frontier agents saturate this subset; the dense reward separates weaker models (gpt-oss-120b scores between 0.03 and 0.37 on every unsolved run instead of 0). v0.2 adds the expert tier below. Every run can be replayed step by step in the replay Space.
Expert tier (v0.2)
A core task is a short list of change requests on a small repo. An expert task is a week of a team's observability queue on a large legacy one: 45 tasks, five families of nine seeds, each at a different fictional company with its own services, teams and history. The starting repo has 62-85 alerting rules in 14-22 files, SLI recording rules, an SLO catalog, an ownership file, ADRs, runbooks, a CI script, and an Alertmanager config shared with two other Prometheus servers (41-53 receivers, legacy match: routes next to matchers:). The queue holds two or three symptom tickets and two to seven routine requests.
What makes it realistic:
- Tickets read like tickets. A symptom ticket says what happened in the reporter's words ("17 pages for one database failover", "KafkaConsumerStalled has never fired") and points at an incident packet: PagerDuty and Slack exports, a postmortem,
query_rangepulls of the night thatbin/range2testturns into promtool tests. It never says which rule is wrong. - Policy lives in the repo, not in the ticket: SLOs and windows in the catalog, the paging policy in the README, ownership in
teams/ownership.yaml, the burn-rate policy in an ADR. A new hire would have to look there too. - The repo has a past: two routing syntaxes, a team split half done, rules that moved with their team, decoy alerts that look broken and aren't, alerts other teams own.
- No two tasks share a company, an org chart or a queue, so a fix can't be copied from a sibling task.
Grading is behavioral and dense, as in the core tier. Promtool replays every hidden scenario group once (plus a re-run where a deadline window needs one) and the grader reads which alerts fire, minute by minute; routing goes through amtool and inhibition through the strict Alertmanager simulator. A symptom ticket is an outcome check: when the incident replays, the right page reaches the right team, and it stays quiet on normal traffic, deploys, flaps and the noise that started the ticket. Alerts are counted after inhibition, so muting everything doesn't make a quiet pager look fixed. Routine requests (routes, inhibitions, label moves, burn alerts, recording rules, annotations, legacy removals) are checked the same way. Untouched rules, receivers and routes have to keep working; a rewritten but equivalent rule is replayed, not compared as text. Each requirement scores s in [-0.25, 1] against the starting repo, and reward = 0.6 · solved + 0.4 · progress as in Reward; expert progress also loses 0.75 weight units for each leave-as-is alert the run changes, and solved requires all of them intact. A symptom ticket weighs 5-10 against 2-3 for a routine request, so the tickets carry 55-82% of a task's weight. diagnosis is the weighted mean s over the symptom tickets and routine the same over the rest: among runs that don't solve a task, they tell finding the cause apart from doing the easy part of the queue. Expert tasks have 128-232 checks and the same anti-hacking defenses, plus tamper checks on reads of ALERTS.
Calibration, honestly: frontier agents solve most expert tasks, in minutes (3-5 minutes per task in our pilot runs). The tier is not built to rank frontier models against each other. It is aimed at separating mid-size and open models, where the dense reward and the diagnosis/routine split show how far a run got, and at RL training on a realistic repo. A harder frontier tier is in development; there is no frontier-hardness result yet.
Leaderboard (v0.2.0, expert: the 15 tasks of 3ambench-expert-lb@0.2.0, 3 per family)
<!-- LEADERBOARD:expert --> | Agent | Model | Runs | Tasks | Mean reward | Solved | Diagnosis | Routine | |---|---|---|---|---|---|---|---| | claude-code | claude-opus-5-5 | 14 | 14/15 | 1.000 | 14/14 | 1.000 | 1.000 | | codex | gpt-6-astra | 15 | 15/15 | 0.960 | 14/15 | 1.000 | 1.000 | | claude-code | claude-fable-5-1 | 14 | 14/15 | 0.956 | 13/14 | 1.000 | 1.000 | | codex | gpt-5.6-sol | 15 | 15/15 | 0.912 | 13/15 | 0.967 | 1.000 | | claude-code | claude-sonnet-5 | 14 | 14/15 | 0.762 | 9/14 | 0.903 | 1.000 | | claude-code | claude-haiku-4-5-20251001 | 14 | 14/15 | 0.346 | 2/14 | 0.473 | 0.927 | | terminus-2 | gpt-oss-120b | 10 | 10/15 | 0.065 | 0/10 | 0.071 | 0.345 |
96 runs on 15 tasks; 2 runs flagged invalid are not counted. <!-- /LEADERBOARD -->
The 15 tasks are 3ambench-expert-lb@0.2.0: the development pilot set (3 per family, scripts/expert_calibrate.py --pilot), not a calibrated or held-out selection. These are unbalanced local development runs on Docker Desktop (agent network enabled, so not sealed), not a held-out model comparison, and the rows cover different tasks (the Tasks column), so compare them with care. The two historical afx-e5-s08 trials (Opus and Astra) are excluded because their survey evidence differs from the current task. Astra and Sol were re-run on the current task, so they cover all 15 tasks; Opus, Fable, Sonnet and Haiku have no valid afx-e5-s08 run yet and cover 14. gpt-oss-120b covers 10: the provider budget cap stopped it before afx-e2-s03, afx-e3-s04, afx-e4-s04 and afx-e5-s03, and it has no afx-e5-s08 run. Every saved output was regraded with the released task's own grader. Six runs score higher than Harbor's verifier gave them at the time, because of the grader fixes listed in the CHANGELOG; the final 0.2.0 changes (reserved domains, ingress series, the two grader fixes) change no recorded score. Infrastructure failures are excluded when present. Every run, with the tickets it solved, can be replayed in the replay Space.
To build the 45 tasks yourself: pip install . from a clone, then alertforge expert --out <dir> --seeds 1-9 (promtool and amtool on the path, or in AF_PROMTOOL and AF_AMTOOL). It runs each task's acceptance gate as well, and it reproduces the released tasks byte for byte.
See calibration status for the calibration protocol and what is still open.
What a task looks like
af-001-slo-onboarding-easy-s1 (instruction excerpt; change requests are shuffled):
- P1.HighErrorRatiois misbehaving; seepostmortems/PM-4043.md. Fix the rule, keeping its purpose. - I1. WhileChatGwDownis firing for a service, suppressChatGwErrorBudgetBurnFast,HighErrorRatiofor that same service only. - A1. AddChatGwErrorBudgetBurnFastforchat-gw(SLO 0.995): fire when bothslo:sli_error:ratio_rate1handslo:sli_error:ratio_rate5mexceed 14.4 × (1 − 0.995),for: 2m, ... - R1. Teampaymentsis now on call forchat-gw, but nothing routes to them yet. ... - C1. Add SLI recording rulesslo:sli_error:ratio_rate5m,slo:sli_error:ratio_rate1hfor thegrpcSLI family ...
The postmortem describes a symptom ("pages at 4 a.m. when there are only a handful of requests"), never the fix. At the practitioner/expert levels, requirements are stated as outcomes ("page payments when chat-gw is burning its error budget fast") and the agent has to find names, windows, for values, labels and routing policy in the repo's README.md, like a new hire would.
Workflows: slo-onboarding, alert-storm-cleanup, missed-page-postmortem (including total outages where series disappear, so only absent() can page), latency-slo (histogram SLIs), team-reorg-migration. Rules come as plain files or as Kubernetes PrometheusRule manifests.
Reward
Each requirement r gets a score q_r from its check families (mean(∅) = 1):
Alert (new or repair): q = E · dup · mean(Fire ∪ Timing) · mean(Silent) · (0.7 + 0.3·mean(Label ∪ Annotation))
(burn alerts: q = ½·q_integrated + ½·q_with_reference_SLI_records)
Recording: q = dup · mean(Value ∪ Cardinality)
Route: q = mean_positive Jaccard(expected, resolved receivers) · mean_negative [no forbidden receiver]
Inhibit: q = mean(Suppressed) · mean(NotSuppressed)
s_r = clip((q_r − q_r^pristine) / (1 − q_r^pristine), −0.25, 1) # pristine → 0, oracle → 1, regressions < 0
progress = max(0, (0.5 + 0.5·preservation) · Σ w_r s_r / Σ w_r) # weights: alert 3, repair 3, others 2
reward = 0.6 · [every s_r = 1 ∧ preservation = 1 ∧ syntax_ok ∧ ¬tamper] + 0.4 · progressWhy products: an always-firing rule fails every silent check, and a renamed or never-firing rule fails every fire check, so both score 0. A catch-all route passes positives and fails negatives. An inhibit-everything rule fails the not-suppressed cases. Partial work still counts: a wrong severity costs the label family only, a missing continue gives Jaccard ½, and a correct burn alert on top of broken SLI records keeps half its credit.
Reward spread (measured on the v0.1.0 tasks, before the 0.1.1 fairness fixes)
Adversaries (all 30 tasks, local gate): always-fire, rename, catch-all route, inhibit-all → targeted req_* ≤ 0.05. ALERTS injection, input-series shadowing, group interval changes → tamper, reward 0. Receiver nulling, over-broad inhibition, band thresholds, per-pod alerts, for-only fixes, deleting untouched rules → outcome 0 and reward ≤ 0.4.
Real-model results are in the two leaderboards above; a dense-vs-outcome GRPO training curve is not published yet.
How it was built: Skill2Env, but procedural
This mirrors NVIDIA's Skill2Env pipeline: skill/prometheus-alerting/SKILL.md → workflows/workflows.yaml (the "planner output", same keys) → sampled axes (archetype, verifier pattern, persona, tone, expertise, tier, rules format) → a deterministic, seeded creator instead of an LLM → acceptance gate (oracle = 1, nop = 0, partial band, 14 adversaries, a "missing for must fail Timing" mutant check, structural leak scan, determinism).
Checks are derived from the world spec through a reference evaluator (promsim.py: Prometheus rate extrapolation, left-open windows, histogram_quantile, staleness, the for state machine, in exact fractions), never from the oracle's rule text. The gate proves the oracle against real promtool, which is the structural equivalent of Skill2Env's freeze boundary.
Anti-hacking
Related work
- NVlabs/Skill2Env: the SkillHub
sre-engineerskill contains a literal 14.4x burn-rate rule, andslo-architectcovers burn-rate alerting. We downloaded ten SRE-adjacent Skill2Env tasks. The closest one (task_slo-architect_lgre556n) grades burn-rate policy by comparing YAML fields, and none run promtool or amtool. - camel-ai/seta-env task 1114: an Alertmanager routing and inhibition task. Its tests run amtool only for
check-configand check routes and inhibitions by reading the YAML. 3amBench's routing checks are end-to-end: the chain case routes the labels the agent's own alert carries. - Community rule sets: samber/awesome-prometheus-alerts, kubernetes-mixin; Google SRE Workbook ch. 5.
Limitations
- Traffic is synthetic (diurnal shape plus jitter), and services are fictional.
- Inhibition is graded by a strict simulator of Alertmanager's semantics, not by the Alertmanager process. Unsupported constructs fail closed.
- Seven alert templates host the defects, not a large upstream corpus.
- The routing chain case uses the agent's static labels, not labels observed in
ALERTS. - Frontier agents solved every run on the two easy tasks measured (af-001, af-013), so the easy tier likely saturates for them.
- Contamination: SkillHub skills (and any skill-augmented agent) already contain near-identical burn-rate recipes, and this repo ships solutions. Evaluate on a private-seed build, and report whether the agent had
SKILL.md. - English only.
- Expert tier (v0.2 candidates): hidden scenarios replay through promtool, whose samples sit exactly on the evaluation ticks. A target that leaves service discovery is marked stale on the next sample; a real Prometheus 3.5 server writes those stale markers about two scrape intervals after the last scrape, so an
absent()Down alert pages 1-2 minutes later in production than in the replay. The graders acceptabsent()Down alerts withforup to 10m, which in production page 11-13 minutes after the targets went away, against the 10 minutes the task README asks for. Window corners that depend on the scrape and evaluation phases are graded as promtool sees them:[90s]on a 1m scrape passes in the replay, but on a real server it is empty for about half of all phase pairs. - Expert tier: when a counter stops (consumer offsets, requests), a stated deadline is counted from the last sample that still increased, the earliest the stop can have happened. When a series goes away, it is counted from the minute its stale marker is written (see above for how that differs from a real server).
- Expert tier: a routing fix that sits below an unfixed catch-all route earns nothing on its own route checks, because the catch-all takes the alerts first. That is Alertmanager's first-match routing, kept on purpose: a run that misses the catch-all also loses the route tickets behind it.
- Expert tier:
clusteris a target label onuponly; request counters and histograms in the hidden data don't carry it. - Expert tier: every hidden scenario holds what the world's targets export at healthy levels: request counters for every service whose pods are up; the histograms, pool gauges and cAdvisor series the repo's rules read about those services; the ingress controller's request counters (
nginx_ingress_controller_requests) for every HTTP service; and the exporter series a ticket needs. Targets that go down or leave discovery get stale markers. The ingress counts the same requests and status codes as the app's own counters, and while a service has no target up it answers 502 (targets down) or 503 (none left) for the traffic clients keep sending. Not simulated: the ingress latency histogram and config-reload gauge, and fleet exporters (node, Postgres, Redis, Kafka, JVM) beyond those a ticket needs. The repo's fleet alerts and its ingress latency and reload alerts therefore never fire in the replay, and an alert that reads only those series cannot be graded on them. A scenario still starts without history: at its first evaluationrate()has no value yet, so a floor rule that reads a missing counter as zero traffic sees zero there, and itsforis what keeps it quiet. - Expert tier: an untouched alert counts as intact when it behaves the same. Its text may change (formatting, operand order, a static label equal to what its expression already yields, an identical second copy) as long as its
for, thresholds, windows and functions stay and it fires with the same labels at every minute of every hidden scenario. An alert that fires in none of them can only be kept as written. - Expert tier: an alert needs the labels the task's README names, with their values; it may carry others, which are judged by where the alert is delivered, not by their presence.
Repository layout
tasks/ (Harbor tasks: this repository holds the 30 core tasks; the Hugging Face dataset adds the 45 expert afx-* tasks that registry.json also lists, so run those with --repo https://huggingface.co/datasets/openenvforge/3ambench@v0.2.0 or build them with alertforge expert) · registry.json (3ambench, 3ambench-mini, 3ambench-expert, 3ambench-expert-lb, all 0.2.0) · manifest.jsonl (one row per core task) · partial/, null/ (reference policies) · skill/, workflows/ · generator/generate.py + src/alertforge/ (the generator and grader source; src/alertforge/expert/ builds the expert tier) · openenv/alertforge_env/ (OpenEnv server) · space/ (static replay viewer) · scripts/ (export_runs.py and leaderboard.py produce the replay data and the leaderboard tables, from a GitHub clone; the dataset does not ship space/) · tests/ (pytest).
License and citation
Apache-2.0; see NOTICE.md for attribution of adapted community rules (CC BY 4.0 / Apache-2.0).
@misc{3ambench2026,
title = {3amBench: a behaviorally graded, dense-reward RL environment for Prometheus alerting as code},
year = {2026},
note = {Harbor dataset and OpenEnv environment; generator package alertforge v0.2.0}
}