shivakunv/devops_incident_sim
OpenEnv DevOps Incident Simulator
This project helps you test DevOps incident handling from real-looking logs. You can paste or upload logs (dmesg/syslog/core dump style), get a predicted issue, apply a mapped fix, and see a score. It now also includes a Phase 2 incident_command scenario for richer multi-step response, communication, and postmortem evaluation.
Deployment On Hugging Face
This app is intended to run as a Hugging Face Docker Space.
- Deployment is triggered manually from the GitHub Actions workflow file [
.github/workflows/deploy-openenv.yml]. - The Space should use port
7860. - The app UI is available at
/,/web, and/log-ui. - If Hugging Face probes
/web, it should still load the same log UI.
Submission Runtime
- The root inference entrypoint is
inference.py. - LLM configuration is read from
API_BASE_URL,MODEL_NAME, andHF_TOKEN. - LLM calls use the OpenAI Python client and fall back to a deterministic heuristic if the remote endpoint is unavailable.
What The App Does
Core behavior:
- Ingests logs (raw text or uploaded log file).
- Predicts incident label (
disk_full,memory_leak,crash,network_issue). - Maps predicted label to an internal remediation action.
- Runs environment episode (
analyze_logs -> take_action). - Returns
reward,done, and graderscore.
Goal: make log triage easy to test and demo.
Phase 2 behavior:
- Runs a richer
incident_commandenvironment task. - Supports delegation, stakeholder communication, and postmortem actions.
- Grades episodes with a configurable reward engine and score breakdown.
Main Capabilities
- Raw log ingestion via JSON (
/ingest_log) - File upload ingestion via multipart (
/ingest_log_file) - Browser UI for paste/upload (
/log-ui) - Configurable reward scoring with component breakdown (
/grader) - Classic OpenEnv endpoints (
/reset,/step,/state,/tasks,/baseline) - Phase 2 multi-step task:
incident_command - Docker-ready deployment
Hackathon Packaging
- Reward evidence artifact: `outputs/reward_evidence/README.md`
- Mini blog and demo outline: `docs/demo/MINI_BLOG.md`
- Hugging Face blog post draft: `docs/demo/HF_BLOG_POST.md`
- Converted reference docs: `docs/README.md`
- Repro command:
python3 evaluation/generate_reward_evidence.py
Submission Links
- Hugging Face Space URL:
https://shivakunv-devops-incident-sim.hf.space/ - Training run notebook in repo: `notebooks/openenv_devops_training_run.ipynb`
- Public Colab link (stable):
https://colab.research.google.com/drive/16jxAyoxadyQglSzYBZe-Vo8Qe9P3oTcz?usp=sharing - Code repository link:
https://github.com/shivakunv/openenv-devops-simulator - Hugging Face blog post draft: `docs/demo/HF_BLOG_POST.md`
- Publish-ready blog editor entry:
https://huggingface.co/new-blog - Final published blog URL:
https://huggingface.co/spaces/shivakunv/devops_incident_sim
Training Evidence Plots
These plots must be committed as real image files and embedded inline for automated validation.
Generate real training metrics and plot images:
source ~/venv/bin/activate && python3 models/train.py --stage all --output-dir outputs/phase2_training && python3 scripts/export_training_metrics.py --output-dir outputs/phase2_training --out outputs/phase2_training/training_metrics.json && python3 scripts/render_training_curves.py --input outputs/phase2_training/training_metrics.json --output-dir docs/assetsExpected metrics JSON path:
outputs/phase2_training/training_metrics.json
Dir Highlights
api/server.py- API routes and ingestion logicenv/- simulator environment and reward mechanicstasks/- easy/medium/hard plusincident_commandgraders/- episode gradingmodels/- log classifierdata/live_logs/- sample realistic log filesscripts/smoke_test.sh- endpoint smoke testsevaluation/run_incident_command_eval.py- Phase 2 scenario evaluator
install Docker
sudo apt update
sudo apt install ca-certificates curl gnupg lsb-release -y
sudo mkdir -p /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
echo \
"deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] \
https://download.docker.com/linux/ubuntu $(lsb_release -cs) stable" | \
sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
sudo apt update
sudo apt install docker-ce docker-ce-cli containerd.io -yBuild and Run (Docker)
apt install python3.12-venv
python3 -m venv venv
source venv/bin/activate
docker build -t openenv-devops .
docker run -p 7860:7860 openenv-devopsBase URL: http://localhost:7860
Log Analysis Workflows
1) Upload log file
curl -s -X POST http://localhost:7860/ingest_log_file \
-F "file=@data/live_logs/dmesg.log" \
-F "source=curl_upload"2) Send pasted/raw log text
curl -s -X POST http://localhost:7860/ingest_log \
-H "Content-Type: application/json" \
-d '{"source":"manual_text","log":"Out of memory: Killed process 2145 (python3)"}'3) Use browser UI
Open: http://localhost:7860/log-ui
You can paste logs or upload .log/.txt and receive result JSON.
Result JSON Meaning
Example response:
{
"source": "curl_upload",
"predicted_label": "memory_leak",
"mapped_task": "medium",
"recommended_fix": "restart_service",
"done": true,
"last_reward": 1.05,
"score": 1.0,
"score_breakdown": {
"recovery": 1.0,
"root_cause": 1.0,
"efficiency": 1.0,
"safety": 1.0,
"coordination": 1.0,
"communication": 1.0,
"learning": 1.0
},
"reward_profile": "medium"
}Field meaning:
predicted_label: detected incident type from logmapped_task: internal scenario selected by labelrecommended_fix: remediation action selecteddone: episode completion statuslast_reward: reward from final action stepscore: weighted episode scorescore_breakdown: per-dimension component scores used by the reward enginereward_profile: scoring profile applied to the task
Label to Action Mapping
disk_full->clear_diskmemory_leak->restart_servicecrash->restart_servicenetwork_issue->scale_up
Other API Endpoints
POST /resetPOST /stepGET /stateGET /tasksPOST /graderGET /baseline
Available Tasks
easymediumhardincident_command
incident_command is the new Phase 2 scenario. It simulates a checkout outage with memory pressure, deployment context, stakeholder updates, and rewards for coordination work beyond simple remediation.
Round 2 Positioning
This repo now targets the April 2026 Round 2 OpenEnv hackathon themes through one unified environment design:
- multi-agent interactions
- long-horizon planning and instruction following
- world modeling for professional workflows
- self-improving agent systems
The incident_command task is the main Phase 2 scenario used to demonstrate those themes.
Training Pipeline
The repo now includes a full TRL-based training pipeline in models/train.py.
It supports:
- supervised fine-tuning on oracle action plans for all tasks
- GRPO fine-tuning using the local environment reward engine
- local evaluation that writes reward metrics to
outputs/phase2_training/eval_metrics.json
Install the optional training stack:
source ~/venv/bin/activate
pip install -r requirements-training.txtRun the full pipeline:
source ~/venv/bin/activate
python3 models/train.py \
--stage all \
--model-name-or-path Qwen/Qwen2.5-0.5B-Instruct \
--output-dir outputs/phase2_trainingRun only SFT:
python3 models/train.py --stage sft --output-dir outputs/phase2_trainingRun only GRPO starting from an existing checkpoint:
python3 models/train.py \
--stage grpo \
--model-name-or-path outputs/phase2_training/sft \
--output-dir outputs/phase2_trainingRun evaluation:
python3 models/train.py \
--stage eval \
--model-name-or-path outputs/phase2_training/grpo \
--output-dir outputs/phase2_trainingMain outputs:
outputs/phase2_training/sft/outputs/phase2_training/grpo/outputs/phase2_training/sft_dataset.jsonloutputs/phase2_training/eval_metrics.json
Phase 2 Actions
The classic tasks still use:
analyze_logstake_action
The new incident_command task also supports:
delegate_investigationcommunicate_statuswrite_postmortem
Example flow:
curl -s -X POST "http://localhost:7860/reset?task_name=incident_command"
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" \
-d '{"action_type":"delegate_investigation","payload":{"role":"sre_agent","objective":"Check memory pressure and recent deploy changes"}}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" \
-d '{"action_type":"communicate_status","payload":{"audience":"stakeholders","summary":"Investigating checkout degradation and elevated latency."}}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" \
-d '{"action_type":"analyze_logs"}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" \
-d '{"action_type":"take_action","payload":{"fix":"restart_service"}}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" \
-d '{"action_type":"write_postmortem","payload":{"summary":"checkout-api memory pressure after deployment caused failed checkouts.","action_items":["Add memory regression guard to rollout checks"]}}'
curl -s -X POST http://localhost:7860/graderQuick Validation
curl -s http://localhost:7860/tasks
./scripts/smoke_test.shHigh-Level Testing Strategy
If you want a clean test flow, use this order.
- Service Health Check
- Start container and check if API is reachable:
docker build -t openenv-devops .
docker run -p 7860:7860 openenv-devops
curl -s http://localhost:7860/tasks- Confirm task list has
easy,medium,hard, andincident_command.
- End-to-End API Regression
- Run the built-in smoke test:
./scripts/smoke_test.sh- This validates
reset -> step -> graderfor easy/medium/hard and baseline. - For the Phase 2 flow, also run:
python3 evaluation/run_incident_command_eval.py- Log Ingestion Path Testing
- JSON text ingestion:
curl -s -X POST http://localhost:7860/ingest_log \
-H "Content-Type: application/json" \
-d '{"source":"manual_text","log":"Out of memory: Killed process 2145 (python3)"}'- File upload ingestion:
curl -s -X POST http://localhost:7860/ingest_log_file \
-F "file=@data/live_logs/dmesg.log" \
-F "source=curl_upload"- Confirm output has:
predicted_label,recommended_fix,done,score.
- Sample Log Coverage
- Test with all sample files:
data/live_logs/dmesg.logdata/live_logs/syslog.logdata/live_logs/coredump.logdata/live_logs/live_logs_bundle.txt- Goal: every input path should work and return valid JSON.
- UI Workflow Validation
- Open
http://localhost:7860/log-ui. - Validate both:
- paste log text flow
- upload file flow
- Confirm returned JSON fields match API output.
- Result Interpretation Validation
- Functional success criteria:
- request succeeds (2xx)
doneistruescoreis present (1.0indicates successful mapped remediation)
Before running the evaluation commands below, start the server:
uv run --active server- Baseline Agent Re-run
python3 evaluation/run_agent_eval.py --agent baseline --runs 5- This re-runs baseline across
easy,medium,hardfor multiple rounds.
- Phase 2 Incident Command Run
python3 evaluation/run_incident_command_eval.py- Runs the new multi-step incident-response scenario and prints grading breakdown.
- Standard LLM Agent Run (all tasks)
export LLM_API_KEY=<your_api_key>
python3 evaluation/run_agent_eval.py \
--agent llm \
--runs 1 \
--model nvidia/llama-3.1-nemotron-70b-instruct- Runs LLM-driven fix selection for all tasks.
- Score Variance Check
python3 evaluation/variance_check.py --agent baseline --runs 10Hackathon Deliverables Checklist
- OpenEnv environment: implemented
- measurable reward model: implemented
- Phase 2 multi-agent scenario: implemented
- full HF TRL training pipeline: implemented in
models/train.py - reward-improvement evidence for demo: packaged in
outputs/reward_evidence/ - mini-blog or short video: packaged as a repo-ready draft in
docs/demo/MINI_BLOG.md
Optional variance check for LLM:
export LLM_API_KEY=<your_api_key>
python3 evaluation/variance_check.py \
--agent llm \
--runs 10 \
--model nvidia/llama-3.1-nemotron-70b-instructDeveloper Debugging Guide
- Confirm latest routes are loaded
curl -s http://localhost:7860/openapi.json | rg "ingest_log|ingest_log_file|log-ui"- Verify server is responding
curl -i http://localhost:7860/tasks
curl -i http://localhost:7860/state- Debug ingestion requests with verbose curl
curl -v -X POST http://localhost:7860/ingest_log \
-H "Content-Type: application/json" \
-d '{"source":"debug","log":"Connection timed out during TLS handshake"}'- Debug file upload path
curl -v -X POST http://localhost:7860/ingest_log_file \
-F "file=@data/live_logs/dmesg.log" \
-F "source=debug_upload"- Validate full environment flow manually
curl -s -X POST "http://localhost:7860/reset?task_name=easy"
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"analyze_logs"}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"take_action","payload":{"fix":"clear_disk"}}'
curl -s -X POST http://localhost:7860/graderSample Log Files
data/live_logs/coredump.logdata/live_logs/dmesg.logdata/live_logs/syslog.logdata/live_logs/live_logs_bundle.txt
Task-wise JSON Testing (easy / medium / hard)
Use this exact flow to validate core environment logic per task.
Easy
curl -s -X POST "http://localhost:7860/reset?task_name=easy"
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"analyze_logs"}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"take_action","payload":{"fix":"clear_disk"}}'
curl -s -X POST http://localhost:7860/graderMedium
curl -s -X POST "http://localhost:7860/reset?task_name=medium"
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"analyze_logs"}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"take_action","payload":{"fix":"restart_service"}}'
curl -s -X POST http://localhost:7860/graderHard
curl -s -X POST "http://localhost:7860/reset?task_name=hard"
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"analyze_logs"}'
curl -s -X POST http://localhost:7860/step -H "Content-Type: application/json" -d '{"action_type":"take_action","payload":{"fix":"scale_up"}}'
curl -s -X POST http://localhost:7860/graderExpected for all three:
- final
take_actionresponse contains"done":true POST /graderreturns"score":1.0
