jester1177/cloudnative-devops-debug-env
Cloud-Native DevOps Debug Environment
An OpenEnv-compatible environment where AI agents learn to debug broken GitHub Actions workflows, Dockerfiles, and Kubernetes manifests. Built for the OpenEnv Hackathon by Scaler School of Technology (partners: Meta, HuggingFace, PyTorch).
Why Cloud-Native Debugging?
Every developer who ships code hits deployment pipeline failures. A misconfigured Dockerfile, a broken GitHub Actions workflow, a missing secret, a Kubernetes selector mismatch — these are the bugs that waste hours of developer time every week. They're hard to debug because:
- Error messages are cryptic ("unable to prepare context: unable to evaluate symlinks")
- The feedback loop is slow (push, wait for CI, read logs, fix, repeat)
- Multiple config files interact in non-obvious ways (Dockerfile + workflow + secrets + K8s manifests)
- Kubernetes errors require cross-resource reasoning (Deployment labels must match Service selectors)
This environment teaches AI agents to do what senior DevOps engineers do: read the error, trace it to the root cause across multiple files, and fix it.
How It Works: The Complete Flow
┌──────────────────────────────────────────────────────────────┐
│ 1. RESET │
│ Agent receives: │
│ - Broken config files (Dockerfile / workflow / K8s YAML) │
│ - Error message from the failed build/deploy │
│ - Available secrets list │
│ - Number of issues to find │
├──────────────────────────────────────────────────────────────┤
│ 2. OBSERVE → THINK → ACT (repeat up to 10 steps) │
│ Agent reads the error, analyzes the files, then: │
│ - edit_file: replace broken content with fixed content │
│ - replace_line: fix a specific line number │
│ - add_line / add_block: insert missing content │
│ - delete_line / delete_block: remove bad content │
│ - request_hint: get a clue (-4% score penalty) │
│ - submit: "I'm done fixing" │
│ │
│ After each action, agent gets: │
│ - Updated file contents │
│ - Reward signal (+0.3 per fix, -0.02 for failed edits) │
│ - How many issues are now fixed │
├──────────────────────────────────────────────────────────────┤
│ 3. GRADE │
│ Deterministic scoring based on: │
│ - What fraction of issues were fixed │
│ - Whether ALL issues were fixed (bonus) │
│ - How many steps it took (efficiency) │
│ - How many hints were used (penalty) │
│ Score range: (0, 1) exclusive — never exactly 0 or 1 │
└──────────────────────────────────────────────────────────────┘The 10 Tasks (50 Scenarios)
Evaluation runs all 50 scenarios deterministically across all 10 tasks for reproducible scoring.
Task 1: Dockerfile Syntax Errors — Easy
Simple typos and instruction errors that break docker build.
Task 2: Dockerfile Runtime Errors — Medium
The Dockerfile builds successfully, but the container crashes at runtime.
Task 3: Workflow Syntax & Structure — Easy
GitHub Actions YAML has structural problems that GitHub rejects before any job runs.
Task 4: Workflow Secrets & Permissions — Medium
Secrets exist but aren't wired correctly to the workflow steps.
Task 5: CI + Docker Integration — Medium
The workflow AND the Dockerfile interact. Fixing one file alone isn't enough.
Task 6: Multi-Stage Pipeline & Matrix — Hard
Complex pipelines with multiple interacting bugs. Agent must find 2-3 issues across files.
Task 7: Kubernetes Pod Failures — Medium
Pod crashes and scheduling failures in Kubernetes deployments.
Task 8: Kubernetes Service & Ingress Issues — Hard
Networking issues where pods run fine but traffic doesn't reach them. Error messages are intentionally vague — the agent must diagnose from kubectl output.
Task 9: CI/CD Build & Push Pipeline — Hard
GHA-to-Docker-to-Registry pipeline failures spanning multiple files.
Task 10: Full Stack Deployment Pipeline — Expert
Multi-error scenarios spanning the entire stack: GHA + Dockerfile + K8s manifests. 2-4 bugs per scenario requiring cross-file reasoning. Error messages are intentionally vague — the agent must trace root causes from symptoms.
Fix Validation: Simulator-Based
Fixes are validated using structural simulators, not string matching. This means:
- Alternative valid fixes are accepted. Setting memory to
512Miinstead of256Miboth resolve the OOM — the simulator accepts either. - Three independent simulators run after every edit:
- DockerSimulator: validates Dockerfile syntax (FROM, COPY, EXPOSE, RUN) and runtime behavior (WORKDIR, CMD/ENTRYPOINT, permissions, ENV)
- WorkflowSimulator: parses YAML, checks triggers, runs-on, step ordering, secrets wiring, permissions, buildx requirements, registry consistency
- KubernetesSimulator: validates manifests, cross-resource dependencies (Service selector ↔ Deployment labels), pod status simulation (OOM, ImagePullBackOff), service endpoint reachability
- 7 granular checks are tracked:
docker_build,docker_run,workflow_parse,workflow_exec,k8s_valid,k8s_pod_running,k8s_service_active - Progress = how many checks flip from fail → pass compared to the initial broken state
Available Actions
Each step, the agent chooses exactly one action:
Important: edit_file requires old_content to match exactly (including whitespace). If it doesn't match, the edit fails and the agent gets a -0.02 reward penalty.
Grading System
Scoring is deterministic (same actions always produce the same score), difficulty-aware (harder tasks are graded more generously), and scores are strictly in (0, 1) exclusive — never exactly 0 or 1.
The Formula
FINAL SCORE = Base + Partial Fixes + Complete Bonus + Difficulty Bonus + Efficiency - Hint Penalty - Failed Edit PenaltyClamped to (0.01, 0.99).
Component Breakdown
Difficulty Modifiers
Evaluation
The evaluation pipeline runs all 50 scenarios across all 10 tasks deterministically:
# Runs all 10 tasks × 5 scenarios = 50 episodes
results = run_baseline_episodes() # num_episodes=None runs all
# Per-episode scores in (0, 1)
# Aggregate = mean of all 50 scores
aggregate = sum(r.score for r in results) / len(results)This ensures:
- Reproducibility: same agent produces same score every time
- Complete coverage: every error pattern is tested
- Fair comparison: all agents face the same 50 scenarios
API Endpoints
Example: Full Episode via API
# 1. Start an episode
curl -X POST http://localhost:7860/reset \
-H "Content-Type: application/json" \
-d '{"task_id": "k8s_pod_failures", "scenario_id": "oom_killed"}'
# 2. Fix the memory limit (any reasonable value works — simulator validates structurally)
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"action_type": "edit_file",
"edits": [{
"file_path": "k8s/deployment.yaml",
"old_content": "memory: \"64Mi\"",
"new_content": "memory: \"512Mi\""
}]
}
}'
# Response: reward=0.3, issues_fixed=1/1, done=trueQuick Start
Local Development
pip install -r requirements.txt
python -m uvicorn server.app:app --host 0.0.0.0 --port 7860Run Tests
pytest tests/ -vDocker
docker build -t cloud-native-devops-env .
docker run -p 7860:7860 cloud-native-devops-envBaseline Inference (with LLM)
export API_BASE_URL=https://router.huggingface.co/v1
export MODEL_NAME=meta-llama/Llama-3.1-70B-Instruct
export HF_TOKEN=your_token_here
python inference.pyProject Structure
cloud-native-devops-env/
├── openenv.yaml # OpenEnv environment specification
├── inference.py # LLM baseline (OpenAI client + HF router)
├── baseline_runner.py # Heuristic baseline — runs all 50 scenarios
├── Dockerfile # Production container
├── requirements.txt # Python dependencies
│
├── server/
│ ├── app.py # FastAPI with 12 endpoints
│ ├── models.py # Pydantic models (type-safe API)
│ ├── environment.py # Core environment loop (reset/step/state)
│ ├── tasks/
│ │ ├── base.py # BaseTask with scenario loading
│ │ ├── task_registry.py # Maps task_id → task class (10 tasks)
│ │ ├── task_1_build_errors.py # 5 Dockerfile syntax scenarios
│ │ ├── task_2_docker_runtime.py # 5 Dockerfile runtime scenarios
│ │ ├── task_3_workflow_syntax.py # 5 workflow structure scenarios
│ │ ├── task_4_workflow_secrets_permissions.py # 5 secrets scenarios
│ │ ├── task_5_ci_docker_integration.py # 5 integration scenarios
│ │ ├── task_6_multi_stage_matrix.py # 5 multi-issue scenarios
│ │ ├── k8s_pod.py # 5 Kubernetes pod failure scenarios
│ │ ├── k8s_networking.py # 5 K8s networking scenarios
│ │ ├── pipeline_build_deploy.py # 5 GHA→Docker→Registry scenarios
│ │ └── pipeline_full.py # 5 full-stack multi-error scenarios
│ ├── graders/
│ │ └── __init__.py # Deterministic trajectory grader
│ └── simulators/
│ ├── docker_simulator.py # Dockerfile build + runtime validation
│ ├── workflow_simulator.py # GHA workflow parse + execution validation
│ └── k8s_simulator.py # K8s manifest + cross-resource validation
│
└── tests/
├── test_endpoints.py # API endpoint tests
├── test_determinism.py # Grader determinism + score range tests
├── test_baseline.py # Heuristic baseline tests
├── test_environment_flow.py # Episode flow tests
└── test_simulators.py # Simulator unit testsDesign Decisions
- Full cloud-native stack: Docker + GitHub Actions + Kubernetes — the three pillars of modern deployment pipelines.
- Simulator-based validation: Structural rule-based simulators validate fixes instead of string matching. Alternative valid fixes are accepted (e.g.,
512Miand256Miboth fix an OOM). Deterministic, fast, no security concerns. - Dense rewards: Partial credit at every step (+0.3 per fix, -0.02 per failed edit) rather than sparse pass/fail.
- Difficulty progression: Easy tasks are single-file, single-issue. Expert tasks are multi-file, multi-issue with interacting bugs across all three layers.
- Vague error messages in harder tasks: Easy tasks have explicit error messages. Hard/Expert tasks have realistic, vague messages that require the agent to actually diagnose the issue from context.
- Deterministic evaluation: All 50 scenarios run every time for reproducible, comparable scores in (0, 1) exclusive.
- 50 scenarios from real bugs: Every scenario is based on actual developer mistakes documented on Stack Overflow, GitHub Issues, and official documentation.
License
MIT
