Team Ai
Apppublic

jester1177/cloudnative-devops-debug-env

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
README.md410 linesDownload Raw Back to root
1---2title: Cloud-Native DevOps Debug Environment3emoji: πŸ”§4colorFrom: blue5colorTo: green6sdk: docker7app_port: 78608pinned: false9---10 11# Cloud-Native DevOps Debug Environment12 13An OpenEnv-compatible environment where AI agents learn to debug broken GitHub Actions workflows, Dockerfiles, and Kubernetes manifests. Built for the OpenEnv Hackathon by Scaler School of Technology (partners: Meta, HuggingFace, PyTorch).14 15## Why Cloud-Native Debugging?16 17Every developer who ships code hits deployment pipeline failures. A misconfigured Dockerfile, a broken GitHub Actions workflow, a missing secret, a Kubernetes selector mismatch β€” these are the bugs that waste hours of developer time every week. They're hard to debug because:18 19- Error messages are cryptic ("unable to prepare context: unable to evaluate symlinks")20- The feedback loop is slow (push, wait for CI, read logs, fix, repeat)21- Multiple config files interact in non-obvious ways (Dockerfile + workflow + secrets + K8s manifests)22- Kubernetes errors require cross-resource reasoning (Deployment labels must match Service selectors)23 24This environment teaches AI agents to do what senior DevOps engineers do: read the error, trace it to the root cause across multiple files, and fix it.25 26---27 28## How It Works: The Complete Flow29 30```31β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”32β”‚  1. RESET                                                     β”‚33β”‚     Agent receives:                                           β”‚34β”‚     - Broken config files (Dockerfile / workflow / K8s YAML)  β”‚35β”‚     - Error message from the failed build/deploy              β”‚36β”‚     - Available secrets list                                  β”‚37β”‚     - Number of issues to find                                β”‚38β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€39β”‚  2. OBSERVE β†’ THINK β†’ ACT  (repeat up to 10 steps)           β”‚40β”‚     Agent reads the error, analyzes the files, then:          β”‚41β”‚     - edit_file: replace broken content with fixed content    β”‚42β”‚     - replace_line: fix a specific line number                β”‚43β”‚     - add_line / add_block: insert missing content            β”‚44β”‚     - delete_line / delete_block: remove bad content          β”‚45β”‚     - request_hint: get a clue (-4% score penalty)            β”‚46β”‚     - submit: "I'm done fixing"                               β”‚47β”‚                                                               β”‚48β”‚     After each action, agent gets:                            β”‚49β”‚     - Updated file contents                                   β”‚50β”‚     - Reward signal (+0.3 per fix, -0.02 for failed edits)   β”‚51β”‚     - How many issues are now fixed                           β”‚52β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€53β”‚  3. GRADE                                                     β”‚54β”‚     Deterministic scoring based on:                           β”‚55β”‚     - What fraction of issues were fixed                      β”‚56β”‚     - Whether ALL issues were fixed (bonus)                   β”‚57β”‚     - How many steps it took (efficiency)                     β”‚58β”‚     - How many hints were used (penalty)                      β”‚59β”‚     Score range: (0, 1) exclusive β€” never exactly 0 or 1     β”‚60β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜61```62 63---64 65## The 10 Tasks (50 Scenarios)66 67Evaluation runs **all 50 scenarios deterministically** across all 10 tasks for reproducible scoring.68 69### Task 1: Dockerfile Syntax Errors β€” Easy70 71Simple typos and instruction errors that break `docker build`.72 73| # | Scenario | What's Broken | Real-World Context |74|---|----------|---------------|-------------------|75| 1 | `typo_filename` | `COPY requirments.txt .` β€” misspelled filename | Most common Docker build error on Stack Overflow |76| 2 | `invalid_base_image` | `FROM python:3.9-slimm` β€” extra 'm' in tag | Happens when copy-pasting image tags |77| 3 | `invalid_run_syntax` | `RUN pip install ... \n && python setup.py` β€” broken line continuation | Formatting multi-line RUN commands is tricky |78| 4 | `copy_missing_source` | `COPY dist/` but build output is in `build/` | Source directory doesn't exist in build context |79| 5 | `missing_from_instruction` | No `FROM` instruction at all | Dockerfile must start with FROM |80 81### Task 2: Dockerfile Runtime Errors β€” Medium82 83The Dockerfile builds successfully, but the container crashes at runtime.84 85| # | Scenario | What's Broken | Real-World Context |86|---|----------|---------------|-------------------|87| 1 | `missing_workdir` | No WORKDIR β€” files scatter to `/` | Container runs but `npm start` can't find `package.json` |88| 2 | `cmd_entrypoint_conflict` | Both ENTRYPOINT and CMD defined as full commands | Process starts incorrectly |89| 3 | `entrypoint_not_executable` | Shell script lacks execute permission | `chmod +x` missing β€” "permission denied" |90| 4 | `missing_required_env` | App needs `DATABASE_URL` but it's not set | Container crashes: "DATABASE_URL is not defined" |91| 5 | `non_root_privileged_port` | Non-root user tries to bind port 80 | Security best practice conflicts with port < 1024 |92 93### Task 3: Workflow Syntax & Structure β€” Easy94 95GitHub Actions YAML has structural problems that GitHub rejects before any job runs.96 97| # | Scenario | What's Broken | Real-World Context |98|---|----------|---------------|-------------------|99| 1 | `checkout_after_build` | `docker build` before `actions/checkout` | No source code β€” "Dockerfile not found" |100| 2 | `missing_runs_on` | Job has no `runs-on` field | Every job needs a runner |101| 3 | `invalid_trigger_syntax` | `branches: main` instead of `branches: [main]` | Must be a YAML list |102| 4 | `missing_step_uses_or_run` | Step has a name but no `uses:` or `run:` | Invalid step |103| 5 | `missing_on_trigger` | No `on:` block at all | Workflow never triggers |104 105### Task 4: Workflow Secrets & Permissions β€” Medium106 107Secrets exist but aren't wired correctly to the workflow steps.108 109| # | Scenario | What's Broken | Real-World Context |110|---|----------|---------------|-------------------|111| 1 | `missing_env_secrets` | `$DOCKER_PASSWORD` without `env:` mapping | Secrets must be passed via `env:` block |112| 2 | `wrong_secret_syntax` | `${ secrets.TOKEN }` instead of `${{ secrets.TOKEN }}` | Single vs double braces |113| 3 | `missing_token_permissions` | Pushing to GHCR without `permissions: packages: write` | GITHUB_TOKEN is read-only by default |114| 4 | `secret_not_in_env` | `$SLACK_WEBHOOK_URL` not in `env:` | Very common mistake |115| 5 | `ghcr_wrong_credentials` | Using `DOCKER_PASSWORD` for GHCR login | GHCR uses `GITHUB_TOKEN` |116 117### Task 5: CI + Docker Integration β€” Medium118 119The workflow AND the Dockerfile interact. Fixing one file alone isn't enough.120 121| # | Scenario | What's Broken | Real-World Context |122|---|----------|---------------|-------------------|123| 1 | `missing_buildx_for_platforms` | Multi-platform build without `setup-buildx-action` | Need BuildKit for cross-compile |124| 2 | `missing_load_true` | `build-push-action` without `load: true` β€” next step can't find image | Buildx doesn't load into local daemon by default |125| 3 | `wrong_build_context` | Context is `./backend` but Dockerfile path is `./Dockerfile` | Path mismatch |126| 4 | `cache_without_mode_max` | GHA cache export missing `mode=max` | Cache doesn't persist |127| 5 | `push_without_login` | `docker push` without `docker login` first | "denied: requested access" |128 129### Task 6: Multi-Stage Pipeline & Matrix β€” Hard130 131Complex pipelines with multiple interacting bugs. Agent must find 2-3 issues across files.132 133| # | Scenario | What's Broken | Real-World Context |134|---|----------|---------------|-------------------|135| 1 | `artifact_path_mismatch` | `COPY --from=builder /app/dist` but React outputs to `/app/build` | CRA uses `build/`, Vite uses `dist/` |136| 2 | `matrix_platform_arg` | `$BUILDPLATFORM` without `ARG BUILDPLATFORM` | Multi-arch needs platform ARGs |137| 3 | `cross_job_artifact` | Test job downloads artifact but missing `needs: build` | Jobs run in parallel by default |138| 4 | `multiple_issues` | Dockerfile typo + workflow secrets not wired (2 bugs) | Problems compound across files |139| 5 | `matrix_version_failure` | Matrix includes Node 14 but code needs >= 16 + missing `needs:` | 2 bugs to find |140 141### Task 7: Kubernetes Pod Failures β€” Medium142 143Pod crashes and scheduling failures in Kubernetes deployments.144 145| # | Scenario | What's Broken | Real-World Context |146|---|----------|---------------|-------------------|147| 1 | `oom_killed` | Memory limit 64Mi too low β€” CrashLoopBackOff/OOMKilled | Most common K8s production issue |148| 2 | `image_pull_backoff` | Image tag typo `nginx:latset` β†’ ImagePullBackOff | Copy-paste tag errors |149| 3 | `wrong_command` | `command: ["python", "workers.py"]` but file is `worker.py` | File name mismatch |150| 4 | `missing_configmap` | `envFrom: configMapRef: app-config` but ConfigMap doesn't exist | CreateContainerConfigError |151| 5 | `liveness_probe_failing` | Liveness probe port 3000 but app listens on 8080 | Probe misconfiguration causes restarts |152 153### Task 8: Kubernetes Service & Ingress Issues β€” Hard154 155Networking issues where pods run fine but traffic doesn't reach them. Error messages are intentionally vague β€” the agent must diagnose from kubectl output.156 157| # | Scenario | What's Broken | Real-World Context |158|---|----------|---------------|-------------------|159| 1 | `selector_mismatch` | Service selector `app: api` but pod label is `app: api-server` | No endpoints β€” most common K8s networking bug |160| 2 | `port_mismatch` | Service targetPort 8080 but container listens on 3000 | Connection refused |161| 3 | `ingress_wrong_service` | Ingress references `api-svc` but service name is `api-service` | Ingress 404 |162| 4 | `network_policy_blocking` | NetworkPolicy with empty ingress rules blocks all traffic | Database unreachable |163| 5 | `missing_ingress_class` | No `ingressClassName: nginx` specified | Ingress controller doesn't pick it up |164 165### Task 9: CI/CD Build & Push Pipeline β€” Hard166 167GHA-to-Docker-to-Registry pipeline failures spanning multiple files.168 169| # | Scenario | What's Broken | Real-World Context |170|---|----------|---------------|-------------------|171| 1 | `registry_mismatch` | Build tags `ghcr.io/...` but push targets `docker.io/...` | Registry URL mismatch between steps |172| 2 | `image_tag_mismatch` | Build uses `github.ref_name` but push uses `github.sha` | "image not found locally" |173| 3 | `inconsistent_tagging` | `docker tag myuser/api:latest` but image was built as `myuser/api:${{ github.sha }}` | Tag source doesn't exist |174| 4 | `build_arg_not_passed` | Dockerfile `ARG APP_VERSION` but no `--build-arg` in workflow | Version file is empty |175| 5 | `dockerfile_path_in_subdirectory` | Workflow points to `./Dockerfile` but it's at `./services/api/Dockerfile` | Monorepo path mismatch |176 177### Task 10: Full Stack Deployment Pipeline β€” Expert178 179Multi-error scenarios spanning the entire stack: GHA + Dockerfile + K8s manifests. 2-4 bugs per scenario requiring cross-file reasoning. Error messages are intentionally vague β€” the agent must trace root causes from symptoms.180 181| # | Scenario | What's Broken | Real-World Context |182|---|----------|---------------|-------------------|183| 1 | `full_pipeline_ghcr_and_selector` | GHCR token not mapped + K8s Service selector mismatch | 2 bugs across workflow + K8s |184| 2 | `full_pipeline_three_bugs` | Missing checkout + no WORKDIR + wrong container/service port | 4 bugs across 4 files |185| 3 | `full_pipeline_ghcr_dockerfile_k8s` | Wrong GHCR secret + base image typo + OOM memory limit | 3 bugs across all layers |186| 4 | `full_pipeline_permissions_image_ingress` | Missing packages:write + hardcoded image placeholder + no ingressClassName | 3 bugs |187| 5 | `full_pipeline_secrets_build_probe` | Docker secrets not wired + wrong build output dir + probe port mismatch | 4 bugs across all layers |188 189---190 191## Fix Validation: Simulator-Based192 193Fixes are validated using **structural simulators**, not string matching. This means:194 195- **Alternative valid fixes are accepted.** Setting memory to `512Mi` instead of `256Mi` both resolve the OOM β€” the simulator accepts either.196- **Three independent simulators** run after every edit:197  - **DockerSimulator**: validates Dockerfile syntax (FROM, COPY, EXPOSE, RUN) and runtime behavior (WORKDIR, CMD/ENTRYPOINT, permissions, ENV)198  - **WorkflowSimulator**: parses YAML, checks triggers, runs-on, step ordering, secrets wiring, permissions, buildx requirements, registry consistency199  - **KubernetesSimulator**: validates manifests, cross-resource dependencies (Service selector ↔ Deployment labels), pod status simulation (OOM, ImagePullBackOff), service endpoint reachability200- **7 granular checks** are tracked: `docker_build`, `docker_run`, `workflow_parse`, `workflow_exec`, `k8s_valid`, `k8s_pod_running`, `k8s_service_active`201- Progress = how many checks flip from fail β†’ pass compared to the initial broken state202 203---204 205## Available Actions206 207Each step, the agent chooses exactly one action:208 209| Action | What It Does | When to Use |210|--------|-------------|-------------|211| `edit_file` | Replace `old_content` with `new_content` in a file | Most common β€” fix a broken line or block |212| `replace_line` | Replace content at a specific line number | When you know exactly which line is wrong |213| `add_line` | Insert a new line into a file | Adding missing instructions (e.g., missing `WORKDIR`) |214| `delete_line` | Remove a specific line | Removing a bad instruction |215| `add_block` | Insert a multi-line block | Adding entire sections (e.g., `env:` block with secrets) |216| `delete_block` | Remove a multi-line block | Removing incorrect sections |217| `request_hint` | Get a clue about what's wrong | Costs -4% on final score β€” use sparingly |218| `submit` | Declare "I'm done" β€” triggers final evaluation | When all fixes are applied |219 220**Important:** `edit_file` requires `old_content` to match **exactly** (including whitespace). If it doesn't match, the edit fails and the agent gets a -0.02 reward penalty.221 222---223 224## Grading System225 226Scoring is **deterministic** (same actions always produce the same score), **difficulty-aware** (harder tasks are graded more generously), and scores are strictly in **(0, 1) exclusive** β€” never exactly 0 or 1.227 228### The Formula229 230```231FINAL SCORE = Base + Partial Fixes + Complete Bonus + Difficulty Bonus + Efficiency - Hint Penalty - Failed Edit Penalty232```233 234Clamped to `(0.01, 0.99)`.235 236### Component Breakdown237 238| Component | Weight | Description |239|-----------|--------|-------------|240| Base score | 5% | Participation credit (guarantees score > 0) |241| Partial fixes | 35% | Proportional to `issues_fixed / issues_total` |242| Complete bonus | 25% | All issues fixed |243| Difficulty bonus | 0-3% | Extra reward for fully solving hard/expert tasks |244| Efficiency | 25% | Decays with extra steps β€” slower decay for harder tasks |245| Hint penalty | -3% to -4% each | Per `request_hint` action (cheaper for hard/expert) |246| Failed edit penalty | -2% each | Per edit with no valid file path |247 248### Difficulty Modifiers249 250| Difficulty | Max Score | Efficiency Decay | Hint Cost |251|------------|-----------|------------------|-----------|252| Easy | 0.90 | 0.03/step (strict) | 4% each |253| Medium | 0.90 | 0.027/step | 4% each |254| Hard/Expert | 0.93 | 0.021/step (forgiving) | 3% each |255 256---257 258## Evaluation259 260The evaluation pipeline runs **all 50 scenarios across all 10 tasks** deterministically:261 262```python263# Runs all 10 tasks Γ— 5 scenarios = 50 episodes264results = run_baseline_episodes()  # num_episodes=None runs all265 266# Per-episode scores in (0, 1)267# Aggregate = mean of all 50 scores268aggregate = sum(r.score for r in results) / len(results)269```270 271This ensures:272- **Reproducibility**: same agent produces same score every time273- **Complete coverage**: every error pattern is tested274- **Fair comparison**: all agents face the same 50 scenarios275 276---277 278## API Endpoints279 280| Endpoint | Method | Description |281|----------|--------|-------------|282| `/` | GET | Root page |283| `/health` | GET | Health check β€” returns `{"status": "healthy"}` |284| `/metadata` | GET | Environment name, description, version, tags |285| `/schema` | GET | Action, observation, and state JSON schemas |286| `/reset` | POST | Start a new episode (optional: `task_id`, `scenario_id`, `seed`) |287| `/step` | POST | Take an action and receive observation + reward |288| `/state` | GET | Get current observation without taking an action |289| `/info` | GET | Task list with metadata |290| `/tasks` | GET | List all tasks with difficulty levels |291| `/grader` | POST | Grade a trajectory (list of step dicts) |292| `/baseline` | POST | Run baseline across all scenarios (optional: `task_id`, `num_episodes`) |293| `/mcp` | POST | JSON-RPC 2.0 MCP endpoint (initialize, tools/list) |294 295### Example: Full Episode via API296 297```bash298# 1. Start an episode299curl -X POST http://localhost:7860/reset \300  -H "Content-Type: application/json" \301  -d '{"task_id": "k8s_pod_failures", "scenario_id": "oom_killed"}'302 303# 2. Fix the memory limit (any reasonable value works β€” simulator validates structurally)304curl -X POST http://localhost:7860/step \305  -H "Content-Type: application/json" \306  -d '{307    "action": {308      "action_type": "edit_file",309      "edits": [{310        "file_path": "k8s/deployment.yaml",311        "old_content": "memory: \"64Mi\"",312        "new_content": "memory: \"512Mi\""313      }]314    }315  }'316 317# Response: reward=0.3, issues_fixed=1/1, done=true318```319 320---321 322## Quick Start323 324### Local Development325 326```bash327pip install -r requirements.txt328python -m uvicorn server.app:app --host 0.0.0.0 --port 7860329```330 331### Run Tests332 333```bash334pytest tests/ -v335```336 337### Docker338 339```bash340docker build -t cloud-native-devops-env .341docker run -p 7860:7860 cloud-native-devops-env342```343 344### Baseline Inference (with LLM)345 346```bash347export API_BASE_URL=https://router.huggingface.co/v1348export MODEL_NAME=meta-llama/Llama-3.1-70B-Instruct349export HF_TOKEN=your_token_here350python inference.py351```352 353---354 355## Project Structure356 357```358cloud-native-devops-env/359β”œβ”€β”€ openenv.yaml              # OpenEnv environment specification360β”œβ”€β”€ inference.py              # LLM baseline (OpenAI client + HF router)361β”œβ”€β”€ baseline_runner.py        # Heuristic baseline β€” runs all 50 scenarios362β”œβ”€β”€ Dockerfile                # Production container363β”œβ”€β”€ requirements.txt          # Python dependencies364β”‚365β”œβ”€β”€ server/366β”‚   β”œβ”€β”€ app.py                # FastAPI with 12 endpoints367β”‚   β”œβ”€β”€ models.py             # Pydantic models (type-safe API)368β”‚   β”œβ”€β”€ environment.py        # Core environment loop (reset/step/state)369β”‚   β”œβ”€β”€ tasks/370β”‚   β”‚   β”œβ”€β”€ base.py           # BaseTask with scenario loading371β”‚   β”‚   β”œβ”€β”€ task_registry.py  # Maps task_id β†’ task class (10 tasks)372β”‚   β”‚   β”œβ”€β”€ task_1_build_errors.py        # 5 Dockerfile syntax scenarios373β”‚   β”‚   β”œβ”€β”€ task_2_docker_runtime.py      # 5 Dockerfile runtime scenarios374β”‚   β”‚   β”œβ”€β”€ task_3_workflow_syntax.py     # 5 workflow structure scenarios375β”‚   β”‚   β”œβ”€β”€ task_4_workflow_secrets_permissions.py  # 5 secrets scenarios376β”‚   β”‚   β”œβ”€β”€ task_5_ci_docker_integration.py        # 5 integration scenarios377β”‚   β”‚   β”œβ”€β”€ task_6_multi_stage_matrix.py           # 5 multi-issue scenarios378β”‚   β”‚   β”œβ”€β”€ k8s_pod.py                   # 5 Kubernetes pod failure scenarios379β”‚   β”‚   β”œβ”€β”€ k8s_networking.py            # 5 K8s networking scenarios380β”‚   β”‚   β”œβ”€β”€ pipeline_build_deploy.py     # 5 GHAβ†’Dockerβ†’Registry scenarios381β”‚   β”‚   └── pipeline_full.py             # 5 full-stack multi-error scenarios382β”‚   β”œβ”€β”€ graders/383β”‚   β”‚   └── __init__.py       # Deterministic trajectory grader384β”‚   └── simulators/385β”‚       β”œβ”€β”€ docker_simulator.py   # Dockerfile build + runtime validation386β”‚       β”œβ”€β”€ workflow_simulator.py # GHA workflow parse + execution validation387β”‚       └── k8s_simulator.py     # K8s manifest + cross-resource validation388β”‚389└── tests/390    β”œβ”€β”€ test_endpoints.py     # API endpoint tests391    β”œβ”€β”€ test_determinism.py   # Grader determinism + score range tests392    β”œβ”€β”€ test_baseline.py      # Heuristic baseline tests393    β”œβ”€β”€ test_environment_flow.py  # Episode flow tests394    └── test_simulators.py    # Simulator unit tests395```396 397## Design Decisions398 3991. **Full cloud-native stack**: Docker + GitHub Actions + Kubernetes β€” the three pillars of modern deployment pipelines.4002. **Simulator-based validation**: Structural rule-based simulators validate fixes instead of string matching. Alternative valid fixes are accepted (e.g., `512Mi` and `256Mi` both fix an OOM). Deterministic, fast, no security concerns.4013. **Dense rewards**: Partial credit at every step (+0.3 per fix, -0.02 per failed edit) rather than sparse pass/fail.4024. **Difficulty progression**: Easy tasks are single-file, single-issue. Expert tasks are multi-file, multi-issue with interacting bugs across all three layers.4035. **Vague error messages in harder tasks**: Easy tasks have explicit error messages. Hard/Expert tasks have realistic, vague messages that require the agent to actually diagnose the issue from context.4046. **Deterministic evaluation**: All 50 scenarios run every time for reproducible, comparable scores in (0, 1) exclusive.4057. **50 scenarios from real bugs**: Every scenario is based on actual developer mistakes documented on Stack Overflow, GitHub Issues, and official documentation.406 407## License408 409MIT410