Team Ai
Apppublic

jester1177/cloud-native-debug-env

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

Cloud-Native DevOps Debug Environment

An OpenEnv-compatible environment where AI agents learn to debug broken GitHub Actions workflows, Dockerfiles, and Kubernetes manifests. Built for the OpenEnv Hackathon by Scaler School of Technology (partners: Meta, HuggingFace, PyTorch).

Why Cloud-Native Debugging?

Every developer who ships code hits deployment pipeline failures. A misconfigured Dockerfile, a broken GitHub Actions workflow, a missing secret, a Kubernetes selector mismatch — these are the bugs that waste hours of developer time every week. They're hard to debug because:

  • —Error messages are cryptic ("unable to prepare context: unable to evaluate symlinks")
  • —The feedback loop is slow (push, wait for CI, read logs, fix, repeat)
  • —Multiple config files interact in non-obvious ways (Dockerfile + workflow + secrets + K8s manifests)
  • —Kubernetes errors require cross-resource reasoning (Deployment labels must match Service selectors)

This environment teaches AI agents to do what senior DevOps engineers do: read the error, trace it to the root cause across multiple files, and fix it.


How It Works: The Complete Flow

┌──────────────────────────────────────────────────────────────┐
│  1. RESET                                                     │
│     Agent receives:                                           │
│     - Broken config files (Dockerfile / workflow / K8s YAML)  │
│     - Error message from the failed build/deploy              │
│     - Available secrets list                                  │
│     - Number of issues to find                                │
├──────────────────────────────────────────────────────────────┤
│  2. OBSERVE → THINK → ACT  (repeat up to 10 steps)           │
│     Agent reads the error, analyzes the files, then:          │
│     - edit_file: replace broken content with fixed content    │
│     - replace_line: fix a specific line number                │
│     - add_line / add_block: insert missing content            │
│     - delete_line / delete_block: remove bad content          │
│     - request_hint: get a clue (-5% score penalty)            │
│     - submit: "I'm done fixing"                               │
│                                                               │
│     After each action, agent gets:                            │
│     - Updated file contents                                   │
│     - Reward signal (+0.3 per fix, -0.02 for failed edits)   │
│     - How many issues are now fixed                           │
├──────────────────────────────────────────────────────────────┤
│  3. GRADE                                                     │
│     Deterministic scoring based on:                           │
│     - What fraction of issues were fixed                      │
│     - Whether ALL issues were fixed (bonus)                   │
│     - How many steps it took (efficiency)                     │
│     - How many hints were used (penalty)                      │
└──────────────────────────────────────────────────────────────┘

The 10 Tasks (50 Scenarios)

Task 1: Dockerfile Syntax Errors — Easy

Simple typos and instruction errors that break docker build.

#ScenarioWhat's BrokenReal-World Context
1typo_filenameCOPY requirments.txt . — misspelled filenameMost common Docker build error on Stack Overflow
2invalid_base_imageFROM python:3.9-slimm — extra 'm' in tagHappens when copy-pasting image tags
3invalid_run_syntaxRUN pip install ... \n && python setup.py — broken line continuationFormatting multi-line RUN commands is tricky
4invalid_exposeEXPOSE "eighty" — string instead of port numberEXPOSE only accepts numeric ports
5missing_from_instructionNo FROM instruction at allDockerfile must start with FROM

Task 2: Dockerfile Runtime Errors — Medium

The Dockerfile builds successfully, but the container crashes at runtime.

#ScenarioWhat's BrokenReal-World Context
1missing_workdirNo WORKDIR — files scatter to /Container runs but npm start can't find package.json
2cmd_entrypoint_conflictBoth ENTRYPOINT and CMD defined as full commandsProcess starts incorrectly
3entrypoint_not_executableShell script lacks execute permissionchmod +x missing — "permission denied"
4missing_required_envApp needs DATABASE_URL but it's not setContainer crashes: "DATABASE_URL is not defined"
5non_root_privileged_portNon-root user tries to bind port 80Security best practice conflicts with port < 1024

Task 3: Workflow Syntax & Structure — Easy

GitHub Actions YAML has structural problems that GitHub rejects before any job runs.

#ScenarioWhat's BrokenReal-World Context
1checkout_after_builddocker build before actions/checkoutNo source code — "Dockerfile not found"
2missing_runs_onJob has no runs-on fieldEvery job needs a runner
3invalid_trigger_syntaxbranches: main instead of branches: [main]Must be a YAML list
4missing_step_uses_or_runStep has a name but no uses: or run:Invalid step
5missing_on_triggerNo on: block at allWorkflow never triggers

Task 4: Workflow Secrets & Permissions — Medium

Secrets exist but aren't wired correctly to the workflow steps.

#ScenarioWhat's BrokenReal-World Context
1missing_env_secrets$DOCKER_PASSWORD without env: mappingSecrets must be passed via env: block
2wrong_secret_syntax${ secrets.TOKEN } instead of ${{ secrets.TOKEN }}Single vs double braces
3missing_token_permissionsPushing to GHCR without permissions: packages: writeGITHUB_TOKEN is read-only by default
4secret_not_in_env$SLACK_WEBHOOK_URL not in env:Very common mistake
5ghcr_wrong_credentialsUsing DOCKER_PASSWORD for GHCR loginGHCR uses GITHUB_TOKEN

Task 5: CI + Docker Integration — Medium-Hard

The workflow AND the Dockerfile interact. Fixing one file alone isn't enough.

#ScenarioWhat's BrokenReal-World Context
1missing_buildx_for_platformsMulti-platform build without setup-buildx-actionNeed BuildKit for cross-compile
2login_secrets_not_wireddocker login missing env: for secrets"unauthorized: authentication required"
3wrong_build_contextContext is ./backend but Dockerfile path is ./DockerfilePath mismatch
4cache_without_mode_maxGHA cache export missing mode=maxCache doesn't persist
5push_without_logindocker push without docker login first"denied: requested access"

Task 6: Multi-Stage Pipeline & Matrix — Hard

Complex pipelines with multiple interacting bugs. Agent must find 2-3 issues across files.

#ScenarioWhat's BrokenReal-World Context
1artifact_path_mismatchCOPY --from=builder /app/dist but React outputs to /app/buildCRA uses build/, Vite uses dist/
2matrix_platform_arg$BUILDPLATFORM without ARG BUILDPLATFORMMulti-arch needs platform ARGs
3cross_job_artifactTest job downloads artifact but missing needs: buildJobs run in parallel by default
4multiple_issuesDockerfile typo + workflow secrets not wired (2 bugs)Problems compound across files
5matrix_version_failureMatrix includes Node 14 but code needs >= 16 + missing needs:2 bugs to find

Task 7: Kubernetes Pod Failures — Medium

Pod crashes and scheduling failures in Kubernetes deployments.

#ScenarioWhat's BrokenReal-World Context
1oom_killedMemory limit 64Mi too low — CrashLoopBackOff/OOMKilledMost common K8s production issue
2image_pull_backoffImage tag typo nginx:latset → ImagePullBackOffCopy-paste tag errors
3wrong_commandcommand: ["python", "workers.py"] but file is worker.pyFile name mismatch
4missing_configmapenvFrom: configMapRef: app-config but ConfigMap doesn't existCreateContainerConfigError
5liveness_probe_failingLiveness probe port 3000 but app listens on 8080Probe misconfiguration causes restarts

Task 8: Kubernetes Service & Ingress Issues — Hard

Networking issues where pods run fine but traffic doesn't reach them.

#ScenarioWhat's BrokenReal-World Context
1selector_mismatchService selector app: api but pod label is app: api-serverNo endpoints — most common K8s networking bug
2port_mismatchService targetPort 8080 but container listens on 3000Connection refused
3ingress_wrong_serviceIngress references api-svc but service name is api-serviceIngress 404
4network_policy_blockingNetworkPolicy with empty ingress rules blocks all trafficDatabase unreachable
5missing_ingress_classNo ingressClassName: nginx specifiedIngress controller doesn't pick it up

Task 9: CI/CD Build & Push Pipeline — Hard

GHA-to-Docker-to-Registry pipeline failures spanning multiple files.

#ScenarioWhat's BrokenReal-World Context
1ghcr_token_not_mapped$GITHUB_TOKEN shell var not mapped from secretsGHCR login fails
2image_tag_mismatchBuild uses github.ref_name but push uses github.sha"image not found locally"
3missing_packages_writeNo permissions: packages: write for GHCR push"permissiondenied: writepackage"
4build_arg_not_passedDockerfile ARG APP_VERSION but no --build-arg in workflowVersion file is empty
5multistage_output_mismatchCOPY --from=builder /app/dist but react-scripts outputs to /app/buildWrong output directory

Task 10: Full Stack Deployment Pipeline — Expert

Multi-error scenarios spanning the entire stack: GHA + Dockerfile + K8s manifests. 2-4 bugs per scenario requiring cross-file reasoning.

#ScenarioWhat's BrokenReal-World Context
1full_pipeline_ghcr_and_selectorGHCR token not mapped + K8s Service selector mismatch2 bugs across workflow + K8s
2full_pipeline_three_bugsMissing checkout + no WORKDIR + wrong container/service port4 bugs across 4 files
3full_pipeline_ghcr_dockerfile_k8sWrong GHCR secret + base image typo + OOM memory limit3 bugs across all layers
4full_pipeline_permissions_image_ingressMissing packages:write + hardcoded image placeholder + no ingressClassName3 bugs
5full_pipeline_secrets_build_probeDocker secrets not wired + wrong build output dir + probe port mismatch4 bugs across all layers

Available Actions

Each step, the agent chooses exactly one action:

ActionWhat It DoesWhen to Use
edit_fileReplace old_content with new_content in a fileMost common — fix a broken line or block
replace_lineReplace content at a specific line numberWhen you know exactly which line is wrong
add_lineInsert a new line into a fileAdding missing instructions (e.g., missing WORKDIR)
delete_lineRemove a specific lineRemoving a bad instruction
add_blockInsert a multi-line blockAdding entire sections (e.g., env: block with secrets)
delete_blockRemove a multi-line blockRemoving incorrect sections
request_hintGet a clue about what's wrongCosts -5% on final score — use sparingly
submitDeclare "I'm done" — triggers final evaluationWhen all fixes are applied

Important: edit_file requires old_content to match exactly (including whitespace). If it doesn't match, the edit fails and the agent gets a -0.02 reward penalty.


Grading System — How Scores Work

Scoring is deterministic (same actions always produce the same score) and dynamic (different strategies get different scores).

The Formula

FINAL SCORE = Base + Partial Fixes + Complete Bonus + Efficiency - Hint Penalty - Failed Edit Penalty

Clamped to (0.01, 0.99).

Component Breakdown

ComponentWeightDescription
Base score5%Participation credit
Partial fixes35%Proportional to issues_fixed / issues_total
Complete bonus25%All issues fixed
Efficiency25%Decays with extra steps beyond optimal
Hint penalty-4% eachPer request_hint action
Failed edit penalty-2% eachPer edit with no valid file path

API Endpoints

EndpointMethodDescription
/GETRoot page
/healthGETHealth check — returns {"status": "healthy"}
/metadataGETEnvironment name, description, version, tags
/schemaGETAction, observation, and state JSON schemas
/resetPOSTStart a new episode (optional: task_id, scenario_id, seed)
/stepPOSTTake an action and receive observation + reward
/stateGETGet current observation without taking an action
/infoGETTask list with metadata
/tasksGETList all tasks with difficulty levels
/graderPOSTGrade a trajectory (list of step dicts)
/baselinePOSTRun built-in heuristic baseline
/mcpPOSTJSON-RPC 2.0 MCP endpoint (initialize, tools/list)

Example: Full Episode via API

bash
# 1. Start an episode
curl -X POST http://localhost:8000/reset \
  -H "Content-Type: application/json" \
  -d '{"task_id": "k8s_pod_failures", "scenario_id": "oom_killed"}'

# 2. Fix the memory limit
curl -X POST http://localhost:8000/step \
  -H "Content-Type: application/json" \
  -d '{
    "action": {
      "action_type": "edit_file",
      "edits": [{
        "file_path": "k8s/deployment.yaml",
        "old_content": "memory: \"64Mi\"",
        "new_content": "memory: \"256Mi\""
      }]
    }
  }'

# Response: reward=0.3, issues_fixed=1/1, done=true

Quick Start

Local Development

bash
pip install -r requirements.txt
python -m uvicorn server.app:app --host 0.0.0.0 --port 8000

Run Tests

bash
pytest tests/ -v

Docker

bash
docker build -t cloud-native-devops-env .
docker run -p 8000:8000 cloud-native-devops-env

Baseline Inference (with LLM)

bash
export API_BASE_URL=https://router.huggingface.co/v1
export MODEL_NAME=meta-llama/Llama-3.1-70B-Instruct
export HF_TOKEN=your_token_here
python inference.py

Project Structure

cloud-native-devops-env/
├── openenv.yaml              # OpenEnv environment specification
├── inference.py              # LLM baseline (OpenAI client + HF router)
├── baseline_runner.py        # Heuristic baseline for /baseline endpoint
├── Dockerfile                # Production container
├── requirements.txt          # Python dependencies
│
├── server/
│   ├── app.py                # FastAPI with 12 endpoints
│   ├── models.py             # Pydantic models (type-safe API)
│   ├── environment.py        # Core environment loop (reset/step/state)
│   ├── tasks/
│   │   ├── base.py           # BaseTask with scenario loading
│   │   ├── task_registry.py  # Maps task_id → task class (10 tasks)
│   │   ├── task_1_build_errors.py        # 5 Dockerfile syntax scenarios
│   │   ├── task_2_docker_runtime.py      # 5 Dockerfile runtime scenarios
│   │   ├── task_3_workflow_syntax.py     # 5 workflow structure scenarios
│   │   ├── task_4_workflow_secrets_permissions.py  # 5 secrets scenarios
│   │   ├── task_5_ci_docker_integration.py        # 5 integration scenarios
│   │   ├── task_6_multi_stage_matrix.py           # 5 multi-issue scenarios
│   │   ├── k8s_pod.py                   # 5 Kubernetes pod failure scenarios
│   │   ├── k8s_networking.py            # 5 K8s networking scenarios
│   │   ├── pipeline_build_deploy.py     # 5 GHA→Docker→Registry scenarios
│   │   └── pipeline_full.py             # 5 full-stack multi-error scenarios
│   ├── graders/
│   │   └── __init__.py       # Deterministic trajectory grader
│   └── simulators/
│       ├── docker_simulator.py   # 15+ Dockerfile validation rules
│       ├── workflow_simulator.py # 15+ workflow validation rules
│       └── k8s_simulator.py     # Kubernetes manifest validator
│
└── tests/
    ├── test_endpoints.py     # API endpoint tests
    ├── test_determinism.py   # Grader determinism + score range tests
    ├── test_baseline.py      # Heuristic baseline tests
    ├── test_environment_flow.py  # Episode flow tests
    └── test_simulators.py    # Simulator unit tests

Design Decisions

  1. 1.Full cloud-native stack: Docker + GitHub Actions + Kubernetes — the three pillars of modern deployment pipelines.
  2. 2.Simulated validation (no real Docker/K8s): Static analysis rules give deterministic results, fast execution, and no security concerns.
  3. 3.Dense rewards: Partial credit at every step (+0.3 per fix, -0.02 per failed edit) rather than sparse pass/fail.
  4. 4.Difficulty progression: Easy tasks are single-file, single-issue. Expert tasks are multi-file, multi-issue with interacting bugs across all three layers.
  5. 5.Exact string matching for edits: Mirrors real file editing — whitespace matters.
  6. 6.50 scenarios from real bugs: Every scenario is based on actual developer mistakes documented on Stack Overflow, GitHub Issues, and official documentation.

License

MIT