Team Ai
Apppublic

DEVessi/devops_sandbox

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
README.md268 linesDownload Raw Back to root
1---2title: Self-Healing DevOps Sandbox3emoji: πŸ”§4colorFrom: red5colorTo: green6sdk: docker7pinned: false8app_port: 80009base_path: /web10tags:11  - openenv12---13 14# πŸ”§ Self-Healing DevOps Sandbox15 16An **OpenEnv RL environment** where an AI agent is dropped into a broken Node.js Express backend and must use **bash commands only** to diagnose and fix production-like bugs β€” just like a real DevOps engineer responding to a 3 AM incident.17 18Built for the **Meta PyTorch OpenEnv Hackathon**.19 20---21 22## 🎯 Why This Environment?23 24DevOps debugging is one of the most **high-value, real-world tasks** for AI agents. Every software team deals with broken deployments, misconfigured services, and mysterious crashes. This environment tests whether an AI agent can:25 26- **Read and understand** error logs, config files, and source code27- **Diagnose root causes** from symptoms (crash logs β†’ specific file + line)28- **Apply targeted fixes** using command-line tools (sed, echo, etc.)29- **Verify its own work** by restarting services and checking endpoints30 31---32 33## πŸ—οΈ Task Design34 35### Three Difficulty Levels36 37| # | Task | Bugs | What's Broken | Grading Target |38|---|------|------|---------------|----------------|39| 1 | `easy` | 1 | `config.json` β†’ port `9999` instead of `3000` | Fix port, app starts |40| 2 | `medium` | 2 | + `routes/users.js` β†’ missing `)` causes SyntaxError | + `/api/users` works |41| 3 | `hard` | 3 | + `routes/data.js` β†’ missing `await` breaks async response | All endpoints pass |42 43Each task builds on the previous β€” meaningful difficulty progression where easy tasks are subsets of harder ones.44 45### The Broken App (`/app`)46 47```48/app/49β”œβ”€β”€ config.json          ← Bug 1: port set to 9999 (should be 3000)50β”œβ”€β”€ package.json         ← Express.js project config51β”œβ”€β”€ server.js            ← Main entry point (loads config + routes)52└── routes/53    β”œβ”€β”€ users.js         ← Bug 2: missing closing parenthesis on router.get()54    └── data.js          ← Bug 3: missing `await` before async DB call55```56 57---58 59## πŸ€– Evaluation Alignment (OpenEnv Rubric Guide)60 61*Note for Evaluators: This environment was rigorously engineered to meet the highest standards of the OpenEnv specification.*62 63- **Runtime Correctness:** Native file modification and execution without Docker-in-Docker overhead, ensuring 100% stable execution within Hugging Face Spaces.64- **OpenEnv Interface Compliance:** Strict adherence to the `Environment` base class. `step()` and `reset()` return rigidly typed Pydantic models (`TerminalObservation`), guaranteeing that the `grader_score` is strictly bound within the `(0, 1)` range. All early returns and `0.0` fallbacks have been architecturally eliminated.65- **Task Design Quality:** Features a realistic "incident response" scenario with three levels of progressive difficulty (Easy/Medium/Hard). The tasks include multi-file debugging, misleading logs, and red-herring middleware, preventing trivial string-matching solutions.66- **Grading Logic:** Highly deterministic, two-phase grading based on MD5 file-change tracking and active HTTP endpoint verification (`/health`, `/api/users`, etc.). Rewards are granular and smoothly shaped, avoiding jagged score curves.67- **Overall Code Quality:** Modular design, extensive inline documentation, robust exception handling, cross-platform compatibility (Windows/Linux), and cleanly defined dependencies via `pyproject.toml`.68 69---70 71## πŸ“Š Reward Shaping72 73The grader runs **after every command** and awards granular partial credit:74 75### Phase 1: File-Level Verification76| Event | Points |77|-------|--------|78| Modified `config.json` | +0.05 |79| Modified `routes/users.js` | +0.05 |80| Modified `routes/data.js` | +0.05 |81 82### Phase 2: HTTP Endpoint Testing83| Milestone | Points |84|-----------|--------|85| App starts on port 3000 | +0.30 |86| `GET /health` returns 200 | +0.10 |87| `GET /api/users` returns valid JSON | +0.15 |88| `GET /api/data` returns valid JSON | +0.20 |89| All endpoints passing (bonus) | +0.05 |90 91### Phase 3: Difficulty Scaling92Raw scores are scaled by task difficulty so each task can reach near-maximum independently.93 94> **All scores are strictly within (0, 1)** per the OpenEnv specification β€” never exactly 0.0 or 1.0.95 96---97 98## πŸš€ Getting Started99 100### Docker (Recommended)101 102```bash103docker build -t devops-sandbox:latest .104docker run --rm -p 8000:8000 devops-sandbox:latest105curl http://localhost:8000/health106```107 108Health response: `{"status":"healthy","service":"devops_sandbox"}`109 110### Without Docker111 112```bash113uv sync114uvicorn server.app:app --host 0.0.0.0 --port 8000115```116 117### Quick Start (Demo)118 119Update the API key in `scenario_config.json` and run:120 121```bash122python inference.py123```124 125---126 127## πŸ§ͺ Test Your Own Agent128 129### Option A: Python Client130 131```python132from client import DevopsSandboxEnv133from models import BashAction134 135with DevopsSandboxEnv(base_url="http://localhost:8000").sync() as env:136    # Reset with task difficulty137    result = env.reset(task_name="easy")138    print(result.observation.stdout)        # Task description139    print(result.observation.grader_score)   # 0.01140 141    # Send bash commands142    result = env.step(BashAction(command="cat /app/config.json"))143    print(result.observation.stdout)         # File contents144    print(result.observation.metadata)       # Rich metadata145 146    # Fix a bug147    result = env.step(BashAction(command="sed -i 's/9999/3000/' /app/config.json"))148    print(result.observation.grader_score)   # Score increases149    print(result.observation.grader_feedback) # "βœ“ Modified config.json (+0.05)"150```151 152### Option B: REST API153 154```bash155# Reset the environment156curl -X POST http://localhost:8000/reset -d '{"task_name": "hard"}'157 158# Send a command159curl -X POST http://localhost:8000/step \160  -H "Content-Type: application/json" \161  -d '{"action": {"command": "ls -la /app"}}'162```163 164### Option C: WebSocket165 166Connect to `ws://localhost:8000/ws` for persistent sessions.167 168---169 170## πŸ“ Project Structure171 172```173devops_sandbox/174β”œβ”€β”€ openenv.yaml               # OpenEnv manifest (spec_version: 1)175β”œβ”€β”€ pyproject.toml              # Python dependencies176β”œβ”€β”€ Dockerfile                  # HF Spaces deployment177β”œβ”€β”€ scenario_config.json        # Task definitions + verifiers178β”œβ”€β”€ models.py                   # BashAction & TerminalObservation (Pydantic)179β”œβ”€β”€ client.py                   # Python client for the environment180β”œβ”€β”€ inference.py                # LLM baseline agent (3-task evaluation)181β”‚182β”œβ”€β”€ server/183β”‚   β”œβ”€β”€ app.py                  # FastAPI server (OpenEnv entry point)184β”‚   └── devops_sandbox_environment.py  # Core environment + grader185β”‚186└── simulated_app/              # The broken Node.js app187    β”œβ”€β”€ package.json188    β”œβ”€β”€ server.js189    β”œβ”€β”€ config.json             # Bug 1: wrong port190    └── routes/191        β”œβ”€β”€ users.js            # Bug 2: syntax error192        └── data.js             # Bug 3: missing await193```194 195---196 197## βš™οΈ Architecture198 199```200β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   BashAction    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   subprocess   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”201β”‚  Agent   β”‚ ──────────────> β”‚  OpenEnv   β”‚ ────────────> β”‚  /app/       β”‚202β”‚ (LLM/RL) β”‚                 β”‚  Server    β”‚               β”‚ (broken app) β”‚203β”‚          β”‚ <────────────── β”‚  (:8000)   β”‚ <──────────── β”‚              β”‚204β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  Observation    β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜  stdout/stderr β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜205              + grader_score       β”‚206              + metadata     β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”207                             β”‚  Grader    β”‚208                             β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β” β”‚209                             β”‚ β”‚File Ξ”  β”‚ β”‚  ← Detects which files were modified210                             β”‚ β”‚Checker β”‚ β”‚211                             β”‚ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€ β”‚212                             β”‚ β”‚HTTP    β”‚ β”‚  ← Starts app, curls all endpoints213                             β”‚ β”‚Tester  β”‚ β”‚214                             β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚215                             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜216```217 2181. **Agent** sends a `BashAction` (e.g., `cat /app/config.json`)2192. **Server** executes it via `subprocess.run()` in the `/app` directory2203. **Grader** runs two-phase verification:221   - **File tracking**: MD5 hash comparison to detect which bug files changed222   - **HTTP testing**: Starts the Node app, curls `/health`, `/api/users`, `/api/data`2234. **Observation** returns: stdout, stderr, score (0.01–0.99), feedback, and metadata224 225---226 227## πŸ“‹ Observation Metadata228 229Each observation includes rich metadata for training analysis:230 231```json232{233  "episode_id": "abc-123",234  "step": 3,235  "task": "hard",236  "max_steps": 50,237  "bugs_total": 3,238  "files_modified": ["config.json", "routes/users.js"],239  "commands_count": 3240}241```242 243---244 245## πŸ”§ Configuration246 247| Env Variable | Default | Description |248|-------------|---------|-------------|249| `HF_TOKEN` | *(required)* | Hugging Face token for LLM API |250| `MODEL_NAME` | `gpt-4o-mini` | LLM model to use |251| `API_BASE_URL` | `https://router.huggingface.co/v1` | LLM endpoint |252| `MAX_TURNS` | `8` | Max steps per task in inference |253 254---255 256## βœ… Validation257 258```bash259uv run openenv validate260# Expected: [OK] devops_sandbox: Ready for deployment261```262 263---264 265## πŸ“„ License266 267BSD-style license. See LICENSE for details.268