training-monkey/dataoncallenv
0
1---2title: DataOnCallEnv3emoji: π4colorFrom: blue5colorTo: indigo6sdk: docker7pinned: true8tags:9 - openenv10base_path: /web11---12 13# DataOnCallEnv14 15An RL benchmark that simulates the workflow of an on-call data analyst debugging broken reports. The agent investigates realistic data pipeline bugs across three difficulty tiers β diagnosing root causes, writing corrected SQL, and earning a multi-dimensional reward score.16 17## Why This Exists18 19AI agents are being deployed as data analysts. But there is no benchmark that measures whether these agents can actually debug β not just query. DataOnCallEnv fills that gap by simulating exactly the workflow a real on-call analyst performs, scored against deterministic ground truth.20 21## Features22 23| Feature | Description |24|---------|-------------|25| Partial Observability | Tables are hidden at reset. Agent must call `list_tables()` to discover schema before inspecting or querying. |26| Query Cost Budget | Each tool has a fixed cost (`run_sql`=2.0, `inspect_schema`=1.0, etc). Total budget: 20.0 per episode. |27| Realistic Logs | Expanded dbt pipeline logs (8β15 rows per task) with noise entries + Airflow DAG run history table. |28| Anti-Cheat | `SELECT *` blocked, results capped at 50 rows, minimum 2 tool calls before submit, memorized-answer detection. |29| Tiered Evaluation | Diagnosis scored by depth of understanding (exact β category β symptom), not binary pass/fail. |30| Deterministic | All tasks and grading are fully deterministic. Same actions always produce the same score. |31 32## Tasks33 34| ID | Difficulty | Title | Root Cause | Optimal Steps | Optimal Cost |35|----|-----------|-------|------------|---------------|-------------|36| 1 | Easy | Revenue shows $0 for international sales | Currency code casing mismatch (`USD` vs `usd`) causes silent NULL JOIN | 5 | 7.0 |37| 2 | Medium | MAU dropped 8% on Feb 1st | UTCβlocal timezone migration double-counts events at month boundary | 7 | 10.0 |38| 3 | Hard | Cloud Storage revenue overstated by 3.7x | Non-unique key in `product_promotions` causes fanout on JOIN | 8 | 13.0 |39 40## Action Space41 42Every action has `tool`, `query`, and an optional `reasoning` field (rewarded by the grader).43 44| Tool | Query Format | Cost | Description |45|------|-------------|------|-------------|46| `list_tables` | `""` | 0.5 | Discover available tables (must be called first) |47| `inspect_schema` | `"table_name"` | 1.0 | Column names and types for a discovered table |48| `check_logs` | `""` | 1.0 | dbt pipeline changelog (ordered by most recent) |49| `check_airflow` | `""` | 1.0 | Airflow DAG run history |50| `run_sql` | `"SELECT col FROM ..."` | 2.0 | Execute a SELECT query (no `SELECT *`) |51| `diff_report` | `"date1,date2"` | 1.5 | Compare revenue totals between two dates |52| `submit` | `"ROOT CAUSE: ... CORRECTED SQL: ..."` | 0.0 | Submit diagnosis and fix. Ends episode. |53 54## Observation Space55 56```json57{58 "task_id": 1,59 "result": { "...tool output..." },60 "steps_taken": 3,61 "done": false,62 "max_steps": 15,63 "cost_spent": 4.5,64 "budget_remaining": 15.565}66```67 68## Reward Function69 70| Component | Range | Description |71|-----------|-------|-------------|72| `diagnosis_correct` | 0.00β0.25 | Tiered: exact root cause (0.25), category match (0.15), symptom only (0.08) |73| `fix_valid` | 0.00β0.25 | Agent's proposed SQL returns correct output vs ground truth |74| `efficiency` | 0.00β0.15 | Combined step + cost efficiency relative to optimal |75| `reasoning_quality` | 0.00β0.10 | Fraction of actions that include a reasoning field |76| `investigation_quality` | 0.00β0.10 | Logical methodology: discovery β schema β logs β hypothesis β verify |77| `false_positive_penalty` | 0.00β0.15 | Deducted for irrelevant table access, duplicate queries, or cheating |78 79**Total: 0.0β1.0** (dense reward β every action contributes signal)80 81## API Endpoints82 83| Method | Endpoint | Description |84|--------|----------|-------------|85| `GET` | `/health` | Health check, returns env metadata and version |86| `GET` | `/` | Root info with available endpoints |87| `GET` | `/tasks` | List all tasks with metadata |88| `GET` | `/state` | Full current episode state |89| `GET` | `/docs` | Interactive Swagger UI |90| `POST` | `/reset` | Start a fresh episode. Body: `{"task_id": 1}` |91| `POST` | `/step` | Send one action. Body: `{"tool": "...", "query": "...", "reasoning": "..."}` |92 93## Setup94 95### Prerequisites96 97- Python 3.10+98- pip99 100### Install and Run Locally101 102```bash103git clone https://github.com/ajaypushparaj5/dataoncallenv.git104cd dataoncallenv105pip install -r requirements.txt106 107# Start the API server108uvicorn api.app:app --reload --port 8000109```110 111### Quick API Test112 113```bash114# Health check115curl http://localhost:8000/health116 117# Start Task 1118curl -X POST http://localhost:8000/reset \119 -H "Content-Type: application/json" \120 -d '{"task_id": 1}'121 122# Discover tables123curl -X POST http://localhost:8000/step \124 -H "Content-Type: application/json" \125 -d '{"tool": "list_tables", "query": "", "reasoning": "Discover available tables"}'126 127# Inspect schema128curl -X POST http://localhost:8000/step \129 -H "Content-Type: application/json" \130 -d '{"tool": "inspect_schema", "query": "sales", "reasoning": "Check sales table structure"}'131```132 133### Run with Docker134 135```bash136docker build -t dataoncallenv .137docker run -p 7860:7860 dataoncallenv138```139 140### Run Baseline Inference141 142Requires an API key for an OpenAI-compatible inference provider.143 144```bash145# Create .env file146cat > .env << EOF147HF_TOKEN=your_hf_token_here148API_BASE_URL=https://router.huggingface.co/v1149MODEL_NAME=Qwen/Qwen2.5-72B-Instruct150EOF151 152python inference.py153```154 155### Run Tests156 157```bash158python test_env.py159# Expected: 36 passed, 0 failed160```161 162## Environment Variables163 164| Variable | Required | Default | Description |165|----------|----------|---------|-------------|166| `HF_TOKEN` | Yes | β | Hugging Face API token (used as `OPENAI_API_KEY`) |167| `API_BASE_URL` | No | `https://router.huggingface.co/v1` | OpenAI-compatible inference endpoint |168| `MODEL_NAME` | No | `Qwen/Qwen2.5-72B-Instruct` | Model identifier for the inference provider |169 170## Hugging Face Spaces Deployment171 1721. Create a new Space on [huggingface.co/new-space](https://huggingface.co/new-space)1732. Select **Docker** as the SDK1743. Upload the project files (or connect via Git)1754. Add `HF_TOKEN` as a secret in Space Settings1765. The Space will auto-build using the `Dockerfile` and expose the API on port 7860177 178## Project Structure179 180```181dataoncallenv/182βββ models.py # Pydantic types: Action, Observation, Reward, EnvState183βββ tasks.py # Task definitions with ground truth and diagnosis tiers184βββ database.py # SQLite database builder, tool implementations, anti-cheat185βββ environment.py # Core RL env: partial obs, query costs, step/reset/state186βββ graders.py # Tiered scoring, investigation quality, penalties187βββ inference.py # Baseline agent using OpenAI-compatible API188βββ test_env.py # 36 integration tests189βββ api/190β βββ app.py # FastAPI server with /health, /reset, /step, /state191βββ openenv.yaml # OpenEnv spec metadata192βββ Dockerfile # HF Spaces container (port 7860)193βββ requirements.txt # Python dependencies194βββ baseline_scores.json # Reproducible baseline results195βββ .env # API credentials (not committed)196```197 198## OpenEnv Spec199 200This environment implements the full [OpenEnv](https://github.com/open-env/openenv) specification:201- Typed Pydantic models for `Action`, `Observation`, `Reward`, and `EnvState`202- `reset()` / `step()` / `state()` API203- `openenv.yaml` with full metadata204- Minimum 3 tasks with agent graders (easy β medium β hard)205- Scores in range 0.0β1.0 with partial progress signals206- Deterministic evaluation207- Baseline inference script with reproducible scores208 209## Anti-Cheat Constraints210 211| Constraint | Effect |212|------------|--------|213| `SELECT *` blocked | Agent must specify columns explicitly |214| Row cap (50) | Large result dumps are truncated |215| Min 2 tools before submit | Prevents skipping investigation |216| Memorized-answer detection | Correct diagnosis without investigation incurs penalty |217| Duplicate query penalty | Repeated identical queries are penalized |218| Irrelevant table penalty | Accessing tables unrelated to the task is penalized |219 