niha647/api-debug-environment
🔧 API Integration Debugger OpenEnv
An OpenEnv-compliant reinforcement learning environment where an agent acts as an Integration Engineer debugging malformed HTTP requests.
🌟 Motivation
API integration is a cornerstone of modern software development. Debugging broken endpoints—whether due to typos, missing authentication, or complex schema mismatches—requires dynamic reasoning and an iterative "test-and-fix" cycle.
This environment provides a deterministic, mathematically grounded benchmark for how well an LLM can:
- Parse server error responses.
- Formulate hypotheses about request fixes based on diagnostic hints.
- Utilize structured MCP (Model Context Protocol) tools to modify and resubmit requests.
- Correct cascading failures (404 → 401 → 400).
🏗️ Environment Description
The agent lives in a loop with a Mock API Server. Every episode initializes with a "broken" request configuration. The agent modifies headers, URL parameters, methods, and the JSON body until the server returns a 200 OK.
To prevent simple memorization, the environment uses Procedural Generation. Every episode features randomized endpoints, randomized authentication tokens/keys, and randomized required JSON body fields.
🛠️ Action Space
The environment exposes 7 structured MCP tools that the agent can invoke:
set_url(url: str): Updates the target API endpoint.set_method(method: str): Changes the HTTP verb (GET, POST, PUT, DELETE).set_header(key: str, value: str): Sets or updates a request header.remove_header(key: str): Deletes a specific header.set_body(body: str): Sets the raw JSON string payload for the request.set_query_param(key: str, value: str): Adds or updates a URL query parameter.send_request(): Submits the current configuration and retrieves a server response.
👁️ Observation Space
After every step, the agent receives a rich state observation:
status_code(int): The HTTP status returned by the mock server (e.g., 404, 401, 400).response_body(dict): The JSON payload from the server, includingerrorand diagnostichintstrings.response_headers(dict): Standard HTTP headers from the response (includingX-Request-Id).current_request(dict): A complete snapshot of the agent's current URL, Method, Headers, and Body.step_number/max_steps(int): Information on the remaining budget for the episode.reward(float): The intermediate reward signal.done(bool): Termination flag.
Tasks
- Easy (
fix_endpoint) - Agent detects a404and corrects the randomized typo in the/usersendpoint. - Easy/Med (
method_mismatch) - Target requires a random method (e.g., POST, PUT, DELETE), but the agent starts with the wrong one (e.g., GET) and hits a405 Method Not Allowed. - Medium (
fix_auth) - Agent detects a401and implements the generated Auth scheme (Bearer/Basic/Key) by reading server hints. - Medium (
query_param_debug) - Target requires a specific query parameter (e.g.,?version=2.0). The agent hits a 400 and must supply it. - Hard (
cascading_debug) - A complex schema validation: 404 followed by 401 followed by 400 where the agent has to parse missing JSON keys and supply them dynamically.
🚀 Setup & Usage
Local Execution
- Install dependencies:
pip install -r server/requirements.txt- Run the baseline evaluation script:
python baseline.pyDocker Deployment
- Build the container:
docker build -t api-debug-env .- Run the server:
docker run -p 8000:8000 api-debug-env The server will be available at http://localhost:8000.
📊 Baseline Scores
The provided baseline.py script uses a rule-based agent that parses error hints via regex.
Note: The reward includes a -0.05 penalty per step to encourage efficiency.
