Team Ai
Apppublic

niha647/api-debug-environment

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

🔧 API Integration Debugger OpenEnv

An OpenEnv-compliant reinforcement learning environment where an agent acts as an Integration Engineer debugging malformed HTTP requests.

🌟 Motivation

API integration is a cornerstone of modern software development. Debugging broken endpoints—whether due to typos, missing authentication, or complex schema mismatches—requires dynamic reasoning and an iterative "test-and-fix" cycle.

This environment provides a deterministic, mathematically grounded benchmark for how well an LLM can:

  1. 1.Parse server error responses.
  2. 2.Formulate hypotheses about request fixes based on diagnostic hints.
  3. 3.Utilize structured MCP (Model Context Protocol) tools to modify and resubmit requests.
  4. 4.Correct cascading failures (404 → 401 → 400).

🏗️ Environment Description

The agent lives in a loop with a Mock API Server. Every episode initializes with a "broken" request configuration. The agent modifies headers, URL parameters, methods, and the JSON body until the server returns a 200 OK.

To prevent simple memorization, the environment uses Procedural Generation. Every episode features randomized endpoints, randomized authentication tokens/keys, and randomized required JSON body fields.


🛠️ Action Space

The environment exposes 7 structured MCP tools that the agent can invoke:

  • —set_url(url: str): Updates the target API endpoint.
  • —set_method(method: str): Changes the HTTP verb (GET, POST, PUT, DELETE).
  • —set_header(key: str, value: str): Sets or updates a request header.
  • —remove_header(key: str): Deletes a specific header.
  • —set_body(body: str): Sets the raw JSON string payload for the request.
  • —set_query_param(key: str, value: str): Adds or updates a URL query parameter.
  • —send_request(): Submits the current configuration and retrieves a server response.

👁️ Observation Space

After every step, the agent receives a rich state observation:

  • —status_code (int): The HTTP status returned by the mock server (e.g., 404, 401, 400).
  • —response_body (dict): The JSON payload from the server, including error and diagnostic hint strings.
  • —response_headers (dict): Standard HTTP headers from the response (including X-Request-Id).
  • —current_request (dict): A complete snapshot of the agent's current URL, Method, Headers, and Body.
  • —step_number / max_steps (int): Information on the remaining budget for the episode.
  • —reward (float): The intermediate reward signal.
  • —done (bool): Termination flag.

Tasks

  1. 1.Easy (fix_endpoint) - Agent detects a 404 and corrects the randomized typo in the /users endpoint.
  2. 2.Easy/Med (method_mismatch) - Target requires a random method (e.g., POST, PUT, DELETE), but the agent starts with the wrong one (e.g., GET) and hits a 405 Method Not Allowed.
  3. 3.Medium (fix_auth) - Agent detects a 401 and implements the generated Auth scheme (Bearer/Basic/Key) by reading server hints.
  4. 4.Medium (query_param_debug) - Target requires a specific query parameter (e.g., ?version=2.0). The agent hits a 400 and must supply it.
  5. 5.Hard (cascading_debug) - A complex schema validation: 404 followed by 401 followed by 400 where the agent has to parse missing JSON keys and supply them dynamically.

🚀 Setup & Usage

Local Execution

  1. 1.Install dependencies:
bash
   pip install -r server/requirements.txt
  1. 1.Run the baseline evaluation script:
bash
   python baseline.py

Docker Deployment

  1. 1.Build the container:
bash
   docker build -t api-debug-env .
  1. 1.Run the server:
bash
   docker run -p 8000:8000 api-debug-env

The server will be available at http://localhost:8000.


📊 Baseline Scores

The provided baseline.py script uses a rule-based agent that parses error hints via regex.

TaskLevelSuccess RateAverage Reward
fix_endpointEasyHigh~0.85
fix_authMediumHigh~0.85
cascading_debugHardMedium~0.55

Note: The reward includes a -0.05 penalty per step to encourage efficiency.