Team Ai
Apppublic

YashCanCode/drug-interaction-safety-agent

sourceHugging Faceupdated 6mo agoView on Hugging Face
1likes
App README

Drug Interaction Safety Agent

OpenEnv environment · Meta Hackathon 2026

A real-world clinical evaluation environment where AI agents must detect, classify, and safely manage drug-drug interactions in patient medication regimens. Grounded in peer-reviewed pharmacology with 35+ interaction pairs across 50+ drugs.


Motivation

Polypharmacy (taking 5+ medications simultaneously) affects over 40% of adults over 65 and is a leading cause of preventable hospitalizations. Drug-drug interactions range from minor nuisances to life-threatening emergencies. This environment evaluates whether language agents can reason about these interactions with clinical accuracy — a task that genuinely challenges even experienced clinicians.


Environment Overview

Three progressive tasks model a realistic clinical workflow:

TaskNameDifficultyMax StepsPass Threshold
task_easySingle drug-pair detectionEasy11.0
task_mediumSeverity classificationMedium10.5
task_hardFull regimen safety reviewHard30.5

Task Descriptions

Task Easy — Binary Interaction Detection

Given a patient profile and two specific medications, determine whether a clinically significant drug-drug interaction exists. This is a binary yes/no judgment.

Example: Warfarin + Ibuprofen → interaction_detected: true (NSAID increases bleeding risk with anticoagulants)

Scoring: 1.0 for correct detection, 0.0 for incorrect. No partial credit.


Task Medium — Severity Classification

Given a confirmed drug interaction (mechanism and clinical effect are revealed, severity is hidden), classify the severity and recommend an appropriate clinical action.

Severity levels: none → minor → moderate → major → contraindicated

Scoring: severity_score × 0.7 + action_score × 0.3

  • —Severity: 1.0 exact match, 0.5 adjacent level, 0.0 off by 2+
  • —Action: 1.0 if optimal, 0.0 otherwise

Task Hard — Full Polypharmacy Regimen Review

Given a full medication regimen of 5–7 drugs, identify all clinically significant interactions, classify each by severity, determine the highest-risk pair, recommend the appropriate clinical action, and suggest a safer alternative for the most dangerous drug.

Multi-step (up to 3 steps): The agent receives feedback after each step and can refine its answer.

Scoring:

interaction_recall × 0.35
+ severity_accuracy × 0.35
+ action_match × 0.20
+ alternative_valid × 0.10
- penalties (−0.30 per missed contraindicated pair)
- hint_penalty (−0.05 per hint used)

Partial rewards: +0.15 per correctly identified interaction (mid-episode)


Action Space

All fields are optional at the schema level; at least one meaningful field must be set.

json
{
  "interaction_detected": true,
  "severity_assessment": "major",
  "interactions_found": [
    {
      "drug_a": "warfarin",
      "drug_b": "amiodarone",
      "severity": "major",
      "mechanism": "CYP2C9 inhibition increases warfarin levels"
    }
  ],
  "highest_severity": "major",
  "recommended_action": "consult_prescriber",
  "alternative_drug": "rosuvastatin",
  "rationale": "Amiodarone inhibits CYP2C9, significantly raising warfarin INR."
}

Allowed values:

  • —severity: none | minor | moderate | major | contraindicated
  • —recommended_action: continue | reduce_dose | switch_drug | consult_prescriber | monitor_closely

Observation Space

json
{
  "task_id": "task_hard",
  "step": 1,
  "patient": {
    "age": 71,
    "weight_kg": 84.0,
    "renal_function": "mild",
    "hepatic_function": "normal",
    "diagnoses": ["post-MI", "atrial fibrillation", "type 2 diabetes"],
    "allergies": []
  },
  "medications": [
    {"name": "warfarin",   "dose": "5mg",    "frequency": "once daily",  "route": "oral"},
    {"name": "amiodarone", "dose": "200mg",  "frequency": "once daily",  "route": "oral"},
    {"name": "simvastatin","dose": "40mg",   "frequency": "once daily",  "route": "oral"}
  ],
  "drug_pair": null,
  "flagged_interaction": null,
  "feedback": "✓ +0.15 — Newly identified: warfarin+amiodarone",
  "partial_score": 0.15,
  "signals_identified": ["warfarin+amiodarone"],
  "hints_available": ["Hint 1: Check interactions involving warfarin and amiodarone"],
  "done": false,
  "reward": null
}

API Reference

EndpointMethodDescription
/healthGETLiveness probe
/tasksGETTask list + action schemas
/resetPOSTStart new episode → {episode_id, observation}
/step/{episode_id}POSTSubmit action → {observation, reward, done, info}
/state/{episode_id}GETFull internal state (ground truth — for grader only)
/grader/{episode_id}GETDetailed reward breakdown
/baselineGETRun oracle agent, return per-task scores

Setup & Usage

Local (Python)

bash
# Install dependencies
pip install -r requirements.txt

# Start the server
uvicorn server:app --host 0.0.0.0 --port 7860

# Verify
curl http://localhost:7860/health

# List tasks
curl http://localhost:7860/tasks

Docker

bash
# Build
docker build -t drug-interaction-env .

# Run
docker run -p 7860:7860 drug-interaction-env

# Test
curl http://localhost:7860/health

Quick API walkthrough

bash
# 1. Start an episode
curl -X POST http://localhost:7860/reset \
  -H "Content-Type: application/json" \
  -d '{"task_id": "task_easy", "seed": 0}'

# → {"episode_id": "abc-123", "observation": {...}}

# 2. Submit an action
curl -X POST http://localhost:7860/step/abc-123 \
  -H "Content-Type: application/json" \
  -d '{"interaction_detected": true, "rationale": "NSAID + anticoagulant = bleeding risk"}'

# → {"observation": {...}, "reward": 1.0, "done": true, "info": {...}}

# 3. Get reward breakdown
curl http://localhost:7860/grader/abc-123

Running the Baseline

Oracle baseline (perfect pharmacological knowledge — for environment verification)

bash
# Via server
curl "http://localhost:7860/baseline?runs_per_task=3"

# Direct (no server needed)
python agent.py --all --direct

LLM baseline (OpenAI GPT-4o)

bash
export OPENAI_API_KEY=sk-...

# All tasks, 3 runs each
python baseline.py --all

# Single task, 5 runs
python baseline.py --task task_hard --runs 5

# Against a specific server
python baseline.py --all --server http://localhost:7860

# Use a different model
OPENAI_MODEL=gpt-4o-mini python baseline.py --all

# Save results to file
python baseline.py --all --out results.json

Baseline Scores

Oracle Agent (expected — verifies environment correctness)

TaskMean ScorePass Rate
task_easy1.000100%
task_medium1.000100%
task_hard~0.90100%

Note: task_hard oracle may score < 1.0 on scenarios where the safe alternative isn't in the drug database.

GPT-4o Baseline (representative — actual scores depend on API version)

TaskMean ScorePass Rate
task_easy~0.87~87%
task_medium~0.72~83%
task_hard~0.55~67%

Project Structure

drug-interaction-safety-agent/
├── drug_database.py     # Ground-truth pharmacological knowledge base
├── models.py            # Pydantic models: Action, Observation, State, Reward
├── tasks.py             # Task definitions, scenario banks, samplers
├── environment.py       # Core env: reset(), step(), get_state(), grade()
├── server.py            # FastAPI server (all REST endpoints)
├── agent.py             # OracleAgent + ClaudeAgent
├── baseline.py          # OpenAI baseline inference script
├── openenv.yaml         # OpenEnv spec metadata
├── Dockerfile           # HuggingFace Spaces compatible container
├── requirements.txt     # Python dependencies
└── README.md            # This file

OpenEnv Compliance

  • —✅ Typed Pydantic models for Action, Observation, State, Reward
  • —✅ reset() → clean initial observation
  • —✅ step(action) → {observation, reward, done, info}
  • —✅ state() → full internal state
  • —✅ 3+ tasks with graders scoring 0.0–1.0
  • —✅ Deterministic and reproducible (seed-controlled)
  • —✅ openenv.yaml metadata
  • —✅ Working Dockerfile (HF Spaces port 7860)
  • —✅ Baseline inference script using OpenAI client

Drug Database Coverage

The environment's pharmacological ground truth covers clinically validated interactions including:

  • —Anticoagulant interactions: warfarin + NSAIDs, fluoroquinolones, amiodarone, statins
  • —Serotonin syndrome: SSRIs + tramadol, MAOIs, triptans
  • —QT prolongation: digoxin + verapamil, amiodarone + macrolides
  • —CYP enzyme interactions: simvastatin + clarithromycin, metoprolol + paroxetine
  • —Immunosuppressant interactions: tacrolimus + fluconazole, phenytoin combinations
  • —Antiplatelet antagonism: aspirin + ibuprofen, clopidogrel + omeprazole

All severities and recommended actions are grounded in standard clinical pharmacology references.


License

MIT License — see LICENSE for details.