YashCanCode/drug-interaction-safety-agent
Drug Interaction Safety Agent
OpenEnv environment · Meta Hackathon 2026
A real-world clinical evaluation environment where AI agents must detect, classify, and safely manage drug-drug interactions in patient medication regimens. Grounded in peer-reviewed pharmacology with 35+ interaction pairs across 50+ drugs.
Motivation
Polypharmacy (taking 5+ medications simultaneously) affects over 40% of adults over 65 and is a leading cause of preventable hospitalizations. Drug-drug interactions range from minor nuisances to life-threatening emergencies. This environment evaluates whether language agents can reason about these interactions with clinical accuracy — a task that genuinely challenges even experienced clinicians.
Environment Overview
Three progressive tasks model a realistic clinical workflow:
Task Descriptions
Task Easy — Binary Interaction Detection
Given a patient profile and two specific medications, determine whether a clinically significant drug-drug interaction exists. This is a binary yes/no judgment.
Example: Warfarin + Ibuprofen → interaction_detected: true (NSAID increases bleeding risk with anticoagulants)
Scoring: 1.0 for correct detection, 0.0 for incorrect. No partial credit.
Task Medium — Severity Classification
Given a confirmed drug interaction (mechanism and clinical effect are revealed, severity is hidden), classify the severity and recommend an appropriate clinical action.
Severity levels: none → minor → moderate → major → contraindicated
Scoring: severity_score × 0.7 + action_score × 0.3
- Severity: 1.0 exact match, 0.5 adjacent level, 0.0 off by 2+
- Action: 1.0 if optimal, 0.0 otherwise
Task Hard — Full Polypharmacy Regimen Review
Given a full medication regimen of 5–7 drugs, identify all clinically significant interactions, classify each by severity, determine the highest-risk pair, recommend the appropriate clinical action, and suggest a safer alternative for the most dangerous drug.
Multi-step (up to 3 steps): The agent receives feedback after each step and can refine its answer.
Scoring:
interaction_recall × 0.35
+ severity_accuracy × 0.35
+ action_match × 0.20
+ alternative_valid × 0.10
- penalties (−0.30 per missed contraindicated pair)
- hint_penalty (−0.05 per hint used)Partial rewards: +0.15 per correctly identified interaction (mid-episode)
Action Space
All fields are optional at the schema level; at least one meaningful field must be set.
{
"interaction_detected": true,
"severity_assessment": "major",
"interactions_found": [
{
"drug_a": "warfarin",
"drug_b": "amiodarone",
"severity": "major",
"mechanism": "CYP2C9 inhibition increases warfarin levels"
}
],
"highest_severity": "major",
"recommended_action": "consult_prescriber",
"alternative_drug": "rosuvastatin",
"rationale": "Amiodarone inhibits CYP2C9, significantly raising warfarin INR."
}Allowed values:
severity:none|minor|moderate|major|contraindicatedrecommended_action:continue|reduce_dose|switch_drug|consult_prescriber|monitor_closely
Observation Space
{
"task_id": "task_hard",
"step": 1,
"patient": {
"age": 71,
"weight_kg": 84.0,
"renal_function": "mild",
"hepatic_function": "normal",
"diagnoses": ["post-MI", "atrial fibrillation", "type 2 diabetes"],
"allergies": []
},
"medications": [
{"name": "warfarin", "dose": "5mg", "frequency": "once daily", "route": "oral"},
{"name": "amiodarone", "dose": "200mg", "frequency": "once daily", "route": "oral"},
{"name": "simvastatin","dose": "40mg", "frequency": "once daily", "route": "oral"}
],
"drug_pair": null,
"flagged_interaction": null,
"feedback": "✓ +0.15 — Newly identified: warfarin+amiodarone",
"partial_score": 0.15,
"signals_identified": ["warfarin+amiodarone"],
"hints_available": ["Hint 1: Check interactions involving warfarin and amiodarone"],
"done": false,
"reward": null
}API Reference
Setup & Usage
Local (Python)
# Install dependencies
pip install -r requirements.txt
# Start the server
uvicorn server:app --host 0.0.0.0 --port 7860
# Verify
curl http://localhost:7860/health
# List tasks
curl http://localhost:7860/tasksDocker
# Build
docker build -t drug-interaction-env .
# Run
docker run -p 7860:7860 drug-interaction-env
# Test
curl http://localhost:7860/healthQuick API walkthrough
# 1. Start an episode
curl -X POST http://localhost:7860/reset \
-H "Content-Type: application/json" \
-d '{"task_id": "task_easy", "seed": 0}'
# → {"episode_id": "abc-123", "observation": {...}}
# 2. Submit an action
curl -X POST http://localhost:7860/step/abc-123 \
-H "Content-Type: application/json" \
-d '{"interaction_detected": true, "rationale": "NSAID + anticoagulant = bleeding risk"}'
# → {"observation": {...}, "reward": 1.0, "done": true, "info": {...}}
# 3. Get reward breakdown
curl http://localhost:7860/grader/abc-123Running the Baseline
Oracle baseline (perfect pharmacological knowledge — for environment verification)
# Via server
curl "http://localhost:7860/baseline?runs_per_task=3"
# Direct (no server needed)
python agent.py --all --directLLM baseline (OpenAI GPT-4o)
export OPENAI_API_KEY=sk-...
# All tasks, 3 runs each
python baseline.py --all
# Single task, 5 runs
python baseline.py --task task_hard --runs 5
# Against a specific server
python baseline.py --all --server http://localhost:7860
# Use a different model
OPENAI_MODEL=gpt-4o-mini python baseline.py --all
# Save results to file
python baseline.py --all --out results.jsonBaseline Scores
Oracle Agent (expected — verifies environment correctness)
Note: task_hard oracle may score < 1.0 on scenarios where the safe alternative isn't in the drug database.
GPT-4o Baseline (representative — actual scores depend on API version)
Project Structure
drug-interaction-safety-agent/
├── drug_database.py # Ground-truth pharmacological knowledge base
├── models.py # Pydantic models: Action, Observation, State, Reward
├── tasks.py # Task definitions, scenario banks, samplers
├── environment.py # Core env: reset(), step(), get_state(), grade()
├── server.py # FastAPI server (all REST endpoints)
├── agent.py # OracleAgent + ClaudeAgent
├── baseline.py # OpenAI baseline inference script
├── openenv.yaml # OpenEnv spec metadata
├── Dockerfile # HuggingFace Spaces compatible container
├── requirements.txt # Python dependencies
└── README.md # This fileOpenEnv Compliance
- ✅ Typed Pydantic models for Action, Observation, State, Reward
- ✅
reset()→ clean initial observation - ✅
step(action)→{observation, reward, done, info} - ✅
state()→ full internal state - ✅ 3+ tasks with graders scoring 0.0–1.0
- ✅ Deterministic and reproducible (seed-controlled)
- ✅
openenv.yamlmetadata - ✅ Working Dockerfile (HF Spaces port 7860)
- ✅ Baseline inference script using OpenAI client
Drug Database Coverage
The environment's pharmacological ground truth covers clinically validated interactions including:
- Anticoagulant interactions: warfarin + NSAIDs, fluoroquinolones, amiodarone, statins
- Serotonin syndrome: SSRIs + tramadol, MAOIs, triptans
- QT prolongation: digoxin + verapamil, amiodarone + macrolides
- CYP enzyme interactions: simvastatin + clarithromycin, metoprolol + paroxetine
- Immunosuppressant interactions: tacrolimus + fluconazole, phenytoin combinations
- Antiplatelet antagonism: aspirin + ibuprofen, clopidogrel + omeprazole
All severities and recommended actions are grounded in standard clinical pharmacology references.
License
MIT License — see LICENSE for details.
