abdulkalamazad07/disaster-response-rl-env
π DisasterResponseEnv: OpenEnv Hackathon Submission
The environment where the simulation IS the dataset. Learn to save lives when every second counts.
Hackathon Theme: #3.1 Professional Tasks & Long-Horizon Planning Goal: Train an LLM agent to efficiently allocate limited emergency resources (ambulances, rescue teams) across multiple disaster-stricken city zones under partial observability and tight deadlines.
π The Problem
During a crisis (like an earthquake or flood), real emergency operations centers face a massive coordination challenge:
- Information Overload: Continuous updates on casualties and hazards across multiple zones.
- Resource Scarcity: Never enough ambulances or rescue teams to cover everyone immediately.
- Cascading Failures: If you rescue everyone but the hospitals overflow, people still die. If you only focus on the worst-hit zone, other zones collapse.
DisasterResponseEnv is a deterministic, OpenEnv-compliant simulation that forces an LLM agent to learn these hard truths through pure Reinforcement Learning (GRPO). There is no human-labeled datasetβthe environment generates the ground truth.
ποΈ Environment Architecture
State (Partial Observability)
The agent sees a JSON representation of:
- Available resources (Ambulances, Rescue Teams).
- Hospital capacity.
- Zone statuses (Critical patients, Hazard levels, Consecutive turns without support).
Action Space
The agent outputs a structured JSON array deciding how many resources go where:
{
"allocations": [
{"zone": "Central", "ambulances": 2, "rescue_teams": 1},
{"zone": "North", "ambulances": 1, "rescue_teams": 0}
]
}The 5 Independent Reward Signals
To prevent reward hacking (e.g., an agent just sending all resources to one zone to get a "good enough" score), the environment calculates 5 competing rewards:
- R1 (Lives Saved): +1.5 for treated, -3.0 for deaths.
- R2 (Resource Efficiency): -0.5 per wasted ambulance sent to a zone with no patients.
- R3 (Coverage / Anti-Neglect): -2.0 penalty if any zone with patients is ignored for 3+ consecutive turns.
- R4 (Response Time): Rewards rapid reduction of hazard levels.
- R5 (Hospital Management): +1.0 for optimal utilization (50-85%), -2.0 for overcrowding.
π Results & Evidence of Learning
We generated synthetic SFT data using a greedy oracle heuristic, and then trained a small LLM (Qwen2.5-1.5B-Instruct via Unsloth) using TRL's GRPOTrainer directly against the environment's reward function.
Before vs After Training (Curriculum Level 2)
(Run `python demo/run_demo.py` to reproduce baseline metrics locally).
π» Running the Code
1. Try the Interactive Demo (HuggingFace Spaces format)
pip install -r requirements.txt
python app.pyThis opens a Gradio app where you can step through the simulation manually or test JSON actions.
2. Generate Warmup Data
python training/generate_sft_data.py3. Run the Training Script (Colab / GPU required)
python training/train_grpo.pyNote: Uses `unsloth` and `trl`. Ensure you are in a Linux/CUDA environment.
π Why This Matters
This environment goes beyond text generation. Every action mutates the state deterministically. Sending an ambulance to the North zone means the South zone waits longer. Overloading the hospital causes a bottleneck. By training an LLM on this environment, we are not teaching it to "sound like a dispatcher"; we are teaching it to make the mathematically optimal decisions of a dispatcher.
