Team Ai
Apppublic

abdulkalamazad07/disaster-response-rl-env

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes
App README

πŸš‘ DisasterResponseEnv: OpenEnv Hackathon Submission

The environment where the simulation IS the dataset. Learn to save lives when every second counts.

Hackathon Theme: #3.1 Professional Tasks & Long-Horizon Planning Goal: Train an LLM agent to efficiently allocate limited emergency resources (ambulances, rescue teams) across multiple disaster-stricken city zones under partial observability and tight deadlines.

🌟 The Problem

During a crisis (like an earthquake or flood), real emergency operations centers face a massive coordination challenge:

  • β€”Information Overload: Continuous updates on casualties and hazards across multiple zones.
  • β€”Resource Scarcity: Never enough ambulances or rescue teams to cover everyone immediately.
  • β€”Cascading Failures: If you rescue everyone but the hospitals overflow, people still die. If you only focus on the worst-hit zone, other zones collapse.

DisasterResponseEnv is a deterministic, OpenEnv-compliant simulation that forces an LLM agent to learn these hard truths through pure Reinforcement Learning (GRPO). There is no human-labeled datasetβ€”the environment generates the ground truth.

πŸ—οΈ Environment Architecture

State (Partial Observability)

The agent sees a JSON representation of:

  • β€”Available resources (Ambulances, Rescue Teams).
  • β€”Hospital capacity.
  • β€”Zone statuses (Critical patients, Hazard levels, Consecutive turns without support).

Action Space

The agent outputs a structured JSON array deciding how many resources go where:

json
{
  "allocations": [
    {"zone": "Central", "ambulances": 2, "rescue_teams": 1},
    {"zone": "North", "ambulances": 1, "rescue_teams": 0}
  ]
}

The 5 Independent Reward Signals

To prevent reward hacking (e.g., an agent just sending all resources to one zone to get a "good enough" score), the environment calculates 5 competing rewards:

  1. 1.R1 (Lives Saved): +1.5 for treated, -3.0 for deaths.
  2. 2.R2 (Resource Efficiency): -0.5 per wasted ambulance sent to a zone with no patients.
  3. 3.R3 (Coverage / Anti-Neglect): -2.0 penalty if any zone with patients is ignored for 3+ consecutive turns.
  4. 4.R4 (Response Time): Rewards rapid reduction of hazard levels.
  5. 5.R5 (Hospital Management): +1.0 for optimal utilization (50-85%), -2.0 for overcrowding.

πŸš€ Results & Evidence of Learning

We generated synthetic SFT data using a greedy oracle heuristic, and then trained a small LLM (Qwen2.5-1.5B-Instruct via Unsloth) using TRL's GRPOTrainer directly against the environment's reward function.

Before vs After Training (Curriculum Level 2)

MetricUntrained (Random)Trained (GRPO / Oracle Policy)Improvement
Total Reward-181.98-8.98+173.0 pts
Total Deaths4231-26.2%
Patients Treated6688+33.3%

(Run `python demo/run_demo.py` to reproduce baseline metrics locally).

πŸ’» Running the Code

1. Try the Interactive Demo (HuggingFace Spaces format)

bash
pip install -r requirements.txt
python app.py

This opens a Gradio app where you can step through the simulation manually or test JSON actions.

2. Generate Warmup Data

bash
python training/generate_sft_data.py

3. Run the Training Script (Colab / GPU required)

bash
python training/train_grpo.py

Note: Uses `unsloth` and `trl`. Ensure you are in a Linux/CUDA environment.

πŸ“ˆ Why This Matters

This environment goes beyond text generation. Every action mutates the state deterministically. Sending an ambulance to the North zone means the South zone waits longer. Overloading the hospital causes a bottleneck. By training an LLM on this environment, we are not teaching it to "sound like a dispatcher"; we are teaching it to make the mathematically optimal decisions of a dispatcher.

abdulkalamazad07/disaster-response-rl-env Β· Team Ai