mohitkourav/DisasterResponseCoordinatorEnv
DisasterResponseCoordinatorEnv
Autonomous Emergency Management Swarm β 8 AI Agents Coordinating Disaster Response
  ![OpenEnv]()
When every second counts, AI must coordinate. This environment trains LLMs to save lives.
Agent Decision Loop
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β COORDINATOR LLM β
β (The model being trained) β
ββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββ
β Observes
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β WORLD STATE β
β β’ Crisis map (zones, roads, hospitals) β
β β’ 8 Agent reports (conflicts, status) β
β β’ Resources (fuel, trucks, medicine) β
β β’ Hour 0β72, Phase: RescueβReliefβRehab β
ββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββ
β Chooses 1 of 8 tools
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β ACTIONS β
β dispatch_team β allocate_resource β re_route β
β request_airlift β order_evacuation β deploy_scout β
β setup_comms β advance_hour β
ββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββ
β Gets reward signal
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β 12-SIGNAL REWARD β
β +0.12 critical patient treated β
β +0.09 person rescued from danger zone β
β -0.12 death before rescue reached β
β -0.08 team dispatched to blocked zone β
ββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββ
β Learns via GRPO
βΌ
Better next decisionThe Problem
India experiences devastating losses during natural disasters due to fragmented coordination in the critical first 72 hours. While AI helps predict weather, no standardized RL environment exists to train AI agents for real-time, dynamic resource allocation under chaos.
The Environment
DisasterResponseCoordinatorEnv is a graph-based sandbox where 8 AI agents coordinate rescue operations across a dynamic crisis zone:
4 India-Specific Tasks
8 MCP Tools (Action Space)
dispatch_team Β· allocate_resource Β· re_route Β· request_airlift Β· order_evacuation Β· deploy_scout Β· setup_comms Β· advance_hour
Hackathon Themes Covered
Theme 1 β Multi-Agent: 8 agents with competing interests (Medical vs Air Support on fuel, Supply vs Medical on conservation). Conflicts visible on dashboard.
Theme 2 β Long-Horizon: 72-hour simulation with 3 phases (RescueβReliefβRehabilitation). Sparse delayed rewards. Early decisions affect late outcomes.
Theme 3 β World Modeling: NetworkX graph with 15-20 nodes. 8 dynamic event types (aftershock, hospital overflow, comms breakdown, new survivors). 6 world state layers.
Theme 4 β Self-Improvement: Adaptive curriculum (auto-generates harder scenarios from failure analysis). Strategy memory (stores learned heuristics). Self-adaptive reward shaping.
Results
Why the regression? The village_flood_rescue task (50 people, 50 steps) is trivially solvable β a random agent saturates the grader at 0.919, leaving no headroom for a 0.5B model to show improvement via GRPO. This is a known limitation of easy-task baselines in RL, not a failure of the environment design. Training on multi_district_cyclone (Medium difficulty) is recommended for future runs to produce positive deltas.
Average episode reward over training.
Random baseline vs GRPO trained agent score comparison.
GRPO training loss over 25 steps β model converges to 0.0766.
How to Run
Try the Live Dashboard
Visit: https://huggingface.co/spaces/mohitkourav/DisasterResponseCoordinatorEnv
Run Locally
git clone https://github.com/MOHITKOURAV01/DisasterResponseCoordinatorEnv.git
cd DisasterResponseCoordinatorEnv
pip install -r requirements.txt
uvicorn server.main:app --port 7860
# Open http://localhost:7860Training (Colab)
Open the Training Notebook and click Run All.
API Quick Reference
Connect to the live environment in 3 lines:
import httpx
obs = httpx.post(
"https://mohitkourav-disasterresponsecoordinatorenv.hf.space/reset",
json={"task_id": "village_flood_rescue"}
).json()["observation"]
# obs now has: zones, hospitals, roads, resources, agent_reportsSee examples/ folder for complete working scripts.
Research Foundations
- Hierarchical MARL for Emergency Responders (ICML 2024)
- ReinforceRouting β Graph-based dynamic routing with RL
- Self-Adaptive Reward Shaping (ICLR 2025)
- Curriculum Learning for RL (JMLR Survey)
- 72-hour PPO Relief Distribution with equity metrics
- Digital Risk Twin for Disaster Management (Nature 2025)
Links
Author
Mohit Kourav β Meta PyTorch OpenEnv Hackathon Γ Scaler School of Technology
Built with FastAPI, NetworkX, Chart.js, and vanilla JS. Zero external cost.
