vighnesh2190/meta-pytorch-openenv-hackathon-v2
Delivery Worker Assignment Environment
Production-grade OpenEnv-style logistics environment for assigning delivery orders to workers under distance, deadline, and capacity constraints.
Problem
An agent must manage a fleet of delivery workers over discrete timesteps. Each task contains a fixed set of orders and workers. Orders have pickup and dropoff coordinates plus delivery deadlines. Some orders are marked high priority, which makes on-time service more valuable and lateness more costly. Workers have locations, capacity limits, and assignment queues. The agent chooses from three validated actions:
assign_order(order_id, worker_id)reassign_order(order_id, worker_id)advance_time()
The environment is deterministic. The same task and action sequence always produce the same state trajectory, rewards, and grader score.
OpenEnv Requirements Covered
step(action) -> observation, reward, done, inforeset() -> observationstate() -> internal state- Typed Pydantic models for
Action,Observation,Reward,State, and task metadata - Root
openenv.yaml - FastAPI service endpoints for environment interaction and evaluation
- Dockerized execution
- Hugging Face Spaces compatibility through
app.pyand Docker
State Design
State tracks:
ordersidpickup_locationdrop_locationdeadlineprioritystatus- assignment and delivery timestamps
workersidcurrent_locationcapacityassigned_ordersstatustime- aggregate delivery metrics
Workers service assigned orders in queue order. Each advance_time step moves a worker by one Manhattan unit toward its current pickup or dropoff target.
Observation Schema
Observations expose:
- all orders with status and assignment metadata
- all worker states
- current timestep and max horizon
- summary counts for pending, assigned, delivered, and remaining time
The schema is available from /schema.
Reward Logic
The reward is dense and additive:
+1.0for each order delivered on time- up to
+0.3for assigning an order to a near-optimal worker +0.1each time a worker moves closer to its current target+0.2extra for on-time high-priority deliveries-0.5when an order misses its deadline-0.2extra when a high-priority order misses its deadline-0.2for invalid actions-0.1per idle worker when active orders remain- fairness penalty when active queue load is too uneven
- reassignment penalty to discourage churn
Each step returns a scalar reward for compatibility, while info.reward_breakdown preserves the full DeliveryReward component model.
Tasks
Easy
- 4 orders
- 3 workers
- 1 high-priority order
- mildly tighter deadlines
- objective: assign cleanly without overloading one worker
Medium
- 5 orders
- 2 workers
- 2 high-priority orders
- distance and sequencing matter
- objective: avoid wasting travel while keeping moderate deadlines
Hard
- 8 orders
- 3 workers with uneven queue capacity limits
- 3 high-priority rush orders
- tighter deadlines and harsher reassignment tradeoffs
- objective: plan across multiple timesteps and use reassignment only when it is worth the cost
List all tasks and the action schema at runtime with:
curl http://localhost:8000/tasksDeterministic Grader
The grader replays a submitted action sequence against the fixed task instance and computes:
- on-time delivery rate
- completion rate
- delivery efficiency score based on distance traveled versus a deterministic lower bound
- worker fairness score based on delivery load balance
- high-priority service rate
- late or missed order ratio
- invalid action ratio
Final score:
score = clamp(
0.60 * on_time_delivery_rate
+ 0.10 * efficiency_score
+ 0.05 * fairness_score
+ 0.10 * priority_service_rate
- 0.10 * late_delivery_ratio
- 0.05 * invalid_action_ratio,
0.0,
1.0
)The same task_id and actions payload always produce the same result.
API Endpoints
POST /resetPOST /stepGET /stateGET /tasksPOST /graderGET /baselineGET /schemaGET /health
Reset Example
curl -X POST http://localhost:8000/reset \
-H "Content-Type: application/json" \
-d '{"task_id":"easy"}'Step Example
curl -X POST http://localhost:8000/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"action_type": "assign_order",
"order_id": "O-101",
"worker_id": "W-1"
}
}'The response shape is:
{
"observation": { "...": "..." },
"reward": 0.2889,
"done": false,
"info": {
"reward_breakdown": {
"total": 0.2889
}
}
}Grader Example
curl -X POST http://localhost:8000/grader \
-H "Content-Type: application/json" \
-d '{
"task_id": "easy",
"actions": [
{"action_type":"assign_order","order_id":"O-101","worker_id":"W-1"},
{"action_type":"advance_time"}
]
}'The grader response includes a top-level scalar score and the detailed deterministic result payload:
{
"score": 0.78,
"result": {
"score": 0.78
}
}Project Structure
api/
baseline/
env/
grader/
models/
tasks/
tests/
app.py
Dockerfile
openenv.yaml
pyproject.toml
README.md
requirements.txtSetup
Local Python
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn api.app:app --host 0.0.0.0 --port 8000Docker
docker build -t delivery-worker-assignment-env .
docker run --rm -p 8000:8000 delivery-worker-assignment-envOpenAI Baseline
The baseline agent lives in baseline/run_baseline.py. It:
- fetches task metadata from
/tasks - resets each task
- asks the OpenAI Responses API for one structured action at a time
- falls back to Gemini if OpenAI is unavailable or out of quota and
GEMINI_API_KEYis configured - falls back to a deterministic heuristic if neither model path is available
- submits the final action trace to
/grader - reports
provider_used,providers_attempted, andproviders_succeededper task so model fallback behavior is explicit
Environment variables:
OPENAI_API_KEYpreferred primary providerGEMINI_API_KEYoptional fallback providerOPENAI_MODELoptional, defaults togpt-5.4-miniGEMINI_MODELoptional, defaults togemini-2.5-flashBASE_URLoptional, defaults tohttp://localhost:8000
Run it with:
python -m baseline.run_baseline --base-url http://localhost:8000Hugging Face Spaces
This repository is compatible with Hugging Face Docker Spaces:
- root
Dockerfileexposes port8000 - root
app.pyexportsapp openenv.yamlpoints toapi.app:app
For a Docker Space, push the repository as-is and set the Space SDK to Docker.
Notes
- Rewards are deterministic and never sampled randomly.
- Tasks are fixed and reproducible.
- Invalid actions are always penalized.
- The environment is stateful across steps until
/resetis called.
