Sanaaashaikh/AI_Food_Safety_Transparency_Environment
AI Food Safety Transparency Environment
An OpenEnv-compliant simulation environment where an AI agent makes decisions about restaurant safety visibility on a food delivery platform — balancing user trust, safety risk, and false positives.
Problem
Consumers using food delivery platforms cannot clearly see verified kitchen hygiene and safety standards. While star ratings exist, verified inspection data is often missing or outdated. This environment simulates a system where an AI agent decides how to expose or act on food safety data to maximize user trust while minimizing health risk.
Environment Design
The environment follows the OpenEnv specification with three core methods:
env.reset() # Initialize / restart the episode
env.step(action) # Apply an action, receive reward + next state
env.state() # Inspect current state without actingState Space
Each restaurant is represented as a RestaurantState dataclass:
Action Space
Reward Logic
The reward is partial (not binary), ranging from -1.0 to +1.0:
Per-Action Rewards
Episode-Level Score (0.0 – 1.0)
final_score = normalize(total_reward)
+ safety_bonus # +0.1 for correctly flagging all risky restaurants
- false_pos_penalty # -0.15 per safe restaurant incorrectly flagged
+ coverage_bonus # +0.05 for requesting inspections on stale dataTasks
Easy Task (easy_task)
- 5 restaurants, 5 steps
- Clear-cut cases: pristine hygiene (8-10) with verified recent inspections vs. severely unsafe kitchens (1-3) with stale data and many complaints
- Optimal agent should score ≥ 0.85
Medium Task (medium_task)
- 7 restaurants, 7 steps
- Mixed signals: decent hygiene with stale inspections, high complaints despite good scores, pending verifications requiring judgment calls
- Optimal agent should score ≥ 0.70
Hard Task (hard_task)
- 10 restaurants, 10 steps
- Adversarial edge cases: near-perfect hygiene scores with 500+ day-old inspections, newly opened restaurants with low scores but fresh checks, high-volume restaurants with hidden complaint spikes
- Optimal agent should score ≥ 0.60
Project Structure
food-safety-env/
├── openenv.yaml # OpenEnv specification
├── inference.py # End-to-end inference script
├── app.py # Gradio Hugging Face Spaces UI
├── Dockerfile # Container definition
├── requirements.txt # Python dependencies
├── README.md # This file
└── env/
├── __init__.py
├── models.py # Pydantic/dataclass typed models
├── environment.py # OpenEnv FoodSafetyEnv class
├── tasks.py # Easy, medium, hard task definitions
└── grader.py # Scoring logic (0.0 – 1.0)Setup and Run
Local (Python)
# 1. Clone / extract project
cd food-safety-env
# 2. Install dependencies
pip install -r requirements.txt
# 3. Set environment variables
export API_BASE_URL=https://api.openai.com/v1
export MODEL_NAME=gpt-4o-mini
export HF_TOKEN=your_openai_or_hf_token
# 4. Run inference
python inference.pyWithout an API key — the script automatically falls back to a deterministic rule-based agent that still demonstrates all environment mechanics.
Docker
# Build
docker build -t food-safety-env .
# Run inference
docker run -e HF_TOKEN=your_token -e MODEL_NAME=gpt-4o-mini food-safety-env
# Run Gradio UI
docker run -p 7860:7860 -e HF_TOKEN=your_token food-safety-env python app.pyGradio UI (Hugging Face Spaces)
python app.py
# Open http://localhost:7860- Select difficulty (easy / medium / hard)
- Press Reset Environment
- Choose an action from the dropdown
- Press Take Action
- Watch scores update in real time
Environment Variables
Expected Output
============================================================
AI FOOD SAFETY TRANSPARENCY ENVIRONMENT — INFERENCE
============================================================
Model: gpt-4o-mini
Seed: 42
TASK: EASY_TASK | Difficulty: easy
Step 1: Golden Spoon
Action: show_safety_badge
Reward: +1.0000 — Correctly showing badge for a safe, verified restaurant.
TASK SCORE: 0.9250
FINAL SCORES SUMMARY
============================================================
easy_task [easy ] Score: 0.9250 ██████████████████
medium_task [medium] Score: 0.7340 ██████████████
hard_task [hard ] Score: 0.6120 ████████████
Overall Average Score: 0.7570