RuDubnium/email-triage-openenv
Email Triage & Response β OpenEnv Environment A real-world OpenEnv environment where AI agents learn to triage corporate email inboxes: categorize, prioritize, reply, forward, and flag emails. π Why Email Triage? Email triage is a task performed by billions of knowledge workers daily. It requires: Reading comprehension β understanding intent and urgency Decision-making β choosing correct actions from a discrete set Context reasoning β considering senderβ¦ See the full description on the dataset page: https://huggingface.co/datasets/RuDubnium/email-triage-openenv.
Email Triage & Response β OpenEnv Environment
A real-world OpenEnv environment where AI agents learn to triage corporate email inboxes: categorize, prioritize, reply, forward, and flag emails.
π Why Email Triage?
Email triage is a task performed by billions of knowledge workers daily. It requires:
- Reading comprehension β understanding intent and urgency
- Decision-making β choosing correct actions from a discrete set
- Context reasoning β considering sender, deadlines, dependencies
- Multi-step planning β processing an inbox of emails in priority order
This environment fills a gap in OpenEnv: no existing environment models text-based decision-making workflows.
π Environment Overview
π― Tasks
Task 1: email_categorization (Easy)
- Emails: 5
- Objective: Categorize each email into the correct category
- Categories:
urgent_business,meeting_request,newsletter,spam,customer_complaint,internal_update - Grading: % of emails correctly categorized (0.0β1.0)
- Max Steps: 15
Task 2: priority_triage (Medium)
- Emails: 10
- Objective: Categorize + assign priority (low/medium/high/urgent) + reply to emails needing response
- Grading: Weighted β 40% category + 30% priority + 30% reply quality
- Max Steps: 35
Task 3: full_inbox_management (Hard)
- Emails: 20
- Objective: Full triage β categorize, prioritize, reply, forward complaints to support, flag deadlines, archive spam
- Grading: 25% category + 25% priority + 20% reply + 15% forward + 15% flag
- Max Steps: 80
Task 4: strategic_inbox_cleanup (Pro)
- Emails: 50
- Objective: Long-running strategic management β handle a massive inbox with consistent decision-making
- Grading: Balanced 20% across all 5 action categories
- Max Steps: 200
π€ Complex Trajectories & Strategic Routing
This environment is designed to test an agent's ability to handle long-running tasks with multiple trajectories.
- Multiple Routes: Agents are not forced into a single path. They can choose to:
- Priority-First: Triage urgent emails immediately, then handle the rest.
- Batch-Processing: Categorize all emails first, then reply to all, then archive all.
- Linear: Process the inbox item-by-item from top to bottom.
- Trajectory Length: With up to 200 steps and 50 emails, agents must maintain state and consistency over long sequences of actions.
- Interdependent Actions: Most emails require multiple actions (e.g.,
categorize->reply->flag->archive), creating deep decision trees.
π‘ Action Space
{
"email_id": "string (required) β ID of the email to act on",
"action_type": "enum: categorize | reply | forward | archive | flag | skip",
"category": "enum: urgent_business | meeting_request | newsletter | spam | customer_complaint | internal_update",
"priority": "enum: low | medium | high | urgent",
"reply_text": "string β reply content (for 'reply' action)",
"forward_to": "string β email address (for 'forward' action)"
}π Observation Space
{
"emails": [{"id", "from_addr", "to_addr", "subject", "body", "timestamp"}],
"inbox_size": "int β unprocessed emails remaining",
"processed_count": "int β emails processed so far",
"current_step": "int",
"max_steps": "int",
"done": "bool",
"reward": "float β reward for last action",
"cumulative_reward": "float β total reward",
"feedback": "string β human-readable feedback",
"task_id": "string",
"task_description": "string"
}π Setup & Usage
Local Development
# Install dependencies
cd email_triage_env
pip install -r server/requirements.txt
# Start the server
uvicorn server.app:app --host 0.0.0.0 --port 8000
# Test health
curl http://localhost:8000/health
# Reset with a task
curl -X POST http://localhost:8000/reset \
-H "Content-Type: application/json" \
-d '{"task_id": "email_categorization", "seed": 42}'
# Take an action
curl -X POST http://localhost:8000/step \
-H "Content-Type: application/json" \
-d '{"email_id": "email_1", "action_type": "categorize", "category": "spam", "priority": "low"}'
# Get state
curl http://localhost:8000/state
# List tasks
curl http://localhost:8000/tasks
# Get grader score (after episode completes)
curl -X POST http://localhost:8000/graderDocker
cd email_triage_env
docker build -f server/Dockerfile -t email-triage-env .
docker run -p 8000:8000 email-triage-envInference
# Set required environment variables
export API_BASE_URL="http://localhost:8000"
export MODEL_NAME="gpt-4o-mini"
export HF_TOKEN="your_huggingface_token"
export OPENAI_API_KEY="sk-..."
# Run inference
python inference.pyThe inference script uses structured logging (START, STEP, END) and supports custom LLM backends via environment variables.
π Baseline Scores
Scores are deterministic with seed=42. LLM-based agents score higher.
π Reward Design
Per-step rewards provide signal throughout the trajectory:
- β Correct category: +0.20
- β Correct priority: +0.15
- β Good reply: +0.15
- β Correct forward: +0.15
- β Correct flag: +0.10
- β Archive spam/newsletters: +0.10
- β‘ Close priority (off by 1): +0.05
- β Wrong category: -0.10
- β Missed urgent email: -0.20
- β Archived urgent/complaint: -0.15 to -0.20
π Project Structure
email_triage_env/
βββ __init__.py # Package exports
βββ models.py # Pydantic Action/Observation/State
βββ email_data.py # Deterministic email generator
βββ tasks.py # Task definitions & graders
βββ inference.py # Main inference script (replaces baseline.py)
βββ openenv.yaml # OpenEnv manifest
βββ pyproject.toml # Package config
βββ README.md # This file
βββ server/
βββ __init__.py
βββ app.py # FastAPI application
βββ email_triage_environment.py # Core Environment class
βββ Dockerfile # Container definition
βββ requirements.txt # Python dependenciesπ API Endpoints
License
MIT
