Team Ai
Datasetpublic

RuDubnium/email-triage-openenv

Email Triage & Response β€” OpenEnv Environment A real-world OpenEnv environment where AI agents learn to triage corporate email inboxes: categorize, prioritize, reply, forward, and flag emails. 🌟 Why Email Triage? Email triage is a task performed by billions of knowledge workers daily. It requires: Reading comprehension β€” understanding intent and urgency Decision-making β€” choosing correct actions from a discrete set Context reasoning β€” considering sender… See the full description on the dataset page: https://huggingface.co/datasets/RuDubnium/email-triage-openenv.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes11downloads
Dataset Card

Email Triage & Response β€” OpenEnv Environment

A real-world OpenEnv environment where AI agents learn to triage corporate email inboxes: categorize, prioritize, reply, forward, and flag emails.

🌟 Why Email Triage?

Email triage is a task performed by billions of knowledge workers daily. It requires:

  • β€”Reading comprehension β€” understanding intent and urgency
  • β€”Decision-making β€” choosing correct actions from a discrete set
  • β€”Context reasoning β€” considering sender, deadlines, dependencies
  • β€”Multi-step planning β€” processing an inbox of emails in priority order

This environment fills a gap in OpenEnv: no existing environment models text-based decision-making workflows.


πŸ“‹ Environment Overview

PropertyValue
DomainCorporate Email Inbox Management
SpecOpenEnv v1 (step/reset/state)
Tasks3 (easy β†’ medium β†’ hard)
Action Typescategorize, reply, forward, archive, flag, skip
RewardPer-step partial rewards (not just end-of-episode)

🎯 Tasks

Task 1: email_categorization (Easy)

  • β€”Emails: 5
  • β€”Objective: Categorize each email into the correct category
  • β€”Categories: urgent_business, meeting_request, newsletter, spam, customer_complaint, internal_update
  • β€”Grading: % of emails correctly categorized (0.0–1.0)
  • β€”Max Steps: 15

Task 2: priority_triage (Medium)

  • β€”Emails: 10
  • β€”Objective: Categorize + assign priority (low/medium/high/urgent) + reply to emails needing response
  • β€”Grading: Weighted β€” 40% category + 30% priority + 30% reply quality
  • β€”Max Steps: 35

Task 3: full_inbox_management (Hard)

  • β€”Emails: 20
  • β€”Objective: Full triage β€” categorize, prioritize, reply, forward complaints to support, flag deadlines, archive spam
  • β€”Grading: 25% category + 25% priority + 20% reply + 15% forward + 15% flag
  • β€”Max Steps: 80

Task 4: strategic_inbox_cleanup (Pro)

  • β€”Emails: 50
  • β€”Objective: Long-running strategic management β€” handle a massive inbox with consistent decision-making
  • β€”Grading: Balanced 20% across all 5 action categories
  • β€”Max Steps: 200

πŸ›€ Complex Trajectories & Strategic Routing

This environment is designed to test an agent's ability to handle long-running tasks with multiple trajectories.

  • β€”Multiple Routes: Agents are not forced into a single path. They can choose to:
  • β€”Priority-First: Triage urgent emails immediately, then handle the rest.
  • β€”Batch-Processing: Categorize all emails first, then reply to all, then archive all.
  • β€”Linear: Process the inbox item-by-item from top to bottom.
  • β€”Trajectory Length: With up to 200 steps and 50 emails, agents must maintain state and consistency over long sequences of actions.
  • β€”Interdependent Actions: Most emails require multiple actions (e.g., categorize -> reply -> flag -> archive), creating deep decision trees.

πŸ“‘ Action Space

json
{
  "email_id": "string (required) β€” ID of the email to act on",
  "action_type": "enum: categorize | reply | forward | archive | flag | skip",
  "category": "enum: urgent_business | meeting_request | newsletter | spam | customer_complaint | internal_update",
  "priority": "enum: low | medium | high | urgent",
  "reply_text": "string β€” reply content (for 'reply' action)",
  "forward_to": "string β€” email address (for 'forward' action)"
}

πŸ‘ Observation Space

json
{
  "emails": [{"id", "from_addr", "to_addr", "subject", "body", "timestamp"}],
  "inbox_size": "int β€” unprocessed emails remaining",
  "processed_count": "int β€” emails processed so far",
  "current_step": "int",
  "max_steps": "int",
  "done": "bool",
  "reward": "float β€” reward for last action",
  "cumulative_reward": "float β€” total reward",
  "feedback": "string β€” human-readable feedback",
  "task_id": "string",
  "task_description": "string"
}

πŸ— Setup & Usage

Local Development

bash
# Install dependencies
cd email_triage_env
pip install -r server/requirements.txt

# Start the server
uvicorn server.app:app --host 0.0.0.0 --port 8000

# Test health
curl http://localhost:8000/health

# Reset with a task
curl -X POST http://localhost:8000/reset \
  -H "Content-Type: application/json" \
  -d '{"task_id": "email_categorization", "seed": 42}'

# Take an action
curl -X POST http://localhost:8000/step \
  -H "Content-Type: application/json" \
  -d '{"email_id": "email_1", "action_type": "categorize", "category": "spam", "priority": "low"}'

# Get state
curl http://localhost:8000/state

# List tasks
curl http://localhost:8000/tasks

# Get grader score (after episode completes)
curl -X POST http://localhost:8000/grader

Docker

bash
cd email_triage_env
docker build -f server/Dockerfile -t email-triage-env .
docker run -p 8000:8000 email-triage-env

Inference

bash
# Set required environment variables
export API_BASE_URL="http://localhost:8000"
export MODEL_NAME="gpt-4o-mini"
export HF_TOKEN="your_huggingface_token"
export OPENAI_API_KEY="sk-..."

# Run inference
python inference.py

The inference script uses structured logging (START, STEP, END) and supports custom LLM backends via environment variables.


πŸ“Š Baseline Scores

TaskDifficultyRule-Based ScoreDescription
email_categorizationEasy~0.80–1.00Keyword heuristics work well
priority_triageMedium~0.60–0.80Priority and reply quality harder
full_inbox_managementHard~0.40–0.60Forwarding and flagging add complexity

Scores are deterministic with seed=42. LLM-based agents score higher.


πŸ† Reward Design

Per-step rewards provide signal throughout the trajectory:

  • β€”βœ… Correct category: +0.20
  • β€”βœ… Correct priority: +0.15
  • β€”βœ… Good reply: +0.15
  • β€”βœ… Correct forward: +0.15
  • β€”βœ… Correct flag: +0.10
  • β€”βœ… Archive spam/newsletters: +0.10
  • β€”βš‘ Close priority (off by 1): +0.05
  • β€”βŒ Wrong category: -0.10
  • β€”βŒ Missed urgent email: -0.20
  • β€”βŒ Archived urgent/complaint: -0.15 to -0.20

πŸ“ Project Structure

email_triage_env/
β”œβ”€β”€ __init__.py                 # Package exports
β”œβ”€β”€ models.py                   # Pydantic Action/Observation/State
β”œβ”€β”€ email_data.py               # Deterministic email generator
β”œβ”€β”€ tasks.py                    # Task definitions & graders
β”œβ”€β”€ inference.py                # Main inference script (replaces baseline.py)
β”œβ”€β”€ openenv.yaml                # OpenEnv manifest
β”œβ”€β”€ pyproject.toml              # Package config
β”œβ”€β”€ README.md                   # This file
└── server/
    β”œβ”€β”€ __init__.py
    β”œβ”€β”€ app.py                  # FastAPI application
    β”œβ”€β”€ email_triage_environment.py  # Core Environment class
    β”œβ”€β”€ Dockerfile              # Container definition
    └── requirements.txt        # Python dependencies

πŸ”Œ API Endpoints

EndpointMethodDescription
/healthGETHealth check
/resetPOSTStart new episode
/stepPOSTExecute action
/stateGETGet episode state
/tasksGETList tasks with schema
/graderPOSTGet grader score
/baselinePOSTRun baseline agent

License

MIT