Team Ai
Apppublic

SHUBHAMOS/meta-pytorch-hackathon

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes
README.md237 linesDownload Raw Back to root
1---2title: SHUBHAMOS - AI Email Triage Environment3emoji: ๐Ÿ“ˆ4colorFrom: blue5colorTo: gray6sdk: docker7pinned: false8license: mit9short_description: OpenEnv AI email triage environment for agent learning10---11 12# SHUBHAMOS: AI Email Operations & Triage Environment13 14[![OpenEnv Compliant](https://img.shields.io/badge/OpenEnv-1.0.0-blue.svg)](https://github.com/openenv/spec)15[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)16 17**SHUBHAMOS** is a production-grade [OpenEnv](https://github.com/openenv/spec)-compliant environment (developed by **Meta Research**) designed for the **Meta PyTorch Hackathon**. It evaluates and trains AI agents on real-world email triage tasks, providing high-density reward signals and deterministic grading for operational AI agents.18 19---20 21## ๐Ÿ–ผ๏ธ Architecture Overview22 23```mermaid24graph TD25    A[Gradio Dashboard] -->|User Trigger| B[Inference Engine]26    B -->|Action| C[EmailTriageEnv]27    C -->|Observation| B28    B -->|Step Logic| D{AI Router}29    D -->|Tier 1| E[Qwen-72B Primary]30    D -->|Tier 2| F[NVIDIA Fallback]31    D -->|Tier 3| G[Smart Heuristics]32    C -->|State| H[Deterministic Grader]33    H -->|Final Score| A34```35 36---37 38## ๐Ÿ› ๏ธ Key Systems39 40> [!IMPORTANT]41> **SHUBHAMOS** implements a rigid three-tier inference safety loop to ensure zero-downtime benchmarking even during API outages.42 43| Feature | Description |44| :--- | :--- |45| **OpenEnv Core** | Full compliance with the OpenEnv v1.0 specification. |46| **Dense Rewards** | Shaped reward signals optimized for RL training and tuning. |47| **Gradio GUI** | Instant interactive benchmarking without writing code. |48| **Multi-AI Switch** | Automatic failover between primary cloud and secret internal AI. |49 50---51 52## ๐Ÿง  Environment Description53 54SHUBHAMOS simulates a dynamic email inbox environment. The environment lifecycle follows the standard RL pattern:55 561.  **Reset:** The environment initializes with a task-specific seed, creating a predictable set of emails with varying categories and priorities.572.  **Step:** The agent receives an **Observation** (inbox state) and submits an **Action**. The environment returns a **Reward**, updated state, and a **Done** flag.583.  **Episode:** The agent continues taking steps until all emails are resolved, escalated, or ignored, or until the `max_steps` limit is reached.59 60---61 62## ๐Ÿš€ Evaluation Center (For Judges)63 64The agent can be verified through three distinct layers of observability:65 66### ๐ŸŽฎ Interactive Verification (UI)671.  Navigate to your **Hugging Face Space**.682.  Paste your **Hugging Face Token** in the dashboard.693.  Choose a **Simulation Task** (Easy / Medium / Hard).704.  Click **Run Agent Benchmark ๐Ÿš€**.715.  Watch the agent process the inbox in real-time and view the final **Grade Report**.72 73### ๐Ÿ“ Automated Startup Audit (Logs)74Every container boot runs a **Pre-flight Health Check**. Inspect the **Hugging Face Container Logs** for:75- [x] **Secret Loading**: `HF_TOKEN loaded successfully โœ…`76- [x] **Inference Test**: `Health Check: AI Provider is HEALTHY โœ…`77- [x] **E2E Validation**: Detailed benchmark metrics for `easy` and `medium` tasks.78 79### ๐Ÿ”Œ Programmatic API (OpenEnv)80The environment exposes standard OpenEnv endpoints via FastAPI. Access the **interactive documentation** at `/docs`.81 82---83 84## โš™๏ธ OpenEnv Specification85 86### ๐Ÿ“ฅ Observation Model87The agent perceives the inbox through a structured `Observation` object:88*   **`emails`**: A list of `ObservationEmail` objects (ID, Subject, Sender, Sentiment, Body Preview).89*   **`pending_count`**: Number of emails requiring action.90*   **`resolved_count`**: Number of emails successfully closed.91*   **`elapsed_ratio`**: A float (0.0 to 1.0) indicating progress toward the step limit.92 93### ๐Ÿ“ค Action Model94The agent interacts using structured discrete actions:95*   **`classify_email(email_id, category)`**: Assigns a label (e.g., `billing_issue`).96*   **`set_priority(email_id, level)`**: Assigns priority (`low`, `medium`, `high`).97*   **`mark_resolved(email_id)`**: Attempts to close the email.98*   **`escalate_email(email_id)`**: Hands off to a human operator.99*   **`ignore_email(email_id)`**: Skips the email (penalized for urgent items).100 101---102 103## ๐ŸŽฏ Tasks104 105| Task | ID | Emails | Steps | Difficulty |106| :--- | :--- | :--- | :--- | :--- |107| **Basic Inbox** | `easy` | 8 | 50 | Low complexity, clear categories. |108| **Support Queue** | `medium` | 20 | 120 | Mixed priorities, requires sequencing. |109| **Enterprise Inbox** | `hard` | 50 | 300 | High ambiguity, spam noise, scale pressure. |110 111---112 113## ๐Ÿงช Grading System (0.0 โ€“ 1.0)114 115SHUBHAMOS uses a **Deterministic Grader** to ensure reproducible scores across different agent architectures. The final score is calculated based on:1161.  **Classification Accuracy (30%):** Matching agent-assigned categories to ground truth.1172.  **Priority Accuracy (30%):** Correctly identifying urgency levels.1183.  **Resolution Rate (30%):** Successfully closing actionable emails.1194.  **Urgent Handling (10%):** Specific bonus for resolving high-priority items early.120 121---122 123## ๐Ÿ† Reward Function124 125The environment provides **dense, shaped rewards** to facilitate reinforcement learning:126*   **+0.3**: Correct classification of an email.127*   **+0.3**: Correct priority assignment.128*   **+0.5**: Successful resolution of an email.129*   **+0.4**: Bonus for resolving urgent emails before the halfway point of `max_steps`.130*   **-0.2**: Penalty for incorrect labels.131*   **-0.5**: Heavy penalty for ignoring or incorrectly handling urgent complaints.132 133---134 135## ๐Ÿค– Baseline Agent (`inference.py`)136 137A reference implementation is provided in `inference.py`. It features:138*   **Multi-Provider Fallback:** Uses a Primary AI (Hugging Face) with a secret Internal Secondary AI fallback for high reliability.139*   **Stateful Memory:** Tracks which emails have already been handled to avoid redundant actions.140*   **Smart Fallback:** A keyword-based heuristic engine triggers if the LLM fails repeatedly or hits rate limits.141*   **Deterministic Logic:** Uses `temperature=0.05` and strict JSON outputs.142 143---144 145## ๐Ÿ› ๏ธ Setup Instructions146 147### 1. Clone & Install148```bash149git clone https://github.com/shubhamos-ai/meta-openenv150cd meta-openenv151pip install -r requirements.txt152```153 154### 2. Configure Secrets155Create a `.env` file in the root directory:156```bash157HF_TOKEN=hf_...158INTERNAL_AI_KEY=nvapi-...159```160> [!TIP]161> On Hugging Face, add these as **Secrets** in the Space settings to ensure high-performance inference fallback is active.162 163---164 165## โ–ถ๏ธ Running the Environment166 167### Run Baseline Agent168```bash169# Run Easy task (default)170python3 inference.py --task easy171 172# Run Medium task173python3 inference.py --task medium --verbose174 175# Run All tasks and save results176python3 inference.py --task all --output results.json177```178 179### Run E2E Validation180SHUBHAMOS includes a validation suite to verify environment health and agent performance:181```bash182chmod +x tests/e2e_runner.sh183./tests/e2e_runner.sh184```185 186---187 188## ๐Ÿณ Docker & Deployment189 190### Local Docker Run191```bash192docker build -t shubhamos .193docker run -p 7860:7860 --env-file .env shubhamos194```195 196### Hugging Face Spaces1971.  Create a new **Docker** Space on Hugging Face.1982.  Select the **Empty** template or push this repository directly.1993.  Add `HF_TOKEN` and `INTERNAL_AI_KEY` under **Settings > Variables and Secrets**.2004.  Deployment will automatically start using the included `Dockerfile`.201 202---203 204## ๐Ÿ“ Project Structure205 206```text207โ”œโ”€โ”€ environment.py      # Core OpenEnv logic208โ”œโ”€โ”€ models.py           # Pydantic state/action models209โ”œโ”€โ”€ reward.py           # Reward signal engine210โ”œโ”€โ”€ tasks/              # Task configurations (Easy/Medium/Hard)211โ”œโ”€โ”€ graders/            # Deterministic scoring logic212โ”œโ”€โ”€ inference.py        # Baseline LLM Agent213โ”œโ”€โ”€ openenv.yaml        # OpenEnv Spec Manifest214โ”œโ”€โ”€ .env                # Private secrets (git-ignored)215โ””โ”€โ”€ tests/              # E2E validation scripts216```217 218---219 220## ๐Ÿ’ก Design Decisions221*   **Determinism:** Fixed seeds and explicit graders ensure that agent improvements are real, not noise.222*   **JSON-First:** Agent communication is forced into pure JSON to minimize parsing variability.223*   **Observability:** Every step is logged with associated rewards and state transitions for easy debugging.224 225---226 227## ๐Ÿš€ Training Purpose228SHUBHAMOS is designed for **Reward-Driven Optimization**. The dense rewards allow for:229*   **Proximal Policy Optimization (PPO)** for triage strategy.230*   **Few-Shot Tuning** of LLM system prompts.231*   **Fine-Tuning** of small models (e.g., Llama-3-8B) on the generated triage trajectories.232 233---234 235## ๐Ÿ“œ License236This project is licensed under the MIT License - see the LICENSE file for details.237