SHUBHAMOS/meta-pytorch-hackathon
0
1# ๐จโโ๏ธ SHUBHAMOS: Evaluation & Verification Guide2 3This guide describes how judges and users can verify that the SHUBHAMOS Email Triage Agent is working correctly according to the **OpenEnv** protocol.4 5---6 7## 1. Interactive Benchmark Dashboard (Gradio)8The easiest way to verify the agent is via the **Hugging Face Space Index**.91. **Open the Space**: Wait for the "Running" status.102. **Configuration**: 11 * Enter your **Hugging Face API Token** (needed to call the Qwen router).12 * Select a **Simulation Task** (Easy, Medium, or Hard).133. **Run**: Click **Run Agent Benchmark ๐**.144. **Verification**: 15 * You will see a live summary of the agent's actions.16 * A final **Grade Report** will appear, showing classification accuracy and total rewards.17 18---19 20## 2. Startup Health Logs21Hugging Face Space logs provide an automated audit trail of the system's integrity on every boot:22* **Token Validation**: The container logs will show `HF_TOKEN loaded successfully โ
`.23* **Inference Test**: It attempts a "Health Check" prompt to verify the LLM router connectivity.24* **Automated E2E Logs**: Before the server starts, it runs a pre-flight benchmark. Look for the `===== SHUBHAMOS E2E LOG =====` section in the console logs.25 26---27 28## 3. Programmatic API Access (OpenEnv)29The agent follows a strict OpenAPI/FastAPI contract. You can test the endpoints manually via the built-in Swagger UI at `/docs`.30 31### Test Flow:321. **Reset**: `POST /reset?task_id=easy` -> Returns initial inbox state.332. **Step**: `POST /step` with an action body like:34 ```json35 {36 "action_type": "classify_email",37 "email_id": "email_1",38 "category": "billing_issue"39 }40 ```413. **Grade**: `GET /grade` -> Returns real-time metrics for the active session.42 43---44 45## 4. Technical Compliance (OpenEnv Requirements)46- **Environment Spec**: View the full OpenEnv definition at `/openenv.yaml`.47- **State Transparency**: All internal ground-truth is available at `/state` (restricted to graders in production, but open here for hackathon transparency).48- **Graceful Failover**: The agent includes a three-tier inference loop (Primary -> Internal Fallback -> Smart Fallback logic) to ensure benchmarks always complete even during API outages.49 50---51 52**Project Repositories:**53- **Hugging Face:** [SHUBHAMOS Space](https://huggingface.co/spaces/SHUBHAMOS/meta-pytorch-hackathon)54- **GitHub:** [shubhamos-ai/meta-openenv](https://github.com/shubhamos-ai/meta-openenv)55 