admesh/agentic-intent-classifier
245
1# Roadmap Issue Drafts2 3These are the next three roadmap issues to open in GitHub once authenticated issue creation is available.4 5## 1. Build External Demo UI For Decision Envelope6 7Suggested title:8 9`Build external demo UI for query -> model_output -> system_decision`10 11Suggested body:12 13```md14## Goal15 16Add a simple external-facing demo interface on top of `/classify` so a user can paste a query and see the full decision envelope in a clean, understandable format.17 18## Scope19 20- add a lightweight UI for entering a raw query21- render `model_output.classification.intent`22- render fallback state when present23- render `system_decision.policy`24- render `system_decision.opportunity`25- include a few preloaded demo prompts26 27## Why28 29The current JSON API is enough for engineering validation, but not enough for partner demos or taxonomy walkthroughs.30 31## Done When32 33- someone can run the demo locally and inspect the full output without using curl34- the UI clearly shows query -> classification -> system decision35```36 37## 2. Add Better Support Handling To Intent-Type Layer38 39Suggested title:40 41`Add dedicated support handling to reduce personal_reflection fallback on account-help prompts`42 43Suggested body:44 45```md46## Goal47 48Reduce the current failure mode where support-like prompts such as login and billing issues collapse into `personal_reflection` or low-confidence fallback behavior.49 50## Scope51 52- review support-like prompts in the current benchmark53- decide whether to add a dedicated `support` intent-type head or a rule-based override layer54- add a fixed support-oriented evaluation set55- document the chosen approach in `known_limitations.md`56 57## Why58 59The `decision_phase` head can already separate `support` reasonably well, but the `intent_type` layer still underperforms on these cases.60 61## Done When62 63- support prompts are no longer commonly labeled as `personal_reflection`64- the combined envelope fails safe for support queries with clearer semantics65```66 67## 3. Add Evaluation Harness And Canonical Benchmark Runner68 69Suggested title:70 71`Add canonical benchmark runner for demo prompts and regression checks`72 73Suggested body:74 75```md76## Goal77 78Turn the current prompt suite and canonical examples into a repeatable regression harness.79 80## Scope81 82- add a script that runs the fixed demo prompts through `combined_inference.py`83- save outputs to a machine-readable artifact84- compare current outputs against expected behavior notes85- flag meaningful regressions in fallback behavior and phase classification86 87## Why88 89The repo now has frozen `v0.1` baselines. A benchmark runner is the clean way to protect demo quality without returning to ad hoc tuning.90 91## Done When92 93- one command runs the prompt suite end to end94- current outputs are easy to inspect and compare over time95- demo regressions become visible before external sharing96```97 