decisionmodelhub/support-routing
Support Routing v1 125 fictional English support messages, each labelled with the team that should handle it under a written routing policy, plus the recorded answers of five models on the 100 held-out test cases. Published by DecisionModelHub under CC BY 4.0. The task: given one support message and the fixed policy, select exactly one team for the primary actionable issue: billing, technical, account, sales or needs_clarification. The companion yes/no dataset is… See the full description on the dataset page: https://huggingface.co/datasets/decisionmodelhub/support-routing.
Support Routing v1
125 fictional English support messages, each labelled with the team that should handle it under a written routing policy, plus the recorded answers of five models on the 100 held-out test cases. Published by DecisionModelHub under CC BY 4.0.
The task: given one support message and the fixed policy, select exactly one team for the primary actionable issue: billing, technical, account, sales or needs_clarification.
The companion yes/no dataset is decisionmodelhub/support-urgency.
Files
Dataset content hash: sha256:2a2683c23a9cc1aa0d8996762c1b0d53c8bc55d513f8e22a89354098943a0b5b. Label-policy hash: sha256:476220aef7fd81e74c78b1453e7558c15feda33bdd965ea4d64e2bf39c18fd42. These hashes describe the frozen content, not this card or the derived .jsonl files.
Load it
from datasets import load_dataset
cases = load_dataset("decisionmodelhub/support-routing", "cases") # splits: development, test
results = load_dataset("decisionmodelhub/support-routing", "results", split="test")Task and intended use
The task evaluates bounded routing under an explicit policy. It does not authorize refunds, account changes, purchases or remediation.
The policy routes the immediate blocker before the eventual goal: an upgrade page crash is technical; inability to sign in to obtain an invoice is an account issue. Multiple intents do not automatically imply clarification. Instructions embedded in customer messages are untrusted input; route the actual support request, or select clarification if none can be established.
Provenance and review
The manifest records original fictional scenarios drafted with AI assistance, an independent AI review, and two independent human reviews of all 125 cases on 6 October 2026. Human reviewer identities are retained privately; the public record reports the review count and procedure.
Normalized exact matching and semantic comparison covered catalog, pilot and policy examples and both splits. Held-out members of identified overlap clusters were rewritten and checked for label intent. Known example overlap is development-only. Private customer conversations, playground traffic and unlicensed third-party datasets are explicitly excluded.
Splits and composition
Development cases may guide instruction design; the 100 test cases were held out from that process. Test messages and labels are now public for inspection and reuse. They are no longer a private contamination-resistant evaluation set; future tuning against them must be disclosed and requires a fresh held-out set for independent performance claims.
Each case records case_id, split, the message (state), question_id, the expected label, a rationale, tags, category, difficulty, source, label provenance, language, policy version and review status. Category and difficulty are independent.
Recorded results
The currently published results are from 9 October 2026: Jev, Clef, Clef Flash, OpenAI's Decisions API and a GPT-6 Luna structured-output baseline on the same 100 test cases and instructions, one request at a time, no retries and caching disabled. Values in parentheses are 95% intervals.
The intervals describe uncertainty in each model's observed rate. This card does not include a paired statistical comparison of the differences between models. Per-label results, every observed failure and a case-by-case explorer are on the benchmark page.
The primary metric in the manifest is macro-F1 with equal weight for all five labels. The results include each model's per-case probability distribution where the provider returns one; the GPT-6 Luna structured-output baseline returns none. Probability coverage, expected calibration error, Brier score and reliability bins are on the benchmark page. Token-based cost figures on the site are dated estimates, not provider bills.
Limitations
- Fictional English-only messages and balanced labels do not estimate production intent prevalence or performance across languages, channels and industries.
- Policy-specific precedence can differ from another company's ownership boundaries. Review labels against your actual routing policy before adopting results.
- Only 100 measured test cases and 20 clarification cases support the current comparison. Read intervals and observed failures; one-case differences do not establish a stable model ordering.
needs_clarificationmeasures insufficient context under this policy. It is not a general out-of-domain or safety detector.- One run cannot establish latency elsewhere or provider reliability under production load.
- Public examples, AI-assisted authoring and disclosed reviews do not prove absence of model training contamination.
- Actual support messages may contain personal information; none were used here. This dataset is not a replacement for privacy review of an organization's own evaluation data.
Reuse and citation
Attribute the frozen dataset to DecisionModelHub, cite this version and label policy, and disclose changes and development or test use. Keep modified datasets under a new identifier or version; do not present a tuned or public-set rerun as unseen test performance.
@misc{decisionmodelhub_support_routing_v1,
title = {Support Routing v1},
author = {DecisionModelHub},
year = {2026},
howpublished = {\url{https://decisionmodelhub.com/benchmarks/support-routing}},
note = {CC BY 4.0. Content hash sha256:2a2683c23a9cc1aa0d8996762c1b0d53c8bc55d513f8e22a89354098943a0b5b}
}