Team Ai
Datasetpublic

decisionmodelhub/support-routing

Support Routing v1 125 fictional English support messages, each labelled with the team that should handle it under a written routing policy, plus the recorded answers of five models on the 100 held-out test cases. Published by DecisionModelHub under CC BY 4.0. The task: given one support message and the fixed policy, select exactly one team for the primary actionable issue: billing, technical, account, sales or needs_clarification. The companion yes/no dataset is… See the full description on the dataset page: https://huggingface.co/datasets/decisionmodelhub/support-routing.

sourceHugging Facecc-by-4.0updated 23h agoView on Hugging Face
0likes117downloads
Dataset Card

Support Routing v1

125 fictional English support messages, each labelled with the team that should handle it under a written routing policy, plus the recorded answers of five models on the 100 held-out test cases. Published by DecisionModelHub under CC BY 4.0.

The task: given one support message and the fixed policy, select exactly one team for the primary actionable issue: billing, technical, account, sales or needs_clarification.

The companion yes/no dataset is decisionmodelhub/support-urgency.

Files

FileContents
cases.jsonAll 125 frozen cases with split, expected label, rationale, tags, category, difficulty and review state
cases_development.jsonl, cases_test.jsonlThe same cases, one file per split, for the dataset viewer and datasets
LABEL-POLICY.mdLabel definitions and precedence rules
manifest.jsonLicense, authoring and review procedure, counts and content hash
results.csv, results.jsonOne row per model per test case: the recorded answer, status, latency, token counts and, where the provider returns one, the probability for each label
results.jsonlThe same result rows, one per line, for the dataset viewer and datasets
results-manifest.jsonThe run's conditions, model revisions and file hashes

Dataset content hash: sha256:2a2683c23a9cc1aa0d8996762c1b0d53c8bc55d513f8e22a89354098943a0b5b. Label-policy hash: sha256:476220aef7fd81e74c78b1453e7558c15feda33bdd965ea4d64e2bf39c18fd42. These hashes describe the frozen content, not this card or the derived .jsonl files.

Load it

python
from datasets import load_dataset

cases = load_dataset("decisionmodelhub/support-routing", "cases")  # splits: development, test
results = load_dataset("decisionmodelhub/support-routing", "results", split="test")

Task and intended use

The task evaluates bounded routing under an explicit policy. It does not authorize refunds, account changes, purchases or remediation.

The policy routes the immediate blocker before the eventual goal: an upgrade page crash is technical; inability to sign in to obtain an invoice is an account issue. Multiple intents do not automatically imply clarification. Instructions embedded in customer messages are untrusted input; route the actual support request, or select clarification if none can be established.

Provenance and review

The manifest records original fictional scenarios drafted with AI assistance, an independent AI review, and two independent human reviews of all 125 cases on 6 October 2026. Human reviewer identities are retained privately; the public record reports the review count and procedure.

Normalized exact matching and semantic comparison covered catalog, pilot and policy examples and both splits. Held-out members of identified overlap clusters were rewritten and checked for label intent. Known example overlap is development-only. Private customer conversations, playground traffic and unlicensed third-party datasets are explicitly excluded.

Splits and composition

SplitCasesCases per labelClearBoundaryAmbiguousAdversarial
Development25512652
Test1002040302010
Total1252552362512

Development cases may guide instruction design; the 100 test cases were held out from that process. Test messages and labels are now public for inspection and reuse. They are no longer a private contamination-resistant evaluation set; future tuning against them must be disclosed and requires a fresh held-out set for independent performance claims.

Each case records case_id, split, the message (state), question_id, the expected label, a rationale, tags, category, difficulty, source, label provenance, language, policy version and review status. Category and difficulty are independent.

Recorded results

The currently published results are from 9 October 2026: Jev, Clef, Clef Flash, OpenAI's Decisions API and a GPT-6 Luna structured-output baseline on the same 100 test cases and instructions, one request at a time, no retries and caching disabled. Values in parentheses are 95% intervals.

ModelValidCorrectOperational successMedian latency95th-percentile latency
Clef100/10092/10092% (85–96)271 ms434 ms
Clef Flash100/10087/10087% (79–92)195 ms292 ms
Decisions API (OpenAI)100/10094/10094% (88–97)111 ms197 ms
GPT-6 Luna (structured-output baseline)100/10098/10098% (93–99)1,325 ms2,476 ms
Jev100/10092/10092% (85–96)138 ms180 ms

The intervals describe uncertainty in each model's observed rate. This card does not include a paired statistical comparison of the differences between models. Per-label results, every observed failure and a case-by-case explorer are on the benchmark page.

The primary metric in the manifest is macro-F1 with equal weight for all five labels. The results include each model's per-case probability distribution where the provider returns one; the GPT-6 Luna structured-output baseline returns none. Probability coverage, expected calibration error, Brier score and reliability bins are on the benchmark page. Token-based cost figures on the site are dated estimates, not provider bills.

Limitations

  • —Fictional English-only messages and balanced labels do not estimate production intent prevalence or performance across languages, channels and industries.
  • —Policy-specific precedence can differ from another company's ownership boundaries. Review labels against your actual routing policy before adopting results.
  • —Only 100 measured test cases and 20 clarification cases support the current comparison. Read intervals and observed failures; one-case differences do not establish a stable model ordering.
  • —needs_clarification measures insufficient context under this policy. It is not a general out-of-domain or safety detector.
  • —One run cannot establish latency elsewhere or provider reliability under production load.
  • —Public examples, AI-assisted authoring and disclosed reviews do not prove absence of model training contamination.
  • —Actual support messages may contain personal information; none were used here. This dataset is not a replacement for privacy review of an organization's own evaluation data.

Reuse and citation

Attribute the frozen dataset to DecisionModelHub, cite this version and label policy, and disclose changes and development or test use. Keep modified datasets under a new identifier or version; do not present a tuned or public-set rerun as unseen test performance.

bibtex
@misc{decisionmodelhub_support_routing_v1,
  title        = {Support Routing v1},
  author       = {DecisionModelHub},
  year         = {2026},
  howpublished = {\url{https://decisionmodelhub.com/benchmarks/support-routing}},
  note         = {CC BY 4.0. Content hash sha256:2a2683c23a9cc1aa0d8996762c1b0d53c8bc55d513f8e22a89354098943a0b5b}
}