Team Ai
Datasetpublic

decisionmodelhub/support-urgency

Support Urgency v1 800 fictional English support messages, each with a yes/no answer to one question, "is this urgent?", plus the recorded answers and probabilities of five models on every case. Published by DecisionModelHub under CC BY 4.0. A message is urgent only when the problem is happening now, stops people from completing a task, and affects more than one person. The customer's own priority, the tone and any instruction inside the message do not change the answer. The… See the full description on the dataset page: https://huggingface.co/datasets/decisionmodelhub/support-urgency.

sourceHugging Facecc-by-4.0updated 1d agoView on Hugging Face
0likes34downloads
Dataset Card

Support Urgency v1

800 fictional English support messages, each with a yes/no answer to one question, "is this urgent?", plus the recorded answers and probabilities of five models on every case. Published by DecisionModelHub under CC BY 4.0.

A message is urgent only when the problem is happening now, stops people from completing a task, and affects more than one person. The customer's own priority, the tone and any instruction inside the message do not change the answer.

The companion routing dataset is decisionmodelhub/support-routing.

Files

FileContents
cases.jsonAll 800 frozen cases with split, expected answer, fact sheet, rationale and tags
cases_development.jsonl, cases_test.jsonlThe same cases, one file per split, for the dataset viewer and datasets
LABEL-POLICY.mdThe rule, what does not change the answer, and the case kinds
manifest.jsonLicense, how the cases were made and checked, counts and content hash
results.csv, results.jsonOne row per model per case: the recorded answer, the probability of yes where the provider returns one, status, latency and token counts
results_development.jsonl, results_test.jsonlThe same result rows, one file per split, for the dataset viewer and datasets
results-manifest.jsonThe run's conditions and model revisions

Dataset content hash: sha256:decc1c2d9ac55655e2f959b4585dc414a3e706f5e5ea627efe00dfbe0de7e24a. Label-policy hash: sha256:416b2735ea0b246ecdc427e96d52bfc9cd1ea594a29360c132456a69069d17f3. These hashes describe the frozen content, not this card or the derived .jsonl files.

Load it

python
from datasets import load_dataset

cases = load_dataset("decisionmodelhub/support-urgency", "cases")      # splits: development, test
results = load_dataset("decisionmodelhub/support-urgency", "results")  # splits: development, test

Task and intended use

Given one English support message, answer whether it is urgent: yes or no.

The dataset is meant for models that answer a yes/no question with a probability. It supports measuring how well that probability separates urgent messages from the rest, where a threshold should sit, and whether a threshold chosen for one model works for another.

How the cases were made

  1. 1.Fact sheets. A seeded script drew 1,120 fact sheets. Each records a case kind (which of the three facts hold), a subject, a sender, and style: tone, length, writing quality, whether the customer claims a priority or embeds an instruction, and whether the message carries an aside.
  2. 2.Labels by rule. Code computed each answer, category, difficulty and rationale from the fact sheet. No person or model judged a label.
  3. 3.Messages. Claude Opus 5.5 wrote one message from each fact sheet.
  4. 4.Read-back. Claude Sonnet 5.5 and Claude Haiku 5.5 each read every message without the fact sheet or the answer, and stated the three facts and the answer. A case was kept only when both matched the fact sheet. 12 cases failed and were removed, not rewritten.
  5. 5.Selection. 1 near duplicate was removed. From the remaining cases a seeded draw took the planned number of each kind and assigned the splits.

No human reviewed the cases. The fact sheets and the selection come from a seeded script; the written messages and read-backs are model output and are not regenerated.

Splits and composition

SplitCasesUrgentAlready fixedNot startedStill completesOnly looks wrongOne personNo problem
Development2001001681682428
Test600300482448247284
Total8004006432643296112

Half of each split is urgent. Of the non-urgent cases, 288 are near misses with exactly one fact false and 112 describe no problem. By computed difficulty: 161 easy, 318 medium, 321 hard.

The development split is for choosing instructions and thresholds. The test split is for measurement only. Both are public, so neither is a private contamination-resistant set.

Style is drawn the same way for every kind, so it does not predict the answer.

Recorded results

The published measurements are from 9 October 2026: Jev, Clef, Clef Flash, OpenAI's Decisions API and a GPT-6 Luna structured-output baseline, one request per case and model, no retries, caching disabled. The figures below are on the 600 test cases, counting an answer as yes when the probability of yes is 0.5 or more. Values in parentheses are 95% intervals.

ModelValidCorrectMissed urgentFalse alarmsAUROCOwn thresholdMedian latency95th-percentile latency
Clef599/600586/600 (96.1–98.6%)13/3000/3000.99970.20343 ms617 ms
Clef Flash600/600570/600 (93.0–96.5%)6/30024/3000.99230.80278 ms467 ms
Decisions API (OpenAI)600/600597/600 (98.5–99.8%)2/3001/3000.99980.50127 ms275 ms
GPT-6 Luna (structured-output baseline)600/600598/600 (98.8–99.9%)0/3002/300not availablenot available1,573 ms2,798 ms
Jev600/600596/600 (98.3–99.7%)0/3004/3000.99990.62135 ms176 ms

"Own threshold" is the cut-off that worked best for that model on the 200 development cases. The baseline returns yes or no without a probability, so it has no threshold, ranking or calibration figures. Thresholds, calibration and results by case kind are on the benchmark page. Cost figures on the site are list-price estimates from reported token counts, not provider bills.

Limitations

  • —The messages are fictional, English-only and balanced. They do not estimate how often real support messages are urgent, or performance across languages, channels and industries.
  • —The rule is one definition of urgency. Another organization may count a single blocked customer, or a slow but working product, as urgent.
  • —The writer and both read-back models come from one model family. None is a measured model, but their shared habits shape how the messages read.
  • —A case that either read-back could not decode was removed, so the hardest phrasings are under-represented. Both read-back models agreed with the rule on every answer, which suggests a strong general-purpose model given the rule will find the yes/no decision easy. The dataset is built to compare probabilities and thresholds, not only yes/no accuracy.
  • —Every fact is stated or shown in the message. Messages where a fact cannot be established are not in the dataset, so it does not measure behaviour on underdetermined requests.
  • —The fact sheets pair senders and product areas at random; some messages explain an unusual pairing in a way a real customer would not need to.

Reuse and citation

Attribute the frozen dataset to DecisionModelHub, cite this version and label policy, and disclose changes and how each split was used. Keep modified datasets under a new identifier or version.

bibtex
@misc{decisionmodelhub_support_urgency_v1,
  title        = {Support Urgency v1},
  author       = {DecisionModelHub},
  year         = {2026},
  howpublished = {\url{https://decisionmodelhub.com/benchmarks/support-urgency}},
  note         = {CC BY 4.0. Content hash sha256:decc1c2d9ac55655e2f959b4585dc414a3e706f5e5ea627efe00dfbe0de7e24a}
}