decisionmodelhub/support-urgency
Support Urgency v1 800 fictional English support messages, each with a yes/no answer to one question, "is this urgent?", plus the recorded answers and probabilities of five models on every case. Published by DecisionModelHub under CC BY 4.0. A message is urgent only when the problem is happening now, stops people from completing a task, and affects more than one person. The customer's own priority, the tone and any instruction inside the message do not change the answer. The… See the full description on the dataset page: https://huggingface.co/datasets/decisionmodelhub/support-urgency.
Support Urgency v1
800 fictional English support messages, each with a yes/no answer to one question, "is this urgent?", plus the recorded answers and probabilities of five models on every case. Published by DecisionModelHub under CC BY 4.0.
A message is urgent only when the problem is happening now, stops people from completing a task, and affects more than one person. The customer's own priority, the tone and any instruction inside the message do not change the answer.
The companion routing dataset is decisionmodelhub/support-routing.
Files
Dataset content hash: sha256:decc1c2d9ac55655e2f959b4585dc414a3e706f5e5ea627efe00dfbe0de7e24a. Label-policy hash: sha256:416b2735ea0b246ecdc427e96d52bfc9cd1ea594a29360c132456a69069d17f3. These hashes describe the frozen content, not this card or the derived .jsonl files.
Load it
from datasets import load_dataset
cases = load_dataset("decisionmodelhub/support-urgency", "cases") # splits: development, test
results = load_dataset("decisionmodelhub/support-urgency", "results") # splits: development, testTask and intended use
Given one English support message, answer whether it is urgent: yes or no.
The dataset is meant for models that answer a yes/no question with a probability. It supports measuring how well that probability separates urgent messages from the rest, where a threshold should sit, and whether a threshold chosen for one model works for another.
How the cases were made
- Fact sheets. A seeded script drew 1,120 fact sheets. Each records a case kind (which of the three facts hold), a subject, a sender, and style: tone, length, writing quality, whether the customer claims a priority or embeds an instruction, and whether the message carries an aside.
- Labels by rule. Code computed each answer, category, difficulty and rationale from the fact sheet. No person or model judged a label.
- Messages. Claude Opus 5.5 wrote one message from each fact sheet.
- Read-back. Claude Sonnet 5.5 and Claude Haiku 5.5 each read every message without the fact sheet or the answer, and stated the three facts and the answer. A case was kept only when both matched the fact sheet. 12 cases failed and were removed, not rewritten.
- Selection. 1 near duplicate was removed. From the remaining cases a seeded draw took the planned number of each kind and assigned the splits.
No human reviewed the cases. The fact sheets and the selection come from a seeded script; the written messages and read-backs are model output and are not regenerated.
Splits and composition
Half of each split is urgent. Of the non-urgent cases, 288 are near misses with exactly one fact false and 112 describe no problem. By computed difficulty: 161 easy, 318 medium, 321 hard.
The development split is for choosing instructions and thresholds. The test split is for measurement only. Both are public, so neither is a private contamination-resistant set.
Style is drawn the same way for every kind, so it does not predict the answer.
Recorded results
The published measurements are from 9 October 2026: Jev, Clef, Clef Flash, OpenAI's Decisions API and a GPT-6 Luna structured-output baseline, one request per case and model, no retries, caching disabled. The figures below are on the 600 test cases, counting an answer as yes when the probability of yes is 0.5 or more. Values in parentheses are 95% intervals.
"Own threshold" is the cut-off that worked best for that model on the 200 development cases. The baseline returns yes or no without a probability, so it has no threshold, ranking or calibration figures. Thresholds, calibration and results by case kind are on the benchmark page. Cost figures on the site are list-price estimates from reported token counts, not provider bills.
Limitations
- The messages are fictional, English-only and balanced. They do not estimate how often real support messages are urgent, or performance across languages, channels and industries.
- The rule is one definition of urgency. Another organization may count a single blocked customer, or a slow but working product, as urgent.
- The writer and both read-back models come from one model family. None is a measured model, but their shared habits shape how the messages read.
- A case that either read-back could not decode was removed, so the hardest phrasings are under-represented. Both read-back models agreed with the rule on every answer, which suggests a strong general-purpose model given the rule will find the yes/no decision easy. The dataset is built to compare probabilities and thresholds, not only yes/no accuracy.
- Every fact is stated or shown in the message. Messages where a fact cannot be established are not in the dataset, so it does not measure behaviour on underdetermined requests.
- The fact sheets pair senders and product areas at random; some messages explain an unusual pairing in a way a real customer would not need to.
Reuse and citation
Attribute the frozen dataset to DecisionModelHub, cite this version and label policy, and disclose changes and how each split was used. Keep modified datasets under a new identifier or version.
@misc{decisionmodelhub_support_urgency_v1,
title = {Support Urgency v1},
author = {DecisionModelHub},
year = {2026},
howpublished = {\url{https://decisionmodelhub.com/benchmarks/support-urgency}},
note = {CC BY 4.0. Content hash sha256:decc1c2d9ac55655e2f959b4585dc414a3e706f5e5ea627efe00dfbe0de7e24a}
}