Team Ai
Datasetpublic

gradium/tts-eval-customer-support-202608

TTS Hard Cases — Customer Support A compact, text-only evaluation set of hard cases for text-to-speech models, focused on the customer support use case. Published by Gradium. Most TTS benchmarks measure naturalness on ordinary prose, where modern models are already close to saturated. Production voice agents fail somewhere else: on the literal payload of a support call — the order number, the email address, the spelled-out surname, the date, the amount refunded. A voice that… See the full description on the dataset page: https://huggingface.co/datasets/gradium/tts-eval-customer-support-202608.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
1likes56downloads
Dataset Card

TTS Hard Cases — Customer Support

A compact, text-only evaluation set of hard cases for text-to-speech models, focused on the customer support use case. Published by Gradium.

Most TTS benchmarks measure naturalness on ordinary prose, where modern models are already close to saturated. Production voice agents fail somewhere else: on the literal payload of a support call — the order number, the email address, the spelled-out surname, the date, the amount refunded. A voice that sounds perfect but reads #ORD-4471-XB as "ord four thousand four hundred seventy-one xb" has failed the call.

This dataset isolates exactly those failure modes, in five languages.

At a glance

Rows500
Unique items100
LanguagesEN, FR, ES, PT, DE
Criteria10, grouped in 2 categories
Structure10 criteria × 10 items × 5 languages
ModalityText only — no reference audio
Splittest (evaluation-only, no training split)
LicenseCC BY 4.0

Dataset structure

One row per (item, language). Fields:

FieldTypeDescription
item_idstringItem identifier, formatted <criterion>-<index> (e.g. 1.4-07). Shared across the 5 languages.
categorystringAtomic Criteria or Composite Criteria.
criterionstringCriterion number (1.1 … 2.3).
criterion_namestringHuman-readable criterion name (e.g. Date).
languagestringEN, FR, ES, PT, DE.
sentencestringThe text to synthesize.

Example row:

csv
item_id,category,criterion,criterion_name,language,sentence
1.7-01,Atomic Criteria,1.7,Email,EN,For questions about your order, contact our support team at help@gmail.com for a reply within a day.

Parallel, not translated

Each item_id poses the same challenge in all five languages, but the content is localized, not translated. The spelling item uses H-I-G-G-I-N-B-O-T-H-A-M in English and S-C-H-W-E-I-N-S-T-E-I-G-E-R in German; the date item uses locale-plausible date formats. This keeps difficulty comparable across languages while avoiding translationese and entity names that a native speaker would never encounter in that language.

The taxonomy

1. Atomic criteria (350 rows)

Short-to-medium utterances that each stress a single verbalization behaviour.

#CriterionItemsWhat it probesEnglish example
1.1Spelling10Letter-by-letter readout, hyphen- and comma-separated, initials with periodsThe town is spelled L, O, U, G, H, B, O, R, O, U, G, H.
1.2Acronyms10Letter-name vs. word pronunciation (FBI vs. NATO), embedded in long sentencesThe FBI shared the classified intelligence with NATO partners last Thursday…
1.3Alphanumerical Tokens10Mixed letter/digit tokens: USB-C, Wi-Fi 6, GPT-4, B2B, 24/7, H2O, ISO-9Our B2B platform now runs smoothly on iOS 17, Android 14, and USB-C connectors…
1.4Date10Ambiguous and mixed formats in one utterance: 04/03/2025, 03-24-25, 18.06.2025, 4 Mar '25, decades (the 1990s)Your cleaning set for 04/03/2025 is moved to Monday, March 17th, and X-rays are on 03-24-25 downtown.
1.5Regular Numbers10Reading mode by context: years, street numbers, room numbers, flight numbers, times, quantities — often several in one sentenceFlight 812 from terminal 4 boards at gate 27, leaving 6:45 PM…
1.6Floating and Large Numbers10Decimals, thousands separators, magnitudes, units, percentages, currencyYour balance is 12,485.67 dollars, and a pending transfer of 2,750.99 processes tomorrow…
1.7Email10Addresses with dots, underscores, digits, and common domains…send a report to first_last@hotmail.com with screenshots.

2. Composite criteria (150 rows)

Realistic support-agent turns that combine several atomic difficulties in a single utterance, with conversational prosody on top: greetings, empathy, reassurance, questions.

#CriterionItemsScenarioEnglish example
2.1Orders10E-commerce: order and SKU references, addresses and postcodes, prices, card digits, emailsI've got your order #ORD-4471-XB right here — that's the navy wool coat, SKU WC-228-NV, shipping to 12 Victoria Street, London SW1A 2AA…
2.2IT Ticket10Telecom/IT support: account and ticket IDs, serials, IP and MAC addresses, plan codes, appointment windowsAccount ACC-44192-TM, I can see your modem, serial SN-882-44719, is offline since 2:14 PM… Can you confirm the IP shows 192.168.1.1 on your end?
2.3Claims10Insurance and healthcare: policy, claim, member and patient IDs, ICD codes, NPI numbers, phone numbers — delivered in emotionally loaded contextI have your policy POL-55821-HM open right now, and I'm filing claim CLM-90217-AC for the accident on April 18th…

Length grows with difficulty: atomic spelling items run 5–19 words, composite items reach 73 words.

Usage

python
from datasets import load_dataset

ds = load_dataset("gradium/tts-eval-customer-support-202608", split="test")

# English composite items only
hard = ds.filter(lambda r: r["language"] == "EN" and r["category"] == "Composite Criteria")
for row in hard:
    audio = my_tts_model(row["sentence"], language=row["language"])
    save(f"{row['item_id']}_{row['language']}.wav", audio)

Or straight from the CSV:

python
import pandas as pd
df = pd.read_csv("data.csv")
df.groupby(["category", "criterion_name"]).size()

Suggested evaluation protocol

The dataset ships text only: there is no reference audio, and there is deliberately no single "correct" phonetic transcription — acceptable verbalizations vary by voice, locale and register. Two protocols work well:

  1. 1.Human scoring (recommended). Synthesize each item, then have a native speaker score the critical tokens — the identifier, the date, the address — as correct / incorrect / ambiguous. Report accuracy per criterion and per language rather than a single aggregate; the interesting signal is which criterion a model breaks on.
  2. 2.ASR-based proxy. Transcribe the synthesized audio with a strong ASR model and compare against a normalized expansion of the source text. Cheap and repeatable, but it inherits the ASR's own biases on exactly these token types, so treat disagreements as items to review by hand.

Because pass/fail here is close to binary per item, per-criterion accuracy over 10 items is coarse — use it to locate failure modes, not to rank two close models.

Limitations and scope

  • —Small by design. 100 items is a diagnostic probe, not a statistical benchmark. Confidence intervals on a 10-item criterion are wide.
  • —No audio, no ground-truth phonemization. Scoring requires a human or an ASR model in the loop.
  • —Five European languages, No tonal languages, no non-Latin scripts, no code-switching items.
  • —Synthetic scenarios. Items are written to be realistic support turns, not sampled from real call transcripts. Names, addresses, order numbers, emails, policy and patient identifiers are all invented; any resemblance to real records is coincidental. Email domains are real providers but the local parts are placeholders.
  • —Written register. Items are clean, well-punctuated text. They do not cover disfluent or LLM-generated input with markdown, emoji, or broken casing — a separate failure mode for production voice agents.

License

Released under CC BY 4.0. You may use, share and adapt the dataset, including commercially, with attribution.

Citation

bibtex
@misc{gradium2026ttshardcases,
  title  = {TTS Hard Cases — Customer Support},
  author = {Gradium},
  year   = {2026},
  url    = {https://gradium.ai},
  note   = {A multilingual diagnostic set of hard text-to-speech cases for customer support}
}

About

Built and published by Gradium.