gradium/tts-eval-customer-support-202608
TTS Hard Cases — Customer Support A compact, text-only evaluation set of hard cases for text-to-speech models, focused on the customer support use case. Published by Gradium. Most TTS benchmarks measure naturalness on ordinary prose, where modern models are already close to saturated. Production voice agents fail somewhere else: on the literal payload of a support call — the order number, the email address, the spelled-out surname, the date, the amount refunded. A voice that… See the full description on the dataset page: https://huggingface.co/datasets/gradium/tts-eval-customer-support-202608.
TTS Hard Cases — Customer Support
A compact, text-only evaluation set of hard cases for text-to-speech models, focused on the customer support use case. Published by Gradium.
Most TTS benchmarks measure naturalness on ordinary prose, where modern models are already close to saturated. Production voice agents fail somewhere else: on the literal payload of a support call — the order number, the email address, the spelled-out surname, the date, the amount refunded. A voice that sounds perfect but reads #ORD-4471-XB as "ord four thousand four hundred seventy-one xb" has failed the call.
This dataset isolates exactly those failure modes, in five languages.
At a glance
Dataset structure
One row per (item, language). Fields:
Example row:
item_id,category,criterion,criterion_name,language,sentence
1.7-01,Atomic Criteria,1.7,Email,EN,For questions about your order, contact our support team at help@gmail.com for a reply within a day.Parallel, not translated
Each item_id poses the same challenge in all five languages, but the content is localized, not translated. The spelling item uses H-I-G-G-I-N-B-O-T-H-A-M in English and S-C-H-W-E-I-N-S-T-E-I-G-E-R in German; the date item uses locale-plausible date formats. This keeps difficulty comparable across languages while avoiding translationese and entity names that a native speaker would never encounter in that language.
The taxonomy
1. Atomic criteria (350 rows)
Short-to-medium utterances that each stress a single verbalization behaviour.
2. Composite criteria (150 rows)
Realistic support-agent turns that combine several atomic difficulties in a single utterance, with conversational prosody on top: greetings, empathy, reassurance, questions.
Length grows with difficulty: atomic spelling items run 5–19 words, composite items reach 73 words.
Usage
from datasets import load_dataset
ds = load_dataset("gradium/tts-eval-customer-support-202608", split="test")
# English composite items only
hard = ds.filter(lambda r: r["language"] == "EN" and r["category"] == "Composite Criteria")
for row in hard:
audio = my_tts_model(row["sentence"], language=row["language"])
save(f"{row['item_id']}_{row['language']}.wav", audio)Or straight from the CSV:
import pandas as pd
df = pd.read_csv("data.csv")
df.groupby(["category", "criterion_name"]).size()Suggested evaluation protocol
The dataset ships text only: there is no reference audio, and there is deliberately no single "correct" phonetic transcription — acceptable verbalizations vary by voice, locale and register. Two protocols work well:
- Human scoring (recommended). Synthesize each item, then have a native speaker score the critical tokens — the identifier, the date, the address — as correct / incorrect / ambiguous. Report accuracy per criterion and per language rather than a single aggregate; the interesting signal is which criterion a model breaks on.
- ASR-based proxy. Transcribe the synthesized audio with a strong ASR model and compare against a normalized expansion of the source text. Cheap and repeatable, but it inherits the ASR's own biases on exactly these token types, so treat disagreements as items to review by hand.
Because pass/fail here is close to binary per item, per-criterion accuracy over 10 items is coarse — use it to locate failure modes, not to rank two close models.
Limitations and scope
- Small by design. 100 items is a diagnostic probe, not a statistical benchmark. Confidence intervals on a 10-item criterion are wide.
- No audio, no ground-truth phonemization. Scoring requires a human or an ASR model in the loop.
- Five European languages, No tonal languages, no non-Latin scripts, no code-switching items.
- Synthetic scenarios. Items are written to be realistic support turns, not sampled from real call transcripts. Names, addresses, order numbers, emails, policy and patient identifiers are all invented; any resemblance to real records is coincidental. Email domains are real providers but the local parts are placeholders.
- Written register. Items are clean, well-punctuated text. They do not cover disfluent or LLM-generated input with markdown, emoji, or broken casing — a separate failure mode for production voice agents.
License
Released under CC BY 4.0. You may use, share and adapt the dataset, including commercially, with attribution.
Citation
@misc{gradium2026ttshardcases,
title = {TTS Hard Cases — Customer Support},
author = {Gradium},
year = {2026},
url = {https://gradium.ai},
note = {A multilingual diagnostic set of hard text-to-speech cases for customer support}
}About
Built and published by Gradium.
