kaushik-systalyze/customer-transcript-analytics
Customer Transcript Analytics Curated customer-support and meeting transcripts mapped to a single fixed "analyze this transcript → compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows range from short support chats… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-analytics.
Customer Transcript Analytics
Curated customer-support and meeting transcripts mapped to a single fixed "analyze this transcript → compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads.
Motivation and intended use
This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows range from short support chats of a few hundred tokens to long multi-party meeting segments, and every row carries token accounting (input_tokens, expected_max_new_tokens, prefill/decode ratio estimates, length_bucket) so throughput can be analyzed by prompt length.
It is not a correctness benchmark for intent, sentiment, resolution, or summarization quality. Model outputs generated from these prompts should not be interpreted as validated labels.
Prompt template
Every row uses this exact template (the static prefix before the transcript is constant across rows; per-row static_prefix_tokens / static_prefix_reuse quantify it for prefix-caching analysis):
Analyze this customer-support transcript.
Return compact JSON with:
- intent
- sentiment
- resolved
- escalation_needed
- summary
Transcript:
{{transcript}}expected_max_new_tokens is a documented heuristic, not a measurement: min(256, max(96, 96 + input_tokens // 64)). Token accounting is approximate and uses the Qwen/Qwen3.5-9B tokenizer (add_special_tokens=False).
Example usage (OpenAI-style request)
Each row's prompt is sent as a single user message, and the row's expected_max_new_tokens sets max_tokens. The model's JSON is generated at benchmark time — it is not stored in the dataset and must not be read as a validated label. The responses below are illustrative only.
Short customer-support row (short_control, ABCD, 189 input tokens):
{
"model": "Qwen/Qwen3.5-9B",
"temperature": 0,
"max_tokens": 98,
"messages": [
{
"role": "user",
"content": "Analyze this customer-support transcript.\n\nReturn compact JSON with:\n- intent\n- sentiment\n- resolved\n- escalation_needed\n- summary\n\nTranscript:\nASSISTANT: Hello, how can i help you today\nUSER: I need to check on my order because the delivery time is wrong\n[... truncated for display ...]\nUSER: No, thank you\nASSISTANT: Okay, have a good day"
}
]
}Illustrative response:
{
"intent": "delivery_time_inquiry",
"sentiment": "neutral",
"resolved": true,
"escalation_needed": false,
"summary": "Customer asked why their delivery time seemed wrong; agent confirmed the order is on schedule for standard 5-day shipping."
}This particular row also carries gold metadata passed through from the source (target_json = {"intent": "status_delivery_time", ...}, has_gold_intent=true), which a future quality check could compare against — the throughput benchmark itself does not score outputs.
Long meeting segment (long_dialogue, AMI window, 4,805 input tokens, 4k_8k bucket): identical request shape, but with ~4.8k tokens of speaker-prefixed meeting transcript in the user message and max_tokens: 171. Such rows are prefill-dominated (prefill_ratio_estimate ≈ 0.97), in contrast to the short case above.
Notes: the 34-token static prefix is identical across all rows, so engines with prefix caching will reuse it (static_prefix_tokens / static_prefix_reuse quantify this per row). If your server supports response_format: {"type": "json_object"}, it keeps outputs compact, though the prompt already encourages that and correctness is not evaluated.
Splits
- `source` — Customer-relevant transcript-analytics prompts (demo believability). Mostly 512–4k-token task-oriented customer-service dialogues.
- `long_dialogue` — Longer transcript-like prompts for long-context throughput and throughput-prediction validation, covering 512 to 16k input tokens with most rows between 2k and 8k.
- `short_control` — Short realistic workloads (128–1024 input tokens) for benchmarking short-prompt serving behaviour.
- `holdout_eval` — Held out for predicted-vs-observed throughput validation ONLY. Not a quality-eval split. Do not use it to fit or select anything. Transcripts are disjoint from all other splits (normalized-hash check) and drawn from all source families and prompt lengths.
Source mix by split
- source (1566 rows): Taskmaster2: 544, SalesBot: 363, Taskmaster3: 331, MULTIWOZ2_2: 232, ABCD: 61, SGD: 35
- long_dialogue (1217 rows): QMSum: 716, ICSI: 273, AMI: 228
- short_control (3000 rows): ABCD: 545, SGD: 495, Taskmaster3: 488, Taskmaster2: 459, MULTIWOZ2_2: 434, bitext/Bitext-customer-support-llm-chatbot-training-dataset: 391, SalesBot: 188
- holdout_eval (1049 rows): QMSum: 200, Taskmaster2: 167, Taskmaster3: 150, MULTIWOZ2_2: 107, SGD: 93, ABCD: 88, SalesBot: 85, bitext/Bitext-customer-support-llm-chatbot-training-dataset: 61, ICSI: 59, AMI: 39
Token statistics
Length-bucket counts
Length buckets on input_tokens: short_under_512 (<512), 512_1k, 1k_2k, 2k_4k, 4k_8k (4096–8192 inclusive), 8k_16k (8193–16384).
Licenses and redistribution
Each row carries license, redistribution_status, and source_attribution columns. Per-source terms:
Notes: SGD-derived rows remain CC BY-SA 4.0 (share-alike); Bitext-derived rows remain CDLA-Sharing-1.0. SalesBot has no upstream license; it is included by explicit dataset-owner decision because its text is fully synthetic (user-simulator/GPT-3 generated) with no third-party human content — treat those rows accordingly if that risk profile does not fit your use.
Excluded sources
The following sources from the original target list were excluded (no raw rows included) because their terms do not clearly allow public redistribution:
- TweetSumm — Summary annotations are CC0, but the underlying tweet text comes from the Kaggle 'Customer Support on Twitter' (TWCS) corpus, licensed CC BY-NC-SA 4.0 (non-commercial). Upstream distributes only tweet IDs for this reason. (https://github.com/guyfe/Tweetsumm ; https://www.kaggle.com/datasets/thoughtvector/customer-support-on-twitter)
- ConvoSumm — Conflicting license signals (repo CC BY-SA 4.0 vs DialogStudio 'No License') and third-party content (NYT comments, Reddit, StackExchange) with unclear redistribution provenance; also off-domain for customer support. (https://github.com/Yale-LILY/ConvoSumm)
- MediaSum — Upstream restricts use to research purposes only (NPR/CNN copyrighted transcripts); not publicly redistributable. (https://github.com/zcgzcgzcg1/MediaSum (README usage restriction))
- DialogSum — Upstream README states CC BY-NC-SA 4.0 (non-commercial); sources include DailyDialog which carries the same restriction. (https://github.com/cylnlp/dialogsum)
- SAMSum — CC BY-NC-ND 4.0 (non-commercial, no derivatives). (https://huggingface.co/datasets/knkarthick/samsum (card: cc-by-nc-nd-4.0))
- tuanhqv123/meeting-summarization — Mixed-source dataset whose card admits roughly half is MeetingBank (CC BY-NC 4.0, non-commercial) and lists license 'other'. (https://huggingface.co/datasets/tuanhqv123/meeting-summarization)
- monodox/customer-support-and-service — License is fine (Apache-2.0) but the dataset contains only 8 nested records — too trivial to contribute. (https://huggingface.co/datasets/monodox/customer-support-and-service)
Construction
- Built from raw
Salesforce/dialogstudioJSON files (configs listed inmetadata.json→source_files_used; for large task-oriented configs a deterministic subset of dialogue files was used) plus the Bitext CSV. - Task-oriented dialogues are flattened from the DialogStudio
logfield intoUSER:/ASSISTANT:turns; meeting transcripts come fromoriginal dialog info → dialog historyasSPEAKER: utterancelines. - Long-meeting segmentation: AMI/ICSI/QMSum meetings run 9k–18k+ tokens, so rows in the 1k–8k buckets are contiguous, non-overlapping segments of a single meeting (never concatenations of unrelated transcripts). QMSum is cut on its own annotated topic spans; AMI/ICSI use utterance-aligned windows. A meeting contributes either windows or one whole-meeting row (≤16k), never both, to avoid shared long prefixes that would skew prefix-cache behaviour.
- Rows over 16384 input tokens were rejected (never truncated) and are counted per source in
source_report.json. Deduplication is by normalized transcript hash;holdout_evaltranscripts are disjoint from all other splits. - Built by a deterministic fixed-seed pipeline (seed 42).
Label policy
- Existing source labels are preserved when trivially available (
reference_intentfrom ABCD subflow / SalesBot intent / Bitext intent;reference_summaryfrom AMI abstractive summaries and QMSum query answers). - No labels are synthesized — nothing is LLM-generated, heuristically inferred, or guessed from text.
has_gold_*flags are true only for fields that came directly from source data. - No source provides sentiment, resolution, or escalation labels, so
reference_sentiment,reference_resolved,reference_escalation_neededare null everywhere and the correspondinghas_gold_*flags are false. target_jsonis assembled only from available source fields and may be partial or null.- The benchmark runs the model on
promptand generates fresh outputs; the reference fields are optional metadata for future quality checks only.
Limitations
- Several sources are synthetic or proxy data (SalesBot and Bitext are fully synthetic; Taskmaster/SGD/MultiWOZ are Wizard-of-Oz or simulated task-oriented dialogues) — they look like customer service but are not real support calls.
- Meeting transcripts (AMI/ICSI/QMSum) are research-meeting recordings, not customer service; they provide realistic long-transcript shape, not customer-support content.
- Long rows are contiguous segments of longer meetings (see Construction).
long_dialogueis modest in size because most long-form sources (MediaSum, MeetingBank, SAMSum, DialogSum) failed license review.- Intended for inference throughput experiments, not model quality evaluation.
Files
data/<split>-*.parquet— the four splitsmetadata.json— build configuration (seed, tokenizer, template, files used)token_stats.json— per-split token statisticssource_report.json— per-source row accounting, license evidence, exclusions
