arpieb/consumer-electronics-support
Dataset Card for Synthetic Multi-Turn Customer Support Tickets (Consumer Electronics) 99,930 wholly synthetic multi-turn customer-support conversations in a consumer-electronics retail domain, each with assigned ticket metadata, a machine coherence score, and provenance linking it to the run that produced it. Dataset Details Dataset Description Every conversation is fabricated by a language model from a committed domain prompt. No real support… See the full description on the dataset page: https://huggingface.co/datasets/arpieb/consumer-electronics-support.
Dataset Card for Synthetic Multi-Turn Customer Support Tickets (Consumer Electronics)
99,930 wholly synthetic multi-turn customer-support conversations in a consumer-electronics retail domain, each with assigned ticket metadata, a machine coherence score, and provenance linking it to the run that produced it.
Dataset Details
Dataset Description
Every conversation is fabricated by a language model from a committed domain prompt. No real support transcript, customer, or agent is involved at any point, and none was used as source material. The conversations are between a customer of a mid-sized online retailer selling consumer electronics — headphones, smart speakers, wearables, small home devices — and a support agent employed by that retailer, covering billing, technical, account, shipping, product and miscellaneous issues.
The dataset exists to support work that needs support-ticket-shaped text at scale — dialogue modelling, routing and triage experiments, evaluation harnesses, pipeline load testing — without the privacy exposure that real support archives carry.
Each record is one complete ticket: an ordered sequence of turns alternating customer and agent, starting with the customer, plus the metadata the conversation was generated to fit.
Composition is assigned before generation rather than measured after it. Each record's category, priority, channel, resolution status, turn count, subdomain and timestamps are seeded choices made from the run seed and the record's position, so the distribution is a property of the request rather than an accident of what the model happened to produce.
- Curated by: Robert Bates (arpieb)
- Funded by [optional]: Unfunded; personal project
- Shared by [optional]: Robert Bates (arpieb)
- Language(s) (NLP): en
- License: CC BY 4.0
Dataset Sources [optional]
- Repository: https://github.com/arpieb/ticket-dataset-generator
- Paper [optional]: None
- Demo [optional]: None
Uses
Direct Use
Suitable for:
- fine-tuning or evaluating multi-turn dialogue models on support-style exchanges — at ~100k records and ~100M characters of conversational content this is large enough to train a small model or adapt a larger one;
- ticket classification, routing, priority prediction and summarisation, where the label is known by construction rather than inferred — every record carries its category, priority, channel, resolution status and subdomain as assigned inputs, not as post-hoc annotation;
- stratified experiments: the 20 declared subdomains are balanced to within ~±2% of each other (4,890–5,121 records each), so per-subdomain slices are directly comparable;
- testing and benchmarking data pipelines that consume conversational JSONL, since every record is schema-validated and carries stable identifiers;
- teaching and demonstration, where a real support archive would be inappropriate to distribute.
Conversations are short by construction — 4 to 8 turns, mean 6.0, mean 1,003 characters of turn content per record. The corpus is not suitable for work on long-horizon dialogue, multi-issue tickets, or conversations that span sessions; none of those exist in it.
Out-of-Scope Use
- Not a sample of real support traffic. Nothing here was observed. Frequencies, phrasing, and problem distributions reflect a prompt, a requested composition and a model, not any real customer population. The composition was specified, so it cannot be evidence of anything. Do not use it to estimate how often anything happens in the world.
- Not a benchmark, and the coherence scores are not ground truth. They were produced by the same model that wrote the conversations, against a rubric that has no calibration record (see Limitations). Treat
coherence_scoreas a filtering artifact of this run. - Not a privacy-safe stand-in for de-identification research. Synthetic-by-construction is not the same problem as de-identified-after-the-fact, and this dataset cannot be used to validate a de-identification method.
- Not a basis for consequential decisions about real people. Do not train systems that triage, prioritise or deny real customer requests on this corpus alone; the priority and resolution labels encode a requested distribution, not a policy anyone validated.
Dataset Structure
One JSON object per line (JSONL, UTF-8). 99,930 records, 196,658,051 bytes. A single train split; no test split is provided, because splitting is the consumer's decision and depends on the task. Records are grouped by nothing in particular, so a random split is sound; a subdomain-stratified split is straightforward from the subdomain field.
Verified across all 99,930 records: every conversation starts with the customer and strictly alternates roles, and resolved_at is present when and only when resolution_status is resolved.
`record_index` is not contiguous. The run requested 100,000 records and wrote 99,930. Indices run 0–99,999 with 70 positions absent. The manifest accounts for 55 of these as attempts_exhausted; the remaining 15 are not accounted for by any recorded discard reason. Do not use record_index as a row number or assume max(record_index) + 1 == len(corpus).
Achieved composition, read from composition_achieved in the manifest and confirmed by counting the corpus. Worst-member drift from the requested distribution is 0.022pp (category.billing, 24.9785% against 25%):
(The composition_drift_pp block in the run report is miscomputed — it reports each member's drift as the negative of its requested share, as though the achieved share were zero. The drift above was recomputed from composition_achieved against composition_requested. See Limitations.)
Turn counts range 4–8, mean 6.006, near-uniform: 4 turns 19,725; 5 turns 20,013; 6 turns 20,009; 7 turns 20,328; 8 turns 19,855.
All 20 of the domain prompt's declared subdomains appear, between 4,890 (technical-device-wont-power-on) and 5,121 (shipping-damaged-on-arrival) records each.
metadata.created_at spans 2025-07-05T00:05:58Z to 2025-12-31T23:59:45Z, distributed across the six months (14,960–17,330 records per month; July is short because the window opens on the 5th).
Dataset Creation
Curation Rationale
Real support tickets are among the densest sources of personal data that exist, which makes them both valuable and largely undistributable. Scrubbing them after the fact is unreliable and irreversible once published. This dataset takes the other route: admit no real identifiers at any point, so there is nothing to scrub. Consumer-electronics support is a good domain for a first release at scale — the issues are mundane, the stakes are low, and nothing in the domain invites the model toward medical, financial or legal content that would carry risks a card like this one could not discharge.
The generator is the deliverable; this corpus is one output of it. The pipeline enforces reproducibility, provenance, and a blocking privacy scan as structural properties rather than review-time checks.
Source Data
Data Collection and Processing
Generated by ollama_chat/gpt-oss:20b, self-hosted, from a committed domain prompt document (consumer-electronics-support.md@8d67eaf56b0a) declaring 20 subdomains.
Per record: the pipeline assigns metadata from the seed, prompts the model for a conversation fitting that assignment, then applies four gates in order — structural validation, a blocking privacy scan, coherence judging, and schema validation. A record failing any gate is discarded and counted by reason; the slot is retried up to 3 times. Output is written to staging and moved to the release path only after the privacy floor is demonstrated against known canaries.
This run: 105,236 generated, 99,930 written — a 5.04% discard rate. Discards by reason: 2,144 coherence_below_threshold, 1,776 structural_invalid, 761 privacy_finding, 552 turn_count_out_of_range, 55 attempts_exhausted, 18 unjudgeable (5,306 total, which reconciles exactly with generated minus written). Zero duplicates. Zero records reached the output with a blocking privacy finding. The run was resumed 3 times and its recorded outcome is completed, verdict pass.
Reproduction inputs: run 8c56469b-f094-4686-90dc-29715bab4228, seed 42, code revision ab247fa42a66, corpus SHA-256 a49bcc9ca0d228d456ae6bd371e8517940629ecd9e97140d65c45cd8c56c755b (verified against the file).
The code revision was dirty. code_revision.dirty is true in the manifest: this corpus was generated from a working tree with uncommitted changes, so it cannot be reproduced from any commit. ab247fa42a66 names the nearest ancestor commit, not the code that ran.
Reproduction is in any case structural, not textual. No model sampling seed was set (models.generator.sampling_seed is null), so a rerun with the same seed, config, prompt and rubric lands every record on the same composition, subdomain, turn count and timestamps, but the conversation text differs on every run. This is a property of the method, not a defect of the run.
The run's resolved configuration is recorded in the manifest and is not the repository's configs/samples/release.toml: it used turns.max = 8 (not 12), time_window.start = 2025-07-05 (not 2025-07-01), max_concurrency = 4, checkpoint_interval = 100, and no budget ceiling. Use config-from-manifest rather than the sample config to recover what actually ran.
Who are the source data producers?
A language model, prompted by an automated pipeline. No humans produced any conversational content, and no human-authored text was used as source material.
Annotations [optional]
Annotation process
Each conversation is scored 0–1 for coherence by a second model call against a committed, versioned rubric (coherence-v2) with four weighted criteria: single_issue (0.30), role_consistency (0.25), conversational_flow (0.25), metadata_fit (0.20). The record's score is the weighted mean. Records below 0.8 are discarded before release; the score of every surviving record is retained. 2,144 records were discarded on this gate.
Score distribution (computed from the corpus — the manifest's coherence_score_distribution field was not populated, see Limitations): mean 0.960, median 1.0, min 0.80, max 1.0, standard deviation 0.053, across 211 distinct values. By bucket — 0.80–0.85: 2,886 (2.89%); 0.85–0.90: 3,952 (3.95%); 0.90–0.95: 28,049 (28.07%); 0.95–1.00: 6,079 (6.08%); exactly 1.00: 58,964 (59.01%). A further 26,018 records (26.04%) sit at exactly 0.90.
Who are the annotators?
ollama_chat/gpt-oss:20b — the same model, at the same version, that generated the conversations. See Limitations.
No human calibration exists for this rubric. calibration/ contains only its README; there is no coherence-v2.calibration.json and no record of any reviewer having read any sample of this corpus. The judge behind every score in this dataset is uncalibrated. The repository documents this as an accepted deferral (checklist item CHK063), not as an oversight — but a consumer of the scores should treat them accordingly.
Personal and Sensitive Information
The dataset contains no real personal data by construction. Identifier-shaped values are drawn from ranges reserved for fiction — RFC 2606 domains (example.com, .test, .invalid), NANP 555-0100–555-0199 phone numbers, published payment-card test numbers — so they cannot refer to anyone. The domain prompt makes this the primary control and forbids bare runs of nine or more digits, so that invented order and serial numbers cannot collide with government identifier or phone number shapes.
Every record was scanned before admission by the datafog-regex detector over turns[].content and scenario, covering CREDIT_CARD, EMAIL, IP_ADDRESS, PHONE and US_SSN. 761 records were blocked and discarded on this gate, and zero records with a blocking finding reached the output. Nothing was quarantined. No privacy exception has ever been approved for this project; privacy/exceptions.json does not exist.
The run report's per-record privacy counters are unusable: records_examined and fields_examined are both 0 and findings is empty, which contradicts the 761 recorded privacy discards. The gate demonstrably ran and blocked; the accounting of what it examined did not survive into the report. This card therefore cannot state how many fields were scanned or how many findings were exempt by reserved range. See Limitations.
The scan does not detect: non-US government identifiers, postal codes, full postal addresses, IBAN bank account numbers, or person names. These are declared gaps, restated here so a clean scan is not mistaken for coverage it does not provide. Person names in particular are present throughout and are fabricated.
Bias, Risks, and Limitations
The judge shares a model with the generator. ollama_chat/gpt-oss:20b wrote and scored every one of the 99,930 conversations. A model evaluating its own output is a known source of inflated agreement, and the score distribution is consistent with that.
The judge did not use the scale. Rubric v2 exists precisely because v1 collapsed to three values, and it instructs the judge to reserve 1.0 for "no observable flaw of this kind". 59.01% of records scored exactly 1.00 and a further 26.04% exactly 0.90 — 85% of the corpus on two values, and only 211 distinct values across ~100k records. The rubric's own stated failure mode has recurred in weaker form. The scores separate the bottom few percent from everything else and should not be read as a quality ordering within the corpus.
Nobody has checked the judge. There is no calibration record for coherence-v2. Every claim this corpus makes about quality rests on one uncalibrated model scoring its own output. This is the single largest gap in the dataset's provenance.
The provenance record has holes. Beyond the dirty code revision, several manifest and report fields are empty or wrong: coherence_score_distribution is {"_count": 0}; privacy.records_examined and fields_examined are 0 with an empty findings list despite 761 privacy discards; composition_drift_pp reports each member's drift as the negative of its requested share; budget.spent_model_calls is 0 and budget.spent_seconds is 2.071 for a run that made on the order of 200,000 model calls; and the single recorded segment spans 4.6 seconds with first_record_index (100000) greater than last_record_index (99929). The corpus itself is internally consistent and its hash verifies, and the discard counts reconcile exactly with generated-minus-written — but a reader should not trust the manifest's timing, budget, coherence or privacy-volume counters, and every number in this card that could be recomputed from the records was recomputed rather than quoted.
Scenario repetition. 98,239 distinct scenario strings cover 99,930 records: 1,002 scenario strings are used more than once, affecting 2,693 records (2.7%). The repeats cluster on a few attractors — "Customer received a Bluetooth speaker instead of the ordered noise-cancelling headphones" and variants appear dozens of times across shipping-wrong-item-received, and duplicate-charge-on-a-smart-speaker dominates billing-duplicate-charge. Exact-duplicate detection found zero duplicate records, so the conversations differ; the situations behind them are narrower than the scenario count suggests.
Single domain, single model, single prompt, single seed, single run. One consumer-electronics prompt, one 20B model, seed 42, one run. Phrasing, product names, escalation language and problem framing will be markedly less varied than real traffic, and stylistic tics of one model run through the entire corpus. The size of the corpus does not mitigate this: 99,930 records from one model are not 99,930 independent samples of anything.
Short conversations only. 4–8 turns by configuration. Real support threads that reopen, span channels, or involve a second agent have no analogue here.
Composition is prescribed, not observed. Every distribution in this card was requested. That 70% of tickets are resolved and 10% are urgent says what the config asked for and nothing about consumer electronics support.
Timestamps are seeded, not causal. created_at and resolved_at are drawn from the run window and a duration range; they do not encode business hours, weekday effects, seasonality, or any relationship between priority and resolution latency. Do not use this corpus for queueing or SLA modelling.
Fabricated products and companies. Product names appearing in conversations are invented. Any resemblance to real products is incidental and not an endorsement or a description.
Privacy coverage is narrower than "scanned" implies. Five entity types, two fields, one regex-based detector. Person names, addresses, postal codes, IBANs and non-US identifiers are outside it by declaration.
Recommendations
Users should be made aware of the risks, biases and limitations of the dataset. Specifically:
- Do not filter or rank on `coherence_score`. Everything in the corpus already passed 0.8, and 85% of it sits on two values from an uncalibrated judge scoring its own work. If you need a quality signal, score the conversations yourself with a different model, or draw a sample and read it.
- Do not treat any distribution as descriptive. Composition, turn counts, subdomain balance and timestamps were all specified before generation. Use them as controlled variables, which is what they are good for, not as observations.
- Deduplicate on `scenario` if scenario diversity matters to your task — dropping to one record per scenario string costs 1,691 records and removes the attractor clusters.
- Do not index on `record_index`. It has 70 gaps.
- Do not rely on the privacy scan for entity types it declares out of scope. If your use requires name- or address-level assurance, run your own detector before use.
- Expect one model's voice. For anything where stylistic diversity matters, regenerate from the repository across several models, prompts and seeds and pool the results rather than resampling this corpus.
- Do not attempt textual reproduction. No sampling seed was set and the tree was dirty; rerun from the manifest to reproduce the structure, and expect different conversations.
Citation [optional]
BibTeX:
@misc{bates_synthetic_support_tickets_2026,
author = {Bates, Robert},
title = {Synthetic Multi-Turn Customer Support Tickets (Consumer Electronics)},
year = {2026},
version = {1.0.0},
howpublished = {\url{https://github.com/arpieb/ticket-dataset-generator}}
}APA:
Bates, R. (2026). Synthetic Multi-Turn Customer Support Tickets (Consumer Electronics) (Version 1.0.0) [Data set]. https://github.com/arpieb/ticket-dataset-generator
Glossary [optional]
- Slot — one unit of generation work. Its metadata is assigned before any model call, which is what makes the composition exact and lets a discarded record be retried without changing the corpus shape.
- Subdomain — a category declared by the domain prompt document (e.g.
billing-invoice-discrepancy). Chosen by seed before dispatch; the model elaborates a specific scenario within it. - Scenario — the specific situation the model invented within its assigned subdomain, stored on the record as free text. Not drawn from a fixed list, which is why it repeats only rarely.
- Exempt by range — a privacy finding whose value comes from a range a standard reserves for fiction. Reported rather than hidden, so the scan never looks cleaner than it was.
- Structural reproduction — same seeded choices at every position, different conversation text. What this pipeline guarantees when no model sampling seed is set.
- Blocking finding — a privacy detection that is not exempt by reserved range. It discards the record; it never appears in the corpus.
More Information [optional]
The manifest (8c56469b-f094-4686-90dc-29715bab4228.manifest.json) and run report (…report.json) for this corpus are the authoritative provenance record: they carry the full resolved configuration, input hashes, code revision, and the complete filter accounting — subject to the counter defects noted in Limitations. ticket-dataset-generator config-from-manifest recovers the exact configuration from the manifest, and ticket-dataset-generator generate --from-manifest reproduces the run, refusing if any recorded input has since changed. Given the dirty code revision, a --from-manifest rerun reproduces the structure under whatever code is checked out, not under the code that produced this corpus.
Dataset Card Authors [optional]
Robert Bates (arpieb), with drafting assistance from Claude Code. Every figure was read from the run manifest, the run report, the domain prompt, the rubric, or recomputed from the corpus itself.
Dataset Card Contact
arpieb via the GitHub repository above.
