adiprog14/lingrow-support-tickets
Lingrow Support Tickets (Synthetic) A synthetic dataset of 10,000 customer-support tickets for Lingrow, a real-time multilingual translation and communication platform. Each ticket contains a customer message (an error report or a how-to question), rich metadata, and a resolution. The data is fully synthetic — no real customer information is included. This dataset was built as the final project for a Data Science course. It powers the Lingrow Support Copilot: a tool that, given… See the full description on the dataset page: https://huggingface.co/datasets/adiprog14/lingrow-support-tickets.
Lingrow Support Tickets (Synthetic)
A synthetic dataset of 10,000 customer-support tickets for Lingrow, a real-time multilingual translation and communication platform. Each ticket contains a customer message (an error report or a how-to question), rich metadata, and a resolution. The data is fully synthetic — no real customer information is included.
This dataset was built as the final project for a Data Science course. It powers the Lingrow Support Copilot: a tool that, given a new customer message, retrieves similar past tickets and drafts a reply.
How it was created
- Schema was derived from the official Lingrow user guide (error flows, roles, session types, audio, translation, user management).
- A synthetic generator combined these building blocks with a hand-written phrasing bank to produce 10,000 varied tickets.
- Text variety was improved with a Hugging Face model (
humarin/chatgpt_paraphraser_on_T5_base): the original 116 base messages were paraphrased into 666 unique customer messages, lowering the duplicate-text rate from ~96.5% to ~80%.
Dataset structure
Exploratory Data Analysis (EDA)
- Total tickets: 10,000
- Error reports: 7,979 | How-to questions: 2,021
- Categories: 7 error categories + how-to
- Languages: 10 (balanced)
- Devices: iOS / Android / Web (balanced)
- Overall resolution rate: 88.2%
- Unique customer messages: 666
Categories and error codes are evenly distributed (a fair, balanced synthetic dataset). How-to questions are always marked resolved; error tickets are resolved ~85% of the time.
Intended use
- Information retrieval / semantic search over support tickets
- Retrieval-augmented reply drafting (RAG)
- Educational / demonstration purposes
Limitations & ethics
- Synthetic data: messages are machine-generated and may occasionally read unnaturally. They do not represent real users.
- No personal data: contains no real customer details, transcripts, credentials, or tokens.
- Error codes and flows are inspired by the Lingrow product but simplified.
License
MIT
