Team Ai
Datasetpublic

adiprog14/lingrow-support-tickets

Lingrow Support Tickets (Synthetic) A synthetic dataset of 10,000 customer-support tickets for Lingrow, a real-time multilingual translation and communication platform. Each ticket contains a customer message (an error report or a how-to question), rich metadata, and a resolution. The data is fully synthetic — no real customer information is included. This dataset was built as the final project for a Data Science course. It powers the Lingrow Support Copilot: a tool that, given… See the full description on the dataset page: https://huggingface.co/datasets/adiprog14/lingrow-support-tickets.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes239downloads
Dataset Card

Lingrow Support Tickets (Synthetic)

A synthetic dataset of 10,000 customer-support tickets for Lingrow, a real-time multilingual translation and communication platform. Each ticket contains a customer message (an error report or a how-to question), rich metadata, and a resolution. The data is fully synthetic — no real customer information is included.

This dataset was built as the final project for a Data Science course. It powers the Lingrow Support Copilot: a tool that, given a new customer message, retrieves similar past tickets and drafts a reply.

How it was created

  1. 1.Schema was derived from the official Lingrow user guide (error flows, roles, session types, audio, translation, user management).
  2. 2.A synthetic generator combined these building blocks with a hand-written phrasing bank to produce 10,000 varied tickets.
  3. 3.Text variety was improved with a Hugging Face model (humarin/chatgpt_paraphraser_on_T5_base): the original 116 base messages were paraphrased into 666 unique customer messages, lowering the duplicate-text rate from ~96.5% to ~80%.

Dataset structure

ColumnDescription
ticket_idUnique ticket identifier (e.g. LG-100000)
created_atTimestamp of the ticket
user_roleadmin / member / guest / audience
session_typequickconversation / groupconversation / live_captions
speaking_stylelive / pushtotalk
context_domaingeneral, medical, academic, technical, legal, sport, logistics, industrial, custom
deviceios / android / web
languageCustomer's language (10 options)
error_codeThe specific error (empty for how-to questions)
topic_keyThe topic key (error code or HOWTO_*)
error_categoryOne of 6 error categories, or how_to
customer_messageThe customer's message
resolution_stepsThe support answer
resolvedWhether the ticket was resolved

Exploratory Data Analysis (EDA)

  • —Total tickets: 10,000
  • —Error reports: 7,979 | How-to questions: 2,021
  • —Categories: 7 error categories + how-to
  • —Languages: 10 (balanced)
  • —Devices: iOS / Android / Web (balanced)
  • —Overall resolution rate: 88.2%
  • —Unique customer messages: 666

Categories and error codes are evenly distributed (a fair, balanced synthetic dataset). How-to questions are always marked resolved; error tickets are resolved ~85% of the time.

Intended use

  • —Information retrieval / semantic search over support tickets
  • —Retrieval-augmented reply drafting (RAG)
  • —Educational / demonstration purposes

Limitations & ethics

  • —Synthetic data: messages are machine-generated and may occasionally read unnaturally. They do not represent real users.
  • —No personal data: contains no real customer details, transcripts, credentials, or tokens.
  • —Error codes and flows are inspired by the Lingrow product but simplified.

License

MIT