datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.Support-Ticket-Router-12K-Cleaned
🔥 Support-Ticket-Router-12K-Cleaned
This dataset is a cleaned and structured version of real-world-like customer support messages designed for intent classification and routing tasks in SaaS / IT support systems.
It is intended for training and evaluating LLM-based or classical NLP intent classifiers for automated customer support ticket routing.
🧪 Data Source
This dataset is synthetically generated using GPT-4-class models (GPT-4 / GPT-4o-style prompting) with… See the full description on the dataset page: https://huggingface.co/datasets/cngchis/Support-Ticket-Router-12K-Cleaned.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.food-delivery-support-tickets
Food Delivery Support Tickets (synthetic)
10,153 synthetic English customer-support conversations for a food delivery
platform (à la Wolt / Uber Eats / DoorDash). Each record is a realistic customer
message with structured labels and a professional agent resolution + reply.
Built for the Food Delivery Support Copilot — an assistant that classifies an
incoming ticket, retrieves similar resolved cases, and drafts a reply.
How it was made
Generated locally with the… See the full description on the dataset page: https://huggingface.co/datasets/OrSabbach/food-delivery-support-tickets.support-json-ru
Support-JSON-RU
Synthetic Russian SaaS support data for policy-conditioned JSON decisions and draft replies. The task supplies customer text, company policies, sourced facts and available capabilities; the model predicts a nine-field decision rather than memorizing a single company's policy.
Русский SaaS-support: обращение + правила + факты → категория, приоритет, настроение, действие, черновик ответа и эскалация.
Model · Dataset files · License
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/A11Sunday/support-json-ru.consumer-electronics-support
Dataset Card for Synthetic Multi-Turn Customer Support Tickets (Consumer Electronics)
99,930 wholly synthetic multi-turn customer-support conversations in a consumer-electronics
retail domain, each with assigned ticket metadata, a machine coherence score, and provenance
linking it to the run that produced it.
Dataset Details
Dataset Description
Every conversation is fabricated by a language model from a committed domain prompt. No real
support… See the full description on the dataset page: https://huggingface.co/datasets/arpieb/consumer-electronics-support.task084_babi_t1_single_supporting_fact_identify_relevant_fact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task084_babi_t1_single_supporting_fact_identify_relevant_fact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task084_babi_t1_single_supporting_fact_identify_relevant_fact.task083_babi_t1_single_supporting_fact_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task083_babi_t1_single_supporting_fact_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task083_babi_t1_single_supporting_fact_answer_generation.mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.IT_Support_V2
Mack: IT Support & Admin Dataset
📋 Dataset Description
This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent.
The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging.
Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.fintech-support-triage
Fintech Support Triage
A multi-turn customer-support dataset with a reward you can compute in code. There is
no LLM judge in the loop.
The setting is Zoomberg Brokerage, a fictional US retail brokerage. At each agent turn
the model writes the reply to the customer and a structured triage action: route or
escalate, which queue, what priority, which flags to raise, and which policies it relied on.
Every action is checked against a written 63-policy pack.
Split
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KartiOS/fintech-support-triage.deep-emotional-support-zh
Deep Emotional Support Dialogue Dataset (Chinese)
深度情感支持对话数据集
Dataset Description
High-quality Chinese emotional support and psychological healing dialogues covering trauma analysis, self-reconstruction, and emotional regulation. Real human-AI interactions, not synthetic.
高质量中文情感支持与心理疗愈对话,涵盖创伤分析、自我重建、情绪调节等深度话题。来源于真实的人机交互,非合成数据。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-emotional-support-zh.Vietnamese-Customer-Support-QA
CSConDa: Customer Support Conversations Dataset for Vietnamese
CSConDa is the first Vietnamese question answering dataset in the customer support domain. It contains over 9,000 QA pairs curated from real interactions between customers and human advisors. The data was collected and approved in collaboration with DooPage, a Vietnamese software company that serves 30,000 customers and 45,000 advisors through its multi-channel support platform.
The questions cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/ura-hcmut/Vietnamese-Customer-Support-QA.food-delivery-support-tickets
Food Delivery Support Tickets (synthetic)
10,153 synthetic English customer-support conversations for a food delivery
platform (à la Wolt / Uber Eats / DoorDash). Each record is a realistic customer
message with structured labels and a professional agent resolution + reply.
Built for the Food Delivery Support Copilot — an assistant that classifies an
incoming ticket, retrieves similar resolved cases, and drafts a reply.
How it was made
Generated locally with the… See the full description on the dataset page: https://huggingface.co/datasets/shaked08/food-delivery-support-tickets.customer-support-th-26.9k
customer-support-th-26.9k
Thai customer-support instruction dataset (~26.9k examples). Thai-localized version of the Bitext customer-support dataset — instruction templates, intent/category labels, and response templates in Thai.
Format
Field
Description
instruction
Customer question template in Thai (may contain {{placeholders}})
response
Support response template in Thai
category
Coarse category (e.g. ORDER)
intent
Fine-grained intent (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/customer-support-th-26.9k.polychart-shown-is-not-supported
Shown Is Not Supported
A chart is a completion claim. It renders cleanly, states a confident finding,
and the underlying data may not support it. A model reading that chart inherits
the gap: it answers from the visual impression because it never extracted the
values.
This adapter reads the data first and corrects the chart when its encoding
misleads. It ships with the first continuous score for how much a chart lies.
Trained with AutoScientist by Adaption for the AutoScientist… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/polychart-shown-is-not-supported.SupportLM-triage-dataset
SupportLM triage v1.0.2
Held-out customer support tickets with five-field triage labels, published so the
SupportLM V1.0.2 adapter's scores are
reproducible rather than self-reported.
Synthetic data. Every ticket was LLM-generated (llm_generated_v6_augmented)
for this project. It has never seen real support traffic, and the synthetic-to-real
gap is unmeasured.
Splits
Split
Rows
Purpose
train
4,000
the rows used for fine-tuning
validation
600… See the full description on the dataset page: https://huggingface.co/datasets/Ompatil19/SupportLM-triage-dataset.MentalHealth-Support
Important Note
This dataset is created from merging two datasets from different sources and has been formatted according to the "messages", "role", "content" chat format. I do not claim any ownership of this dataset.
Keep in mind that this dataset is entirely synthetic. It is not fully representative of real therapy situations. If you are training an LLM therapist keep in mind the limitations of LLMs and highlight those limitations to users in a responsible manner.
Since Mental… See the full description on the dataset page: https://huggingface.co/datasets/ShivomH/MentalHealth-Support.saas-erp-support-ru-train-v13
Capstone support: F training shards
Train-only synthetic Russian SaaS / fictional ERP / telecom scenarios:
430 dialogues, 1113 target turns, 13 local JSONL shards. No real company or customer data.
No validation/test split is published here. Files preserve project-relative paths
under data/quality90_v1/train; restore those paths in a checkout to reuse them.
Each dialogue contains context and turn-level target labels/replies. The exact
file list and SHA-256 hashes are in… See the full description on the dataset page: https://huggingface.co/datasets/AkanaYB/saas-erp-support-ru-train-v13.synthetic-b2b-saas-support-dialogues-sample
Synthetic B2B SaaS Support Dialogues (Sample)
Free sample: 100 dialogues from a larger dataset of 484 synthetic customer support conversations for B2B SaaS products.
What's inside
100 complete dialogues (6–8 messages each)
7 issue categories: auth, billing, integration, data, account, technical, onboarding
Rich metadata: resolution_status, customer_sentiment, agent_actions, escalation_needed
Realistic technical details: error codes, URLs, button names, account… See the full description on the dataset page: https://huggingface.co/datasets/Jurgen1161/synthetic-b2b-saas-support-dialogues-sample.customer-support-chatml
Customer Support ChatML Dataset
This dataset is a curated and preprocessed version of the
Bitext Customer Support Dataset.
Dataset Description
The dataset has been converted to ChatML format for fine-tuning conversational AI models.
Format
Each example contains:
text: The complete conversation in ChatML format
messages: JSON string of the conversation as a list of messages
instruction: The original user query
response: The original assistant response… See the full description on the dataset page: https://huggingface.co/datasets/Shivam271089/customer-support-chatml.datainventor-support-general
DataInventor: General customer support
A synthetic instruction dataset of 200 rows in English, built with DataInventor, an open-source recreation of the Invent a Dataset workflow. Describe the data you want and DataInventor writes it. This is one of eight showcase datasets.
Request
A dataset of customer support inquiries and corresponding responses, covering issues such as product information, technical troubleshooting, billing disputes, and service returns.
The… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/datainventor-support-general.cloudsync-support-sft
CloudSync Pro support demonstrations (SFT)
1,931 chat conversations showing a perfect first-line support agent for a
fictional product: read the customer's message, search a knowledge base, answer
from what came back, and hand over to a human when the conversation belongs to
one.
This is the pile that trained
monte-inc/qwen2.5-1.5b-cloudsync-support
(11.79% → 87.19% on its dev exam, before GRPO took it to 96.07%).
One row
Chat messages plus the tools the agent may… See the full description on the dataset page: https://huggingface.co/datasets/monte-inc/cloudsync-support-sft.it-support-l1-ticket-classification
IT Support L1 Multilingual Dataset
Dataset Summary
IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping.
This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.IT_Support
Mack IT Support Datasets
The Mack dataset is a collection of high-quality IT support data curated for developing and benchmarking agentic language models, digital helpdesk assistants, and troubleshooting bots.It contains seven .jsonl files with diverse coverage:
A_identity.jsonl: Agent identity and persona modeling.
B_troubleshooting.jsonl: Stepwise troubleshooting dialogs and solutions.
C_steps.jsonl: IT procedures and diagnostic workflow data.
D_reasoning.jsonl: Support agent… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support.sea-ecommerce-customer-support-sample
SEA Multilingual E-commerce Customer Support Sample
This public sample contains 1,000 synthetic, AI-generated customer-support
conversations for Southeast Asian e-commerce scenarios.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, prototyping, multilingual testing,
and intent-classification experiments.
Important Limitations
This is synthetic… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-ecommerce-customer-support-sample.customer-support
Description
Topic: Customer Support Interactions
Domains: E-commerce, Telecommunications, Software Services
Number of Entries: 1,000
Dataset Type: Raw Dataset
Model Used: Meta Llama4 Maverick 17B Instruct V1
Language: English
customer-support-dpo-100k
Customer Support DPO 100K
A synthetic Direct Preference Optimization (DPO) dataset of 100,000 customer support interactions with chosen (high-quality) and rejected (poor-quality) response pairs. Designed to train AI models to provide genuinely helpful, specific, and empathetic customer support.
Dataset Description
This dataset covers 23 real-world customer support scenarios across B2B and B2C contexts. Each record includes a customer message, a high-quality chosen… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/customer-support-dpo-100k.datainventor-support-banking
DataInventor: Banking customer support
A synthetic instruction dataset of 200 rows in English, built with DataInventor, an open-source recreation of the Invent a Dataset workflow. Describe the data you want and DataInventor writes it. This is one of eight showcase datasets.
Request
Dataset of customer service responses guiding users through credit card activation, blocking, and mortgage inquiries.
The wording is one of the eight example requests in the Invent a… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/datainventor-support-banking.
