datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.Customer-Support-Responsescontrolled_anchor_v1_support_switch
Controlled ICIL Anchor V1 Support Switch
LeRobot conversion of the Anchor V1 controlled ICIL collection.
Source HDF5:
/ibex/project/c2090/jian/icil_openpi/ICIL/data/manifest_collection_v1/controlled_anchor_v1_60a_6p_3j_6obj_res256_lzf_merged.hdf5
OpenPI sidecars are stored under meta/controlled_icil/.
unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.customer-support-on-twitter-conversationSupportBench
SupportBench
A multilingual benchmark for evaluating case extraction from real-world tech support group chats.
SupportBench contains 60,000 messages across 6 datasets in 3 languages (English, Spanish, Ukrainian), spanning 6 technical domains. All messages are sourced from public Telegram support groups.
Datasets
Dataset
Language
Domain
Messages
Users
Reply%
Media
Ardupilot-UA
Ukrainian
UAV / Drones
10,000
319
51.8%
1,440
MikroTik-UA
Ukrainian
Networking
10… See the full description on the dataset page: https://huggingface.co/datasets/pavelshpagin/SupportBench.Customer_Support_on_Twitterit-support-tickets
Classification of IT Support Tickets
This dataset contains real support tickets collected from an IT support
company in the Florianópolis region of Brazil. It contains 2,229 ticket texts,
manually classified into seven categories by three IT support professionals.
The source authors' train/test split is preserved exactly; Tasksource joins
each text and label by the original ID without reshuffling or editing text.
The tickets are mainly in English, German, Portuguese, and Spanish… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/it-support-tickets.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.Support-Ticket-Router-12K-Cleaned
🔥 Support-Ticket-Router-12K-Cleaned
This dataset is a cleaned and structured version of real-world-like customer support messages designed for intent classification and routing tasks in SaaS / IT support systems.
It is intended for training and evaluating LLM-based or classical NLP intent classifiers for automated customer support ticket routing.
🧪 Data Source
This dataset is synthetically generated using GPT-4-class models (GPT-4 / GPT-4o-style prompting) with… See the full description on the dataset page: https://huggingface.co/datasets/cngchis/Support-Ticket-Router-12K-Cleaned.pydreg-supporting-data
pydreg vs. dREG benchmark outputs
Raw benchmark artifacts backing the performance and accuracy comparisons in
pydreg, a from-scratch Python port of
dREG (Danko Lab). This dataset holds the
paired outputs of running both tools' full peak-calling pipeline
(run_dREG/pydreg) on the same 12 real PRO-seq/GRO-seq/ChRO-seq libraries,
plus the /usr/bin/time -v logs used to compare wall-clock time and peak
memory. It is data, not code — see the pydreg repo for the package itself and
for… See the full description on the dataset page: https://huggingface.co/datasets/adamyhe/pydreg-supporting-data.support-ticket-dataset
Support Tickets with an AI-Automation Counterfactual
1.5M synthetic support tickets in two matched worlds — one where an AI assistant handles part
of the queue, one staffed entirely by humans. Same tickets, same customers, same attributes.
Only the routing differs.
Most support datasets give you one world and leave you guessing about the other. This one gives
you both, so questions like "what would this have cost without automation?" are measured
rather than estimated.
from… See the full description on the dataset page: https://huggingface.co/datasets/s2pidape/support-ticket-dataset.lingrow-support-tickets
Lingrow Support Tickets (Synthetic)
A synthetic dataset of 10,000 customer-support tickets for Lingrow,
a real-time multilingual translation and communication platform. Each ticket
contains a customer message (an error report or a how-to question), rich
metadata, and a resolution. The data is fully synthetic — no real customer
information is included.
This dataset was built as the final project for a Data Science course. It powers
the Lingrow Support Copilot: a tool that, given… See the full description on the dataset page: https://huggingface.co/datasets/adiprog14/lingrow-support-tickets.synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.food-delivery-support-tickets
Food Delivery Support Tickets (synthetic)
10,153 synthetic English customer-support conversations for a food delivery
platform (à la Wolt / Uber Eats / DoorDash). Each record is a realistic customer
message with structured labels and a professional agent resolution + reply.
Built for the Food Delivery Support Copilot — an assistant that classifies an
incoming ticket, retrieves similar resolved cases, and drafts a reply.
How it was made
Generated locally with the… See the full description on the dataset page: https://huggingface.co/datasets/OrSabbach/food-delivery-support-tickets.mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.consumer-electronics-support
Dataset Card for Synthetic Multi-Turn Customer Support Tickets (Consumer Electronics)
99,930 wholly synthetic multi-turn customer-support conversations in a consumer-electronics
retail domain, each with assigned ticket metadata, a machine coherence score, and provenance
linking it to the run that produced it.
Dataset Details
Dataset Description
Every conversation is fabricated by a language model from a committed domain prompt. No real
support… See the full description on the dataset page: https://huggingface.co/datasets/arpieb/consumer-electronics-support.customer_support_conversations_dataset
💬 Customer Support Conversation Dataset — Powered by Syncora.ai
A free synthetic dataset for chatbot training, LLM fine-tuning, and synthetic data generation research.Created using Syncora.ai’s privacy-safe synthetic data engine, this dataset is ideal for developing, testing, and benchmarking AI customer support systems.
It serves as a dataset for chatbot training and a dataset for LLM training, offering rich, structured conversation data for real-world simulation.
🌟… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/customer_support_conversations_dataset.e-commerce-customer-support-qa
Dataset Card for Dataset Name
from: NebulaByte/E-Commerce_Customer_Support_Conversations
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/rjac/e-commerce-customer-support-qa.controlled_anchor_v0_support_switch
Controlled ICIL Anchor v0 Support Switch
LeRobot conversion of the controlled ICIL anchor-v0 atomic demonstrations, with
the challenge support-switch pair manifests attached under
meta/controlled_icil/pair_manifests/anchor_v0_challenge_split.
Counts
Episodes: 2379
Frames: 236436
Skipped source episodes: 0
Train anchors: 165
Test anchors: 20
Excluded incomplete anchors: 15
Train support-target pairs: 1980
Test support-target pairs: 240
OpenPI paths… See the full description on the dataset page: https://huggingface.co/datasets/daixianjie/controlled_anchor_v0_support_switch.support-json-ru
Support-JSON-RU
Synthetic Russian SaaS support data for policy-conditioned JSON decisions and draft replies. The task supplies customer text, company policies, sourced facts and available capabilities; the model predicts a nine-field decision rather than memorizing a single company's policy.
Русский SaaS-support: обращение + правила + факты → категория, приоритет, настроение, действие, черновик ответа и эскалация.
Model · Dataset files · License
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/A11Sunday/support-json-ru.lfqa_support_docsSupport documents for building https://huggingface.co/vblagoje/bart_lfqa model
Bitext-customer-support-llm-chatbot-training-dataset-spanish
Spanish Customer Support LLM Chatbot Training Dataset
Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset.
This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models.
Dataset Details
Dataset Description
This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.insuff_supported_argumentstask084_babi_t1_single_supporting_fact_identify_relevant_fact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task084_babi_t1_single_supporting_fact_identify_relevant_fact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task084_babi_t1_single_supporting_fact_identify_relevant_fact.customer-support-client-agent-conversations
Customer Support Client-Agent Conversations Dataset
A synthetic context-summarized multi-turn customer-service question-answering dataset for banking domain conversations, designed for training and evaluating small language models on dialogue continuity and contextual understanding tasks.
Dataset Description
This dataset contains 183,337 context-summarized multi-turn customer-service conversations spanning various banking scenarios including account management… See the full description on the dataset page: https://huggingface.co/datasets/Lakshan2003/customer-support-client-agent-conversations.E-Commerce_Customer_Support_Conversations
Dataset Card for "E-Commerce_Customer_Support_Conversations"
The dataset is synthetically generated with OpenAI ChatGPT model (gpt-3.5-turbo).
More Information needed
task083_babi_t1_single_supporting_fact_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task083_babi_t1_single_supporting_fact_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task083_babi_t1_single_supporting_fact_answer_generation.
