datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
email-spam-classification
Email Spam Classification
The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems.
The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.Email_Intent_Classification
Dataset Information
This is a dataset of English sentences used in emails with six basic categories: request, informational, transaction, feedback.
An example looks as follows: {"Email": "Your subscription renewal is confirmed. Thank you for staying with us!", "Intent": "Transaction"}
Dataset Sources
Instances generated and annotated by ChatGPT 4.
Uses
Demo for email intent classification tasks.
haldonmez_spam-or-ham-a-dataset-for-email-classification
Spam Email Classification Dataset
Mirror of the Kaggle dataset haldonmez/spam-or-ham-a-dataset-for-email-classification by Halil Dönmez, released under CC0: Public Domain. All credit goes to the original author; please cite and link the Kaggle page when using this data.
_ A dataset containing each of the 6 cleaned versions of the spam mail set._
Original description (from Kaggle)
This dataset is a comprehensive collection of email data, categorized into ‘spam’… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/haldonmez_spam-or-ham-a-dataset-for-email-classification.german-english-email-ticket-classification
Customer Support Tickets (Short Version)
This dataset is a simplified version of the Customer Support Tickets dataset.
Dataset Details:
The dataset includes combinations of the following columns:
type
queue
priority
language
Modifications:
Shortened Version: This version only includes the first three rows for each combination of the above columns (i.e., 'type', 'queue', 'priority', 'language').
This reduction makes the dataset smaller and more manageable… See the full description on the dataset page: https://huggingface.co/datasets/ale-dp/german-english-email-ticket-classification.email-intent-classification
