datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sms_spam
Dataset Card for [Dataset Name]
Dataset Summary
The SMS Spam Collection v.1 is a public set of SMS labeled messages that have been collected for mobile phone spam research.
It has one collection composed by 5,574 English, real and non-enconded messages, tagged according being legitimate (ham) or spam.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
Data Instances
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/ucirvine/sms_spam.verifiable-code-reasoning
Verifiable Code Reasoning
Execution-verified Python problems with chain-of-thought
Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text
Overview
Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests.
Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if:
a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.sms-spam-collection
SMS Spam Collection v.1
DESCRIPTION
The SMS Spam Collection v.1 (hereafter the corpus) is a set of SMS tagged messages that have been collected for SMS Spam research. It contains one set of SMS messages in English of 5,574 messages, tagged acording being ham (legitimate) or spam.
1.1. Compilation
This corpus has been collected from free or free for research sources at the Web:
A collection of between 425 SMS spam messages extracted manually from the Grumbletext Web… See the full description on the dataset page: https://huggingface.co/datasets/codesignal/sms-spam-collection.execution-verified-codework
Execution-Verified CodeWork
Only code that passes the tests ships
Sandbox-executed · ≥6 unit tests · implement / repair / harden · instance-deduplicated
One-sentence pitch
Training traces for writing, fixing, and hardening Python functions — every kept solution was actually run against unit tests and passed.
What you get
Field
Role
kind
implement · repair · harden
problem
Clear developer task
reasoning
Numbered… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/execution-verified-codework.RIFA-Artbook-datasetSMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset
Collection of Multilingual SMS messages tagged as spam or legitimate
About Dataset
Context
The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French.
The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dbarbedillo/SMS_Spam_Multilingual_Collection_Dataset.india-spam-smssms-otp-spam-dataset
📲 SMS OTP Spam Dataset
A synthetic dataset of 10,000 OTP-style SMS messages for spam classification tasks. The dataset includes both valid and spam-like messages, with labels for message validity and delivery status.
📊 Dataset Summary
Total samples: 10,000
Valid messages: 90%
Not valid (spam-like): 10%
Status types: delivered, failed, spam, bounced, expired
Each entry includes:
phone_id: Synthetic phone number
sms_text: Message content
label: valid or not valid… See the full description on the dataset page: https://huggingface.co/datasets/alusci/sms-otp-spam-dataset.Telecommunication_SMS_time_series
SMS Time series data for traffic and fraud forecasting. TeleWhale vendor collected data
This dataset contains various time series from vendors.
Shashkov A.A.
Vendor A: 01.03.23-14.08.23
TS_*_all - Count of all SMS
Vendor A: January
TS_*_fraud - Count of fraud
TS_*_all - Count of all SMS
TS_*_hlrDelay - Mean values of hlr delay
Vendor B: January 1-8
1-8_TS_*_fraud - Count of fraud
1-8_TS_*_all - Count of all SMS
1-8_TS_*_hlrDelay -… See the full description on the dataset page: https://huggingface.co/datasets/Shashkovich/Telecommunication_SMS_time_series.SMS-spamThis dataset can be found on Kaggle, Huggingface and many other websites, but the source is from the research paper [Tiago], whose authors contributed it to
the Machine Learning repository at [UCI]. It contains 5,574 SMS messages, of which 747 massages are labeled as spam.
By nature, the messages are short and, in some cases, quite cryptic and personal.
The CSV file is a straightforward representation of the data.
References
[UCI] https://archive.ics.uci.edu/dataset/228/sms+spam+collection… See the full description on the dataset page: https://huggingface.co/datasets/bvk/SMS-spam.sms-spam-ham-dataset
Dataset
The dataset is composed of messages labeled by ham or spam, merged from three data sources:
SMS Spam Collection https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset
Telegram Spam Ham https://huggingface.co/datasets/thehamkercat/telegram-spam-ham/tree/main
Enron Spam: https://huggingface.co/datasets/SetFit/enron_spam/tree/main (only used message column and labels)
The necessary preprocessing of the datasets and the project is located in my github. You can… See the full description on the dataset page: https://huggingface.co/datasets/subashdhamee/sms-spam-ham-dataset.Spam_SMS
Description
The Spam SMS is a set of SMS-tagged messages that have been collected for SMS Spam research. It contains one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam.
Source: uciml/sms-spam-collection-dataset
sms-spam-classificationsms_spam_collectionspam-sms-collection-01
📦 SMS Spam Detection Dataset
A curated dataset of SMS messages labeled as Spam or Ham (Not Spam).This dataset is ideal for building and testing spam detection models using Machine Learning or Deep Learning.
🧠 Overview
File Name: spam.csv
Total Entries: 5,000+ SMS messages
Format: CSV (Comma Separated Values)
Columns:
label → Indicates whether the message is spam or ham
message → The SMS text content
📊 Dataset Features
Feature… See the full description on the dataset page: https://huggingface.co/datasets/DarkNeuron-AI/spam-sms-collection-01.bengali-sms-smishing-dataset
Bengali SMS Smishing Dataset
A multilingual SMS dataset for phishing (smishing) detection, covering Bengali, English, Banglish, and Code-Mixed linguistic varieties. Developed as part of the SmishDetect-LLM research framework.
Hugging Face: shariul-islam/bengali-sms-smishing-dataset
Dataset Summary
This dataset contains 7,005 SMS messages annotated across three classification categories and four linguistic varieties. It is the first publicly available smishing… See the full description on the dataset page: https://huggingface.co/datasets/shariul-islam/bengali-sms-smishing-dataset.sms_spam_categoryindian-scam-sms-synthetic-audited
Indian Scam SMS (synthetic, audited)
1,580 short messages that imitate SMS and WhatsApp scams and their genuine look-alikes in Indian
English, Hindi (Devanagari), Hinglish and four Roman-script code-mixed styles (Tamil, Telugu, Bengali,
Marathi with English). Every row was written by a large language model and then audited for label noise.
It exists to train and stress-test scam detectors on the hard negatives that public datasets lack:
real-looking bank, courier, bill and job… See the full description on the dataset page: https://huggingface.co/datasets/Ridham115/indian-scam-sms-synthetic-audited.SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset
Collection of Multilingual SMS messages tagged as spam or legitimate
About Dataset
Context
The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French.
The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/KumarSahil299885/SMS_Spam_Multilingual_Collection_Dataset.vietnamese_sms_dataset
Bộ dữ liệu SMS lừa đảo tiếng Việt được đảm bảo chất lượng (Official Release)
(English Below)
Chào mừng bạn đến với kho lưu trữ chính thức của Bộ dữ liệu SMS lừa đảo tiếng Việt được đảm bảo chất lượng.
Đây là một bộ dữ liệu được xây dựng nhằm phục vụ nghiên cứu trong các lĩnh vực an ninh mạng, xử lý ngôn ngữ tự nhiên (NLP) và học máy, với trọng tâm là bài toán phát hiện tin nhắn SMS rác/lừa đảo.
Bộ dữ liệu này được tổng hợp từ các tin nhắn SMS thực tế trong cuộc sống. Không… See the full description on the dataset page: https://huggingface.co/datasets/trannguyenthaituan/vietnamese_sms_dataset.sm-stuffmalicious-benign-sms-mms-dataset
Dataset v3 Changelog
Changes from dataset v2 (model_datasets/v4/) to v3 (model_datasets/v2-4/).
Summary
v3 is a curated, rebalanced, and feature-enriched derivative of v2. The goal was to improve training signal quality by removing noisy examples, correcting mislabelled data, fixing the short-message class imbalance, and adding 23 engineered text features.
v2
v3 (base)
v3 (DeBERTa)
File
dataset_v4_dual_cleaned_v2.csv
dataset_v2.4.csv
dataset_v2-4_deberta.csv… See the full description on the dataset page: https://huggingface.co/datasets/notd5a/malicious-benign-sms-mms-dataset.sms-spam-enriched
SMS Spam Enriched Dataset
An enriched version of the classic SMS Spam Collection Dataset from UC Irvine with additional engineered features and semantic embeddings.This dataset is designed for spam detection, feature engineering experiments, and model interpretability research.
Dataset Overview
Total samples: 5,171
Classes:
0: Ham (non-spam)
1: Spam
Enrichments Added
Alongside the raw SMS text (sms) and labels (label), we engineered multiple new… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/sms-spam-enriched.sms-spam
SMS Spam
The SMS Spam Collection, deduplicated.
The original corpus, assembled by Almeida and Gómez Hidalgo in 2011, contains 5,574 SMS messages tagged as ham (legitimate) or spam. About seven percent of the rows are exact duplicates: the same message text appearing more than once. This release removes them and repairs a small set of encoding artifacts. The messages and labels are otherwise unchanged.
5,159 messages remain: 4,517 ham, 642 spam.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/sms-spam.kinshield-sms-20261005
Kinshield SMS — 20261005
Artifacts SFT synthetic từ campaign Sonnet của Kinshield, cập nhật verdict/key findings và evidence có scope rõ ràng. Giữ SMS hiện có; không sinh lại toàn bộ corpus.
Split
Requests
train
73,917
dev
997
test
767
Total
75,681
So với bản 20260917: 78,126 →75,681 unique rows, giảm2,445 sau selection/QA gates. Bản cũ77,358 train +384 test_ood_scenario +384 test_ood_entity; bản mới73,917 train +997 dev +767 test. Dev được tách từ old… See the full description on the dataset page: https://huggingface.co/datasets/jacky-qualgo/kinshield-sms-20261005.sms-spam-combinedsmsaSmSA is a sentence-level sentiment analysis dataset (Purwarianti and Crisdayanti, 2019) is a collection of comments and reviews
in Indonesian obtained from multiple online platforms. The text was crawled and then annotated by several Indonesian linguists
to construct this dataset. There are three possible sentiments on the SmSA dataset: positive, negative, and neutralsms-spam-collectionnus_sms_corpussms-spam-balanced
SMS Spam Collection (Balanced)
A balanced SMS spam dataset for text classification.
Overview
Property
Value
Total Samples
1,494
Train
1,045 (70%)
Validation
149 (10%)
Test
300 (20%)
Classes
ham (0), spam (1)
Dataset Description
This is a balanced version of the UCI SMS Spam Collection dataset. Originally, the dataset had 5,572 messages with an imbalanced distribution (4,825 ham, 747 spam). We balanced it to 747 ham and 747 spam for… See the full description on the dataset page: https://huggingface.co/datasets/Mustafaege/sms-spam-balanced.
