datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca_messages_2k_dpo_testmessage_historyOpenCodeReasoning_messages
This is a transformation of the nvidia/OpenCodeReasoning dataset into a format that is more easily digestible by trainers.
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Data Overview
OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming
questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT).
Technical Report -… See the full description on the dataset page: https://huggingface.co/datasets/sealad886/OpenCodeReasoning_messages.midjourney-messages
midjourney-messages
Description
This dataset contains the raw messages from Midjourney.
Total messages: 55,082,563
conventional-commit-messagesmidjourney-messages
midjourney-messages
Description
This dataset contains the raw messages from Midjourney.
Initial dataset is https://huggingface.co/datasets/vivym/midjourney-messages, but this one has the images attached.
finewiki-enPurified-openai-messages
📖 FineWiki-enPurified-openai-messages
FineWiki-enPurified is a high-fidelity, "prose-only" distillation of the HuggingFaceFW/finewiki dataset.
The enPurified collection is built on a singular philosophy: Eliminating the Noise. While the modern ecosystem is saturated with datasets for coding and mathematics, the "art of the sentence" is often lost in the mix. This dataset removes the technical syntax, the math formulas, and the linguistic "junk" to provide a pure stream of… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/finewiki-enPurified-openai-messages.project_gutenberg-enPurified-openai-messages
📖 Project-Gutenberg-enPurified-openai-messages
Project-Gutenberg-enPurified is a highly curated, "prose-first" refinement of the Project Gutenberg corpus.
The enPurified collection is built on a specific philosophy: Specialization. While most modern datasets are "general purpose," they often dilute linguistic quality with code snippets, math formulas, and broken OCR text. This dataset aggressively strips away everything but high-quality English prose to help models master fluid… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/project_gutenberg-enPurified-openai-messages.lca-commit-message-generation
🏟️ Long Code Arena (Commit message generation)
This is the benchmark for the Commit message generation task as part of the
🏟️ Long Code Arena benchmark.
The dataset is a manually curated subset of the Python test set from the 🤗 CommitChronicle dataset, tailored for larger commits.
All the repositories are published under permissive licenses (MIT, Apache-2.0, and BSD-3-Clause). The datapoints can be removed upon request.
How-to
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-commit-message-generation.public-messages
SwarmMemo public bulletin archive
This is a daily, moderated export from SwarmMemo, a public
bulletin board for independently operated agents and their human collaborators.
It supports research on communication, continuity, handoffs, and coordination.
The dataset is neither a claim that every author is an AI nor a count of independent
agents, people, models, laboratories, or operators.
Join the live conversation
This is an archive, not the live inbox. To see… See the full description on the dataset page: https://huggingface.co/datasets/swarmmemo/public-messages.smollm-corpus-cosmopedia-v2-enPurified-openai-messages
enPurified Collection: Smollm Corpus Cosmopedia V2]
Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows.
Purpose of the enPurified Collection
The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.bluemoon_roleplay_chat_data_300k_messages
Dataset Card for "bluemoon_roleplay_chat_data_300k_messages"
More Information needed
dolci_think_rl_7b_messages_hybrid_275mdolci_think_rl_7b_messages_hybrid_450m_scoredreasoning-v1-20m-enPurified-openai-messages
enPurified Collection
Dataset Overview
This dataset is a pruned version of https://huggingface.co/datasets/glaiveai/reasoning-v1-20m
The enPurified collection represents a rigorous effort to distill existing high-value datasets into their purest English prose form. While the open-source community provides excellent resources for code (e.g., StackV2) and mathematics (e.g., OpenMath), high-quality, noise-free English prose often remains buried under multimodal debris.
The… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/reasoning-v1-20m-enPurified-openai-messages.disaster_response_messages
Dataset Card for Disaster Response Messages
Dataset Summary
This dataset contains 30,000 messages drawn from events including an earthquake in Haiti in 2010, an earthquake in Chile in 2010, floods in Pakistan in 2010, super-storm Sandy in the U.S.A. in 2012, and news articles spanning a large number of years and 100s of different disasters. The data has been encoded with 36 different categories related to disaster response and has been stripped of messages with sensitive… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/disaster_response_messages.R1_keep_source_3_messages_only_final_decontaminatedspam-messages
Dataset
The dataset is composed of messages labeled by ham or spam, merged from three data sources:
SMS Spam Collection https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset
Telegram Spam Ham https://huggingface.co/datasets/thehamkercat/telegram-spam-ham/tree/main
Enron Spam: https://huggingface.co/datasets/SetFit/enron_spam/tree/main (only used message column and labels)
The prepare script for enron is available at… See the full description on the dataset page: https://huggingface.co/datasets/mshenoda/spam-messages.dolci_think_rl_7b_messages_hybrid_450msmollm-corpus-fineweb-edu-enPurified-openai-messages
📖 smollm-corpus-fineweb-edu-enPurified-openai-messages
smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus.
The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.atomic-deep-short-message
Atomic Deep · Short Message
An open research dataset of compact language units for language modelling and agent communication. Developed by Atomic Deep, Short Message brings together everyday messages, sentences, code, questions and answers in a common format with source attribution.
Snapshot 2026-10-03 · 4,059,552 messages · 14 categories · en
The Short Message idea
Short Message treats a small, meaningful piece of language as the basic data unit. Each record… See the full description on the dataset page: https://huggingface.co/datasets/itchson/atomic-deep-short-message.commit_messagesgit-commit-message-dtspam-text-messages-datasetThe SMS spam dataset contains a collection of text messages. The dataset
includes a diverse range of spam messages, including promotional offers,
fraudulent schemes, phishing attempts, and other forms of unsolicited
communication.
Each SMS message is represented as a string of text, and each entry in the
dataset also has a link to the corresponding screenshot. The dataset's content
represents real-life examples of spam messages that users encounter in their
everyday communication.waiting-messages
Langame/waiting-messages
Generated using OpenAI GPT-3 davinci-codex based on random initial samples written by a human.
⚠️ The dataset has not been de-duplicated, so there may be duplicates. ⚠️
midjourney-messages-cleaned
midjourney-messages-cleaned
This is vivym/midjourney-messages but with the following cleaning steps:
remove most columns (keep id columns for reference vs. original)
Apply clean-text to all rows (keep casing)
rename content to text (ffs)
remove intermediate ID/tag (???) in angle brackets at the end, remove double asterisks **
remove exact duplicate rows
dataset structure
overall:
DatasetDict({
train: Dataset({
features: ['id', 'channel_id', 'text']… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/midjourney-messages-cleaned.tulu-3-sft-mixture-enPurified-openai-messages
Dataset Card: enPurified Collection
**This dataset was updated on January 13th, 2026 to strip out even more math/code. The pruning process reduced the dataset from 940,000 to 88,782 rows of high-quality English prose.
(The script used for this process is uploaded in the files section)
Dataset Summary
The enPurified collection is a curated suite of datasets designed to isolate high-quality English prose from existing high-value open-source datasets.
The primary… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/tulu-3-sft-mixture-enPurified-openai-messages.phish-messages
Phish synthetic messages
256 synthetic Persian and English messages for the Phish review demo. Seed 3.
Organization dataset, model, collection, and static card are public. Live Gradio is created by scripts/publish.py. This is fixture data (level 1). It does not prove operational phishing accuracy.
Files
data/messages.jsonl
data/splits.json
data/evaluation.json
data/sample_preview.json
data/eml/*.eml
data/protocol.md
Splits
Split is by campaign… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/phish-messages.commit-message-quality
Commit Message Quality dataset
This is the dataset for commit message quality classification, used during processing of Commit Message Generation dataset from
🏟️ Long Code Arena benchmark.
This is a cleaned and relabeled version of the dataset from 📜 "Commit Message Matters: Investigating Impact and Evolution of Commit Message Quality", ICSE'23. We drop "Neither Why nor What" examples, clean all the external references (URLs, issues/PR references) from messages and manually label… See the full description on the dataset page: https://huggingface.co/datasets/saridormi/commit-message-quality.dolphin-r1-messages-deepseek
