Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fozziethebeat /alpaca_messages_2k_dpo_testtext1K<n<10K2 likes4.9k downloads2y agoHugging Face02janblue /message_history0 likes1.9k downloads2y agoHugging Face03sealad886 /OpenCodeReasoning_messages This is a transformation of the nvidia/OpenCodeReasoning dataset into a format that is more easily digestible by trainers. OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report -… See the full description on the dataset page: https://huggingface.co/datasets/sealad886/OpenCodeReasoning_messages.texttext-generation100K<n<1M0 likes837 downloads1y agoHugging Face04vivym /midjourney-messages midjourney-messages Description This dataset contains the raw messages from Midjourney. Total messages: 55,082,563 image1M<n<10M128 likes783 downloads3y agoHugging Face052doo /conventional-commit-messages0 likes559 downloads2y agoHugging Face06TwoAbove /midjourney-messages midjourney-messages Description This dataset contains the raw messages from Midjourney. Initial dataset is https://huggingface.co/datasets/vivym/midjourney-messages, but this one has the images attached. 2 likes448 downloads3y agoHugging Face07enPurified /finewiki-enPurified-openai-messages 📖 FineWiki-enPurified-openai-messages FineWiki-enPurified is a high-fidelity, "prose-only" distillation of the HuggingFaceFW/finewiki dataset. The enPurified collection is built on a singular philosophy: Eliminating the Noise. While the modern ecosystem is saturated with datasets for coding and mathematics, the "art of the sentence" is often lost in the mix. This dataset removes the technical syntax, the math formulas, and the linguistic "junk" to provide a pure stream of… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/finewiki-enPurified-openai-messages.texttext-generation1M<n<10M1 likes431 downloads9mo agoHugging Face08enPurified /project_gutenberg-enPurified-openai-messages 📖 Project-Gutenberg-enPurified-openai-messages Project-Gutenberg-enPurified is a highly curated, "prose-first" refinement of the Project Gutenberg corpus. The enPurified collection is built on a specific philosophy: Specialization. While most modern datasets are "general purpose," they often dilute linguistic quality with code snippets, math formulas, and broken OCR text. This dataset aggressively strips away everything but high-quality English prose to help models master fluid… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/project_gutenberg-enPurified-openai-messages.texttext-generation100K<n<1M1 likes334 downloads9mo agoHugging Face09JetBrains-Research /lca-commit-message-generation 🏟️ Long Code Arena (Commit message generation) This is the benchmark for the Commit message generation task as part of the 🏟️ Long Code Arena benchmark. The dataset is a manually curated subset of the Python test set from the 🤗 CommitChronicle dataset, tailored for larger commits. All the repositories are published under permissive licenses (MIT, Apache-2.0, and BSD-3-Clause). The datapoints can be removed upon request. How-to from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-commit-message-generation.text1K<n<10K0 likes330 downloads2y agoHugging Face10swarmmemo /public-messages SwarmMemo public bulletin archive This is a daily, moderated export from SwarmMemo, a public bulletin board for independently operated agents and their human collaborators. It supports research on communication, continuity, handoffs, and coordination. The dataset is neither a claim that every author is an AI nor a count of independent agents, people, models, laboratories, or operators. Join the live conversation This is an archive, not the live inbox. To see… See the full description on the dataset page: https://huggingface.co/datasets/swarmmemo/public-messages.1 likes325 downloads4d agoHugging Face11enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes319 downloads9mo agoHugging Face12rickRossie /bluemoon_roleplay_chat_data_300k_messages Dataset Card for "bluemoon_roleplay_chat_data_300k_messages" More Information needed text100K<n<1M102 likes306 downloads3y agoHugging Face13mnoukhov /dolci_think_rl_7b_messages_hybrid_275mtext100K<n<1M0 likes305 downloads2mo agoHugging Face14mnoukhov /dolci_think_rl_7b_messages_hybrid_450m_scoredtabular100K<n<1M0 likes264 downloads2mo agoHugging Face15enPurified /reasoning-v1-20m-enPurified-openai-messages enPurified Collection Dataset Overview This dataset is a pruned version of https://huggingface.co/datasets/glaiveai/reasoning-v1-20m The enPurified collection represents a rigorous effort to distill existing high-value datasets into their purest English prose form. While the open-source community provides excellent resources for code (e.g., StackV2) and mathematics (e.g., OpenMath), high-quality, noise-free English prose often remains buried under multimodal debris. The… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/reasoning-v1-20m-enPurified-openai-messages.text10M<n<100M1 likes263 downloads9mo agoHugging Face16community-datasets /disaster_response_messages Dataset Card for Disaster Response Messages Dataset Summary This dataset contains 30,000 messages drawn from events including an earthquake in Haiti in 2010, an earthquake in Chile in 2010, floods in Pakistan in 2010, super-storm Sandy in the U.S.A. in 2012, and news articles spanning a large number of years and 100s of different disasters. The data has been encoded with 36 different categories related to disaster response and has been stripped of messages with sensitive… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/disaster_response_messages.tabulartext-classification10K<n<100K10 likes255 downloads2y agoHugging Face17meoconxinhxan /R1_keep_source_3_messages_only_final_decontaminatedtext100K<n<1M0 likes240 downloads2y agoHugging Face18mshenoda /spam-messages Dataset The dataset is composed of messages labeled by ham or spam, merged from three data sources: SMS Spam Collection https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset Telegram Spam Ham https://huggingface.co/datasets/thehamkercat/telegram-spam-ham/tree/main Enron Spam: https://huggingface.co/datasets/SetFit/enron_spam/tree/main (only used message column and labels) The prepare script for enron is available at… See the full description on the dataset page: https://huggingface.co/datasets/mshenoda/spam-messages.text10K<n<100K3 likes238 downloads1y agoHugging Face19mnoukhov /dolci_think_rl_7b_messages_hybrid_450mtext100K<n<1M0 likes235 downloads2mo agoHugging Face20enPurified /smollm-corpus-fineweb-edu-enPurified-openai-messages 📖 smollm-corpus-fineweb-edu-enPurified-openai-messages smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus. The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.texttext-generation10M<n<100M1 likes222 downloads9mo agoHugging Face21itchson /atomic-deep-short-message Atomic Deep · Short Message An open research dataset of compact language units for language modelling and agent communication. Developed by Atomic Deep, Short Message brings together everyday messages, sentences, code, questions and answers in a common format with source attribution. Snapshot 2026-10-03 · 4,059,552 messages · 14 categories · en The Short Message idea Short Message treats a small, meaningful piece of language as the basic data unit. Each record… See the full description on the dataset page: https://huggingface.co/datasets/itchson/atomic-deep-short-message.tabulartext-generation1M<n<10M0 likes191 downloads7d agoHugging Face22dev-analyzer /commit_messagestext1K<n<10K0 likes158 downloads2y agoHugging Face23Tavernari /git-commit-message-dttext1K<n<10K5 likes149 downloads2y agoHugging Face24UniqueData /spam-text-messages-datasetThe SMS spam dataset contains a collection of text messages. The dataset includes a diverse range of spam messages, including promotional offers, fraudulent schemes, phishing attempts, and other forms of unsolicited communication. Each SMS message is represented as a string of text, and each entry in the dataset also has a link to the corresponding screenshot. The dataset's content represents real-life examples of spam messages that users encounter in their everyday communication.text-classification10K<n<100K2 likes146 downloads1y agoHugging Face25Langame /waiting-messages Langame/waiting-messages Generated using OpenAI GPT-3 davinci-codex based on random initial samples written by a human. ⚠️ The dataset has not been de-duplicated, so there may be duplicates. ⚠️ textn<1K0 likes133 downloads5y agoHugging Face26pszemraj /midjourney-messages-cleaned midjourney-messages-cleaned This is vivym/midjourney-messages but with the following cleaning steps: remove most columns (keep id columns for reference vs. original) Apply clean-text to all rows (keep casing) rename content to text (ffs) remove intermediate ID/tag (???) in angle brackets at the end, remove double asterisks ** remove exact duplicate rows dataset structure overall: DatasetDict({ train: Dataset({ features: ['id', 'channel_id', 'text']… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/midjourney-messages-cleaned.texttext-generation10M<n<100M8 likes119 downloads10mo agoHugging Face27enPurified /tulu-3-sft-mixture-enPurified-openai-messages Dataset Card: enPurified Collection **This dataset was updated on January 13th, 2026 to strip out even more math/code. The pruning process reduced the dataset from 940,000 to 88,782 rows of high-quality English prose. (The script used for this process is uploaded in the files section) Dataset Summary The enPurified collection is a curated suite of datasets designed to isolate high-quality English prose from existing high-value open-source datasets. The primary… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/tulu-3-sft-mixture-enPurified-openai-messages.text-generation2 likes118 downloads9mo agoHugging Face28AriaAICompany /phish-messages Phish synthetic messages 256 synthetic Persian and English messages for the Phish review demo. Seed 3. Organization dataset, model, collection, and static card are public. Live Gradio is created by scripts/publish.py. This is fixture data (level 1). It does not prove operational phishing accuracy. Files data/messages.jsonl data/splits.json data/evaluation.json data/sample_preview.json data/eml/*.eml data/protocol.md Splits Split is by campaign… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/phish-messages.tabulartext-classificationn<1K0 likes112 downloads20d agoHugging Face29saridormi /commit-message-quality Commit Message Quality dataset This is the dataset for commit message quality classification, used during processing of Commit Message Generation dataset from 🏟️ Long Code Arena benchmark. This is a cleaned and relabeled version of the dataset from 📜 "Commit Message Matters: Investigating Impact and Evolution of Commit Message Quality", ICSE'23. We drop "Neither Why nor What" examples, clean all the external references (URLs, issues/PR references) from messages and manually label… See the full description on the dataset page: https://huggingface.co/datasets/saridormi/commit-message-quality.texttext-classification1K<n<10K0 likes110 downloads3y agoHugging Face30heegyu /dolphin-r1-messages-deepseektext100K<n<1M0 likes98 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.