Team Ai
Datasetpublic

itchson/atomic-deep-short-message

Atomic Deep · Short Message An open research dataset of compact language units for language modelling and agent communication. Developed by Atomic Deep, Short Message brings together everyday messages, sentences, code, questions and answers in a common format with source attribution. Snapshot 2026-10-03 · 4,059,552 messages · 14 categories · en The Short Message idea Short Message treats a small, meaningful piece of language as the basic data unit. Each record… See the full description on the dataset page: https://huggingface.co/datasets/itchson/atomic-deep-short-message.

sourceHugging Faceotherupdated 7d agoView on Hugging Face
0likes191downloads
Dataset Card

<p align="center"><img src="assets/banner.svg" alt="Atomic Deep Short Message dataset overview" width="100%"></p>

Atomic Deep · Short Message

An open research dataset of compact language units for language modelling and agent communication. Developed by Atomic Deep, Short Message brings together everyday messages, sentences, code, questions and answers in a common format with source attribution.

Snapshot 2026-10-03 · 4,059,552 messages · 14 categories · en

The Short Message idea

Short Message treats a small, meaningful piece of language as the basic data unit. Each record pairs its text with a category, origin and source information. These units support research into compact language models, controlled data mixtures and concise communication between agents.

[image]

Dataset contents

ConfigurationRecordsContent
messages4,059,552Messages, sentences, code and question–answer items
reasoning_experimental192Short prompts, answers and explanations

Messages are distributed as 443.08 MB of compressed Parquet. Records retain their text, category, language, origin, source terms and recorded quality checks.

[image]

Sources

OriginMessagesShare
Synthetic / generated4,007,54098.72%
Extracted public text52,0121.28%

Source and generator attribution is available in each record and the source ledger. Synthetic origin describes how a record was produced; generated text may reproduce existing material.

Examples

<details> <summary><strong>SMS-style text</strong> · synthetic</summary>

<pre>Hi, running 10 mins late to coffee. Start without me!</pre>

Terms: synthetic-model-output · Record ID: <code>84ced58c5acf1ebf0d96977585835a9f8f88372dc18c7e29dfc52119861ccbbc</code>

</details>

<details> <summary><strong>Code snippets</strong> · synthetic</summary>

<pre>def greet(name): return f&#x27;Hello, {name}!&#x27;

print(greet(&#x27;Alice&#x27;))</pre>

Terms: synthetic-model-output · Record ID: <code>8b2586b6fd4a883d3128533119a78ab2c60511fd334ffc3786feccde3e1dcf27</code>

</details>

<details> <summary><strong>Trivia</strong> · extracted</summary>

<pre>Q: How many islands does Kuwait have? A: 9</pre>

Terms: CC-BY-SA-4.0 · Record ID: <code>30a31214cef4aa27e8feb479c8c6ac9c518a56d6fb216d0d8fff5a07dc1c2b5e</code>

</details>

Short reasoning · experimental

The separate pilot contains 192 scripted examples in arithmetic, unit conversion and ordering. Each combines a short prompt, an answer and an authored explanation, with an independent answer verifier.

PromptAnswerExplanation
A price is 360 credits. Apply a 20% discount. What is the final price in credits?288Keep 80%: 360 × 80/100 = 288.

The pilot is experimental and marked training_eligible: false. Template families stay within one split; domains and language patterns remain shared. See the reasoning manifest for the schema, character limits and verification details.

Quick start

python
from datasets import load_dataset

revision = "snapshot-2026-10-03-r3"
messages = load_dataset("itchson/atomic-deep-short-message", "messages",
                        split="train", revision=revision, streaming=True)
print(next(iter(messages)))

Use "reasoning_experimental" as the configuration name to load the reasoning pilot.

SplitMessagesExperimental reasoning
train3,859,33696
validation120,31048
test79,90648

This is a versioned snapshot. File checksums and release metadata are in the release manifest.

Licence and data quality

Terms vary by source. See LICENSE.md and the upstream licences for attribution and usage conditions. The authored reasoning fixtures are CC0-1.0.

The corpus is predominantly synthetic and automatically checked. Accuracy, originality and freedom from benchmark overlap are not guaranteed; review record provenance and quality fields for downstream use.

Contributing

Contributions are welcome: new public sources, corrections, broader language coverage, compact reasoning tasks and evaluation tools. Open a Discussion or submit a pull request following CONTRIBUTING.md.

Citation

bibtex
@misc{atomicdeep_short_message_2026,
  author = {Atomic Deep},
  title = {Atomic Deep Short Message},
  year = {2026},
  url = {https://huggingface.co/datasets/itchson/atomic-deep-short-message},
  note = {Snapshot 2026-10-03; revision snapshot-2026-10-03-r3}
}