Gugu8/Token-Efficiency
token_efficiency_corpus A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs. Progresses from basic filler removal to expert-level nested reasoning compression. Contents verbose_output - The padded, wasteful version of the text efficient_output - The compressed, token-efficient equivalent technique - Compression strategy used subcategory - Specific variant of the technique difficulty - Tier 1 (easiest)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Token-Efficiency.
tokenefficiencycorpus
A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs. Progresses from basic filler removal to expert-level nested reasoning compression.
Contents
verboseoutput - The padded, wasteful version of the text efficientoutput - The compressed, token-efficient equivalent technique - Compression strategy used subcategory - Specific variant of the technique difficulty - Tier 1 (easiest) -> Tier 8 (hardest) notes - Reserved for future metadata
Difficulty tiers
Tier 1 — removefillerwords Strip padding phrases like "it is important to note that", "in order to", "due to the fact that", etc.
Tier 2 — useshortersynonyms Replace long phrases with single-word equivalents (utilize -> use, subsequently -> then, commence -> start ...).
Tier 3 — removeredundantcontext Delete information the reader already knows from the surrounding text or conversation history.
Tier 4 — usestructuredformat Convert paragraph prose into key=value, lists, and machine-readable pipe-delimited strings.
Tier 5 — omittrivialsteps Remove deductions the reader can derive themselves; jump straight to the conclusion.
Tier 6 — deduplicatecrossfields Reference shared attributes instead of repeating them across multiple sentences or fields.
Tier 7 — domain_shorthand Use field-standard abbreviations (MI, PCI, DAPT, TTL, TAR) that practitioners understand but novices do not need spelled out.
Tier 8 — compressnestedreasoning Collapse multi-step chains of logic (cause -> effect -> decision) into a single supported conclusion with the key constraints only.
Training usage
Load with pandas:
import pandas as pd df = pd.readcsv("tokenefficiency_corpus.csv")
Supervised compression task:
X = df["verboseoutput"] y = df["efficientoutput"]
Classification task (identify compression technique):
cls_y = df[["technique", "difficulty"]]
Fine-tuning an LLM with the pair as input -> target:
# input = "Compress this:\n{verboseoutput}" # target = efficientoutput
Design notes
- Streaming-first generation — rows are flushed to disk continuously, keeping peak memory small.
- QUOTE_ALL CSV mode ensures commas and newlines inside fields never break parsing.
- Bidirectional training is possible with a reverse pass (efficient -> verbose) to teach both compression and expansion.
- Harder tiers get more rows because each example encodes a more nuanced compression insight.
