modernbert
midtraining_mix_modernbert_filtered_documentsmailroom-modernbert-training
mailroom-modernbert-training
Cleaned + prepared hierarchical-classification training set for the
ModernBERT-base ingest fast-path — the fine-tuning surface of the
mailroom-ml synthetic-data layer.
The layer (end to end)
Lucius-Morningstar/mailroom-dataset corpus (GT labels, canonical v9)
│ (working copy, pinned)
▼
THIS REPO (mailroom-modernbert-training) @ pinned revision
│ training/train_modernbert.py (--data <repo>)
▼… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/mailroom-modernbert-training.dolma3-hq-2M-modernbert
Dolma3 High-Quality 2M (ModernBERT Filtered)
A curated subset of 2 million high-quality text samples from allenai/dolma3_dolmino_mix-100B-1125, filtered to fit within ModernBERT's 8192 token context window.
Dataset Description
This dataset is designed for pretraining diffusion language models based on ModernBERT. Each sample has been:
Source filtered: Only from ingredient1-common_crawl-high-quality folders (highest quality web text)
Length filtered: Minimum 200… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/dolma3-hq-2M-modernbert.data_ablation_full59K-modernbert-split-kmeans-dim768-20250218wikitext-tags-modernbertfineweb-10b-512-modernbert
FineWeb-Edu — ModernBERT continuous packed chunks
Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files).
Tokenizer: answerdotai/ModernBERT-large. No truncation or padding.
Each nonempty document contributes CLS (50281), document IDs, SEP (50282).
The concatenated stream is split into 512-token rows. Documents
may span chunks; a chunk need not begin with CLS or end with SEP.
Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.
