datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midtraining_mix_modernbert_filtered_documentsmailroom-modernbert-training
mailroom-modernbert-training
Cleaned + prepared hierarchical-classification training set for the
ModernBERT-base ingest fast-path — the fine-tuning surface of the
mailroom-ml synthetic-data layer.
The layer (end to end)
Lucius-Morningstar/mailroom-dataset corpus (GT labels, canonical v9)
│ (working copy, pinned)
▼
THIS REPO (mailroom-modernbert-training) @ pinned revision
│ training/train_modernbert.py (--data <repo>)
▼… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/mailroom-modernbert-training.dolma3-hq-2M-modernbert
Dolma3 High-Quality 2M (ModernBERT Filtered)
A curated subset of 2 million high-quality text samples from allenai/dolma3_dolmino_mix-100B-1125, filtered to fit within ModernBERT's 8192 token context window.
Dataset Description
This dataset is designed for pretraining diffusion language models based on ModernBERT. Each sample has been:
Source filtered: Only from ingredient1-common_crawl-high-quality folders (highest quality web text)
Length filtered: Minimum 200… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/dolma3-hq-2M-modernbert.data_ablation_full59K-modernbert-split-kmeans-dim768-20250218wikitext-tags-modernbertfineweb-10b-512-modernbert
FineWeb-Edu — ModernBERT continuous packed chunks
Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files).
Tokenizer: answerdotai/ModernBERT-large. No truncation or padding.
Each nonempty document contributes CLS (50281), document IDs, SEP (50282).
The concatenated stream is split into 512-token rows. Documents
may span chunks; a chunk need not begin with CLS or end with SEP.
Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.nfcorpus_modernbert_xtr
nfcorpus_modernbert_xtr
Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with
robro612/ModernBERT-XTR at revision 8c06fce0b8d2387be98183582eb9600af7bd1b8a.
Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded through ir_datasets, not… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_modernbert_xtr.Russian-toxic-modernbert
Tokenized Russian toxic text
Tokenized version of Mnwa/russian-toxic dataset with modernbert base model tokenizer
ni-ood-dataset-20250131-modernbert-train-kmeans-dim768-20250318paws-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the google-research-datasets/paws dataset
This is the google-research-datasets/paws dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/paws-gte-modernbert-pooled.chonkiepedia-modernbert-tokenizedni-ood-dataset-20250131-modernbert-train-kmeans-dim128-20250312pubmedqa-query-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the qiaojin/PubMedQA dataset
This is the qiaojin/PubMedQA dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/pubmedqa-query-gte-modernbert-pooled.ni-ood-dataset-10p-20250127-modernbert-kmeans-dim128-20250128msmarco-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/msmarco-corpus dataset
This is the sentence-transformers/msmarco-corpus dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-gte-modernbert-pooled.final_train_test_split_modernbert_largeSlimPajama-6B-modernbert-split-kmeans-dim768-20250316modernbert_encoder_sp_seq_512_csedm_fold1ni-20-clustered-fulltext-modernbert-sweep-20250107mdlr-query-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/mldr dataset
This is the sentence-transformers/mldr dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to the… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/mdlr-query-gte-modernbert-pooled.snli-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the stanfordnlp/snli dataset
This is the stanfordnlp/snli dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/snli-gte-modernbert-pooled.ni-unique-20-tasks-modernbert-512-20250111databricks-dolly-15k-modernbert-kmeans-dim768-normalize-20250130modernbert_encoder_sp_seq_512_dbe22kt_fold1eval-gliner2-modernbert_pasteproof-uni-20260621ni-ood-dataset-10p-20250127-modernbert-split-kmeans-dim128-20250128english-word-definitions-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the MongoDB/english-words-definitions dataset
This is the MongoDB/english-words-definitions dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/english-word-definitions-gte-modernbert-pooled.GTE-ModernBERT-RedPajama-Data-1T-100k-SubSample-max-1k-tokensmr-tydi-query-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/mr-tydi dataset
This is the sentence-transformers/mr-tydi dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/mr-tydi-query-gte-modernbert-pooled.dolly-15k-clustered-modernbert-768-20250111
