Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes386 downloads10mo agoHugging Face02Lucius-Morningstar /mailroom-modernbert-training mailroom-modernbert-training Cleaned + prepared hierarchical-classification training set for the ModernBERT-base ingest fast-path — the fine-tuning surface of the mailroom-ml synthetic-data layer. The layer (end to end) Lucius-Morningstar/mailroom-dataset corpus (GT labels, canonical v9) │ (working copy, pinned) ▼ THIS REPO (mailroom-modernbert-training) @ pinned revision │ training/train_modernbert.py (--data <repo>) ▼… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/mailroom-modernbert-training.tabulartext-classification1K<n<10K1 likes293 downloads15d agoHugging Face03Ayushnangia /dolma3-hq-2M-modernbert Dolma3 High-Quality 2M (ModernBERT Filtered) A curated subset of 2 million high-quality text samples from allenai/dolma3_dolmino_mix-100B-1125, filtered to fit within ModernBERT's 8192 token context window. Dataset Description This dataset is designed for pretraining diffusion language models based on ModernBERT. Each sample has been: Source filtered: Only from ingredient1-common_crawl-high-quality folders (highest quality web text) Length filtered: Minimum 200… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/dolma3-hq-2M-modernbert.texttext-generation1M<n<10M0 likes162 downloads9mo agoHugging Face04albertge /data_ablation_full59K-modernbert-split-kmeans-dim768-20250218tabular10K<n<100K0 likes111 downloads2y agoHugging Face05hriaz /wikitext-tags-modernberttext1M<n<10M0 likes106 downloads1y agoHugging Face06dme5245 /fineweb-10b-512-modernbert FineWeb-Edu — ModernBERT continuous packed chunks Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files). Tokenizer: answerdotai/ModernBERT-large. No truncation or padding. Each nonempty document contributes CLS (50281), document IDs, SEP (50282). The concatenated stream is split into 512-token rows. Documents may span chunks; a chunk need not begin with CLS or end with SEP. Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.10M<n<100M0 likes86 downloads2d agoHugging Face07robro612 /nfcorpus_modernbert_xtr nfcorpus_modernbert_xtr Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with robro612/ModernBERT-XTR at revision 8c06fce0b8d2387be98183582eb9600af7bd1b8a. Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded through ir_datasets, not… See the full description on the dataset page: https://huggingface.co/datasets/robro612/nfcorpus_modernbert_xtr.tabular10K<n<100K0 likes61 downloads7d agoHugging Face08Mnwa /Russian-toxic-modernbert Tokenized Russian toxic text Tokenized version of Mnwa/russian-toxic dataset with modernbert base model tokenizer text-classification100K<n<1M0 likes60 downloads2y agoHugging Face09NamburiSrinath /ni-ood-dataset-20250131-modernbert-train-kmeans-dim768-20250318tabular1M<n<10M0 likes59 downloads2y agoHugging Face10stephantulkens /paws-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the google-research-datasets/paws dataset This is the google-research-datasets/paws dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/paws-gte-modernbert-pooled.text1M<n<10M0 likes59 downloads11mo agoHugging Face11feyninc /chonkiepedia-modernbert-tokenized1M<n<10M0 likes56 downloads1y agoHugging Face12rchu233 /ni-ood-dataset-20250131-modernbert-train-kmeans-dim128-20250312tabular1M<n<10M0 likes50 downloads2y agoHugging Face13stephantulkens /pubmedqa-query-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the qiaojin/PubMedQA dataset This is the qiaojin/PubMedQA dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/pubmedqa-query-gte-modernbert-pooled.text100K<n<1M0 likes47 downloads11mo agoHugging Face14rchu233 /ni-ood-dataset-10p-20250127-modernbert-kmeans-dim128-20250128tabular100K<n<1M0 likes42 downloads2y agoHugging Face15stephantulkens /msmarco-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/msmarco-corpus dataset This is the sentence-transformers/msmarco-corpus dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-gte-modernbert-pooled.text1M<n<10M0 likes42 downloads11mo agoHugging Face16waleedk7-fyp /final_train_test_split_modernbert_large0 likes42 downloads5mo agoHugging Face17albertge /SlimPajama-6B-modernbert-split-kmeans-dim768-20250316tabular1M<n<10M0 likes41 downloads2y agoHugging Face18Unggi /modernbert_encoder_sp_seq_512_csedm_fold1tabularn<1K0 likes40 downloads2y agoHugging Face19albertge /ni-20-clustered-fulltext-modernbert-sweep-20250107tabular10K<n<100K0 likes36 downloads2y agoHugging Face20stephantulkens /mdlr-query-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/mldr dataset This is the sentence-transformers/mldr dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to the… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/mdlr-query-gte-modernbert-pooled.text10K<n<100K0 likes36 downloads11mo agoHugging Face21stephantulkens /snli-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the stanfordnlp/snli dataset This is the stanfordnlp/snli dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/snli-gte-modernbert-pooled.text100K<n<1M0 likes33 downloads11mo agoHugging Face22albertge /ni-unique-20-tasks-modernbert-512-20250111tabular10K<n<100K0 likes29 downloads2y agoHugging Face23albertge /databricks-dolly-15k-modernbert-kmeans-dim768-normalize-20250130tabular10K<n<100K0 likes28 downloads2y agoHugging Face24Unggi /modernbert_encoder_sp_seq_512_dbe22kt_fold1tabularn<1K0 likes28 downloads2y agoHugging Face25AITeamUIT /eval-gliner2-modernbert_pasteproof-uni-202606210 likes28 downloads4mo agoHugging Face26rchu233 /ni-ood-dataset-10p-20250127-modernbert-split-kmeans-dim128-20250128tabular100K<n<1M0 likes27 downloads2y agoHugging Face27stephantulkens /english-word-definitions-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the MongoDB/english-words-definitions dataset This is the MongoDB/english-words-definitions dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/english-word-definitions-gte-modernbert-pooled.text100K<n<1M0 likes27 downloads11mo agoHugging Face28AmanPriyanshu /GTE-ModernBERT-RedPajama-Data-1T-100k-SubSample-max-1k-tokenstext100K<n<1M0 likes26 downloads2y agoHugging Face29stephantulkens /mr-tydi-query-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/mr-tydi dataset This is the sentence-transformers/mr-tydi dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/mr-tydi-query-gte-modernbert-pooled.text1K<n<10K0 likes26 downloads11mo agoHugging Face30albertge /dolly-15k-clustered-modernbert-768-20250111tabular10K<n<100K0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.