Team Ai
20 results

modernbert

geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes386 downloads10mo agoHugging FaceLucius-Morningstar /mailroom-modernbert-training mailroom-modernbert-training Cleaned + prepared hierarchical-classification training set for the ModernBERT-base ingest fast-path — the fine-tuning surface of the mailroom-ml synthetic-data layer. The layer (end to end) Lucius-Morningstar/mailroom-dataset corpus (GT labels, canonical v9) │ (working copy, pinned) ▼ THIS REPO (mailroom-modernbert-training) @ pinned revision │ training/train_modernbert.py (--data <repo>) ▼… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/mailroom-modernbert-training.tabulartext-classification1K<n<10K1 likes293 downloads15d agoHugging FaceAyushnangia /dolma3-hq-2M-modernbert Dolma3 High-Quality 2M (ModernBERT Filtered) A curated subset of 2 million high-quality text samples from allenai/dolma3_dolmino_mix-100B-1125, filtered to fit within ModernBERT's 8192 token context window. Dataset Description This dataset is designed for pretraining diffusion language models based on ModernBERT. Each sample has been: Source filtered: Only from ingredient1-common_crawl-high-quality folders (highest quality web text) Length filtered: Minimum 200… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/dolma3-hq-2M-modernbert.texttext-generation1M<n<10M0 likes162 downloads9mo agoHugging Facealbertge /data_ablation_full59K-modernbert-split-kmeans-dim768-20250218tabular10K<n<100K0 likes111 downloads2y agoHugging Facehriaz /wikitext-tags-modernberttext1M<n<10M0 likes106 downloads1y agoHugging Facedme5245 /fineweb-10b-512-modernbert FineWeb-Edu — ModernBERT continuous packed chunks Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files). Tokenizer: answerdotai/ModernBERT-large. No truncation or padding. Each nonempty document contributes CLS (50281), document IDs, SEP (50282). The concatenated stream is split into 512-token rows. Documents may span chunks; a chunk need not begin with CLS or end with SEP. Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.10M<n<100M0 likes86 downloads2d agoHugging Face