Team Ai
Datasetpublic

LakoreAI/bert-mlm-experiments-en

Unified English MLM Pre-training Corpus (80M Rows) This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings. Dataset Details Repository ID: 8Opt/bert-mlm-experiments-en Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.

sourceHugging Faceopenrailupdated 4mo agoView on Hugging Face
1likes1.4kdownloads
4 commits on main
a94283d4mo ago

Upload dataset (part 00001-of-00002)

minhleduc
a2215cc4mo ago

Upload dataset (part 00000-of-00002)

minhleduc
229615b4mo ago

Create README.md

minhanhlt
e2beeae4mo ago

initial commit

minhleduc