dme5245/fineweb-10b-512-modernbert
FineWeb-Edu — ModernBERT continuous packed chunks Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files). Tokenizer: answerdotai/ModernBERT-large. No truncation or padding. Each nonempty document contributes CLS (50281), document IDs, SEP (50282). The concatenated stream is split into 512-token rows. Documents may span chunks; a chunk need not begin with CLS or end with SEP. Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.
FineWeb-Edu — ModernBERT continuous packed chunks
Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files). Tokenizer: answerdotai/ModernBERT-large. No truncation or padding. Each nonempty document contributes CLS (50281), document IDs, SEP (50282). The concatenated stream is split into 512-token rows. Documents may span chunks; a chunk need not begin with CLS or end with SEP. Original sorted-file and row order is preserved throughout processing.
Schema: input_ids: list<int32>. Every row has 512 tokens. Chunks: 19,392,828. Packed tokens: 9,929,127,936. The final 345 tokens are retained locally in remainder.npy and are not included in the fixed-length training split.
from datasets import load_dataset
from torch.utils.data import DataLoader
ds = load_dataset("dme5245/fineweb-10b-512-modernbert", split="train").with_format("torch")
for batch in DataLoader(ds, batch_size=32, num_workers=4):
input_ids = batch["input_ids"]