Team Ai
Datasetpublic

dme5245/fineweb-10b-512-modernbert

FineWeb-Edu — ModernBERT continuous packed chunks Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files). Tokenizer: answerdotai/ModernBERT-large. No truncation or padding. Each nonempty document contributes CLS (50281), document IDs, SEP (50282). The concatenated stream is split into 512-token rows. Documents may span chunks; a chunk need not begin with CLS or end with SEP. Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.

sourceHugging Faceodc-byupdated 2d agoView on Hugging Face
0likes86downloads
Dataset Card

FineWeb-Edu — ModernBERT continuous packed chunks

Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files). Tokenizer: answerdotai/ModernBERT-large. No truncation or padding. Each nonempty document contributes CLS (50281), document IDs, SEP (50282). The concatenated stream is split into 512-token rows. Documents may span chunks; a chunk need not begin with CLS or end with SEP. Original sorted-file and row order is preserved throughout processing.

Schema: input_ids: list<int32>. Every row has 512 tokens. Chunks: 19,392,828. Packed tokens: 9,929,127,936. The final 345 tokens are retained locally in remainder.npy and are not included in the fixed-length training split.

python
from datasets import load_dataset
from torch.utils.data import DataLoader

ds = load_dataset("dme5245/fineweb-10b-512-modernbert", split="train").with_format("torch")
for batch in DataLoader(ds, batch_size=32, num_workers=4):
    input_ids = batch["input_ids"]