Team Ai
Datasetpublic

verify-ppt/smollm3-stack-v2-Java

Synthetic Pre-pretraining Datasets This repository contains the pre-processed datasets released with the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Java.

sourceHugging Facemitupdated 22h agoView on Hugging Face
0likes73downloads
7 commits on main
d7ea32622h ago

Add dataset card with metadata and links (#1)

atsuki-yamaguchi, nielsr
14b6b3e7mo ago

Add files using upload-large-folder tool

atsuki-yamaguchi
febd37a7mo ago

Add files using upload-large-folder tool

atsuki-yamaguchi
c5f65337mo ago

Add files using upload-large-folder tool

atsuki-yamaguchi
37edcd37mo ago

Add files using upload-large-folder tool

atsuki-yamaguchi
8b22b1d7mo ago

Add files using upload-large-folder tool

atsuki-yamaguchi
843022a7mo ago

initial commit

atsuki-yamaguchi