verify-ppt/smollm3-stack-v2-Java
Synthetic Pre-pretraining Datasets This repository contains the pre-processed datasets released with the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Java.
0610
1---2license: mit3task_categories:4- text-generation5tags:6- synthetic-data7- pre-training8- language-modeling9---10 11# Synthetic Pre-pretraining Datasets12 13This repository contains the pre-processed datasets released with the paper [Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior](https://huggingface.co/papers/2609.39827).14 15The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by configuration.16 17- Paper: [https://huggingface.co/papers/2609.39827](https://huggingface.co/papers/2609.39827)18- Code: [https://github.com/gucci-j/verify-ppt-at-scale](https://github.com/gucci-j/verify-ppt-at-scale)19- Project page: [https://huggingface.co/verify-ppt](https://huggingface.co/verify-ppt)20 21For details on dataset composition, preprocessing, training, and evaluation, please refer to the [GitHub repository](https://github.com/gucci-j/verify-ppt-at-scale).