verify-ppt/smollm3-stack-v2-Java
Synthetic Pre-pretraining Datasets This repository contains the pre-processed datasets released with the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Java.
Synthetic Pre-pretraining Datasets
This repository contains the pre-processed datasets released with the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior.
The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by configuration.
- Paper: https://huggingface.co/papers/2609.39827
- Code: https://github.com/gucci-j/verify-ppt-at-scale
- Project page: https://huggingface.co/verify-ppt
For details on dataset composition, preprocessing, training, and evaluation, please refer to the GitHub repository.
