Team Ai
Datasetpublic

verify-ppt/smollm3-stack-v2-Cpp

Synthetic Pre-pretraining Datasets This dataset contains the pre-processed synthetic pre-pretraining (PPT) and pre-training (PT) data used in the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. It includes a range of PPT tasks (e.g., k-Shuffle Dyck, Set, MP-Struct Core, NCA) and standard pre-training mixtures (e.g., C4, SmolLM3, Olmo3, Marin) used to evaluate PPT at scale. Code: GitHub repository Project page: Hugging Face Organization For a… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Cpp.

sourceHugging Facemitupdated 5d agoView on Hugging Face
0likes475downloads
Dataset Card

Synthetic Pre-pretraining Datasets

This dataset contains the pre-processed synthetic pre-pretraining (PPT) and pre-training (PT) data used in the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior.

It includes a range of PPT tasks (e.g., k-Shuffle Dyck, Set, MP-Struct Core, NCA) and standard pre-training mixtures (e.g., C4, SmolLM3, Olmo3, Marin) used to evaluate PPT at scale.

For a detailed list of all released datasets (pre-pretraining and pre-training), please refer to the GitHub README.