Team Ai
Datasetpublic

verify-ppt/smollm3-stack-v2-Java

Synthetic Pre-pretraining Datasets This repository contains the pre-processed datasets released with the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Java.

sourceHugging Facemitupdated 4d agoView on Hugging Face
0likes298downloads
stack-v2-Java_00007_00000_shuffled.ds4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:592a5b8afee7dd2cdf8b2da71dfe79e6de5dbb74bf1785a766ff2a14caf40ad93size 127885604