verify-ppt/smollm3-stack-v2-Java
Synthetic Pre-pretraining Datasets This repository contains the pre-processed datasets released with the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Java.
0610
1p � � �
� l. �. �1 "2 +6 M8 �: �A ,B �B rD 2H �H �H sI �L �N �O ,P CS �S �T �\ ~e �e =f g h �h �h �r �r t �v �x �y �{ g� �� � � � � p� ^� � �� �� g� ,� � B� �� #� IJ �� J� �� Ź � � @� �� t� �� �� &� F� w� �� �� =� �� � ] � k � r! �! 3" h/ �/ �1 f3 4 �4 �5 J il �m /} �} � �� �� �� ڂ ޅ �� ю � �� � � �� �� u� � q� � � � � 9� ͮ X� $ �4 � � 5 � � , 6 � � � 4 E � n m � " �" D$ �+ �@ �A B fC �F G AG oJ Sy �y | �~ n� �� _� �� � ݇ j� �� ȉ <� ^� c� �� b� � � � -� �� �� <� ռ (� �� �� #� t� �� p� }� �� q� 4� �� �� �� �� �� �� D� B� j� � �� z� �� 7� (� R� �� �� � �� ?� c5