verify-ppt/smollm3-stack-v2-Java
Synthetic Pre-pretraining Datasets This repository contains the pre-processed datasets released with the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. The datasets cover both pre-pretraining (PPT) tasks and pre-training (PT) mixtures used in the paper. PPT tasks include k-Shuffle Dyck, Set, MP-Struct Core, and NCA; PT mixtures include C4, SmolLM3, Olmo3, Marin, and FineWeb-Edu configurations. The pre-processed data are tokenized and organized by… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Java.
0610
15 Z � � � < � � @ � � � U � � � ! 1! �* w- �- [. �0 �3 95 �9 �? ?A �A M 5M +N 'O :Q �Q vR pT �U �V ?W �W .X s �v 1w �y �y ^z �z 9} ' � '� d� �� _� n� � ;� �� ȉ {� �� A� � O� Ĕ Й c� �� �� *� 6� �� 4� �� E� U� �� &� o� �� �� �� � �� �� �� Z� �� �� �� j� �� �� �� �� �� @� �� �� � {� � � � m � � � k e � � �! �! �" �/ 7 �7 D8 1@ �@ GA JB �T �U $X �X �Y [ �[ /\ S^ �^ B_ �a c `d .f 4q �r Hs �v �w bx R 4� � I� �� � $� �� �� � 4� n� �� � �� � �� _� a� �� � `� �� �� $� p� � �� �� �� � $� �� �� �� -� � "� b� �� � 3� v�
� � � � { � 7 �6 �7 �9 �= F �F J �K �L �P �S �T �V �q 0u Gv �v �w �x Ry g| L} � :� )� D� ~� �� �� �� � A� �� �� �� R� �� [� � �� S� � e� �� W� � k� �� /� �� � �� �� 7� C� �� ȩ խ � �� � � �� \� h� �� 2� =� �� �� Q� n� �� >� W� �� � � e� �� � *� Z� �� �� '� S� r� �� [� 3� W� � �� �� �� �� F� �� ( o V z � V 8 . Q �4 � � � $ &$ �� � �� �� ߡ �� � )� �� z� �� �� �� %� =� B� �� �� �� �� � x� *� �� @� A� �� �� D5