nicholasKluge/patterns
All datasets used for the patterns project. Language datasets (c4, Fineweb, Open-Web-Math, and Codeparrot) are tokenized with the HuggingFaceTB/SmolLM2-135M tokenizer. See runs in https://wandb.ai/bonn/Patterns.
0762
All datasets used for the patterns project.
Language datasets (c4, Fineweb, Open-Web-Math, and Codeparrot) are tokenized with the HuggingFaceTB/SmolLM2-135M tokenizer.
See runs in https://wandb.ai/bonn/Patterns.
