Team Ai
18 results

toksuite

toksuitebackup /meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes1k downloads11mo agoHugging Facetoksuite /toksuite_pretraining_datatext100M<n<1B0 likes899 downloads6mo agoHugging Facetoksuitebackup /aya-expanse-8b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes745 downloads11mo agoHugging Facetoksuitebackup /gpt-4o-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M1 likes655 downloads11mo agoHugging Facetoksuite /toksuite_italian Dataset Card for Tokenization Robustness TokSuite Benchmark (Italian Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Italian language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_italian.textmultiple-choice1K<n<10K0 likes513 downloads9mo agoHugging Facetoksuite /toksuite_english Dataset Card for Tokenization Robustness TokSuite Benchmark (English Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness in isolation. This specific collection contains English multiple-choice text completion questions paired with a wide range of real-world surface-form perturbations that are known to interact… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_english.textmultiple-choice1K<n<10K0 likes359 downloads9mo agoHugging Face