datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.RULER-8192-Qwen2.5-3B-tokenizerRULER-32768-Qwen2.5-3B-tokenizerall-tokenizersRULER-luciole_tokenizer_128k-arab-regional_v2tokenizers
SlayerLab Tokenizers
Normalized tokenizer artifacts collected from the contributor directories in
slayerlabs/tokenizer,
pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1.
The dataset contains one row per tokenizer: the 38 workshop submissions plus
the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset
Viewer to sort, filter, and compare tokenizers without navigating folders.
Columns
author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.fixed-tokenizer-morphscore-segmentsRULER-4096-llama-3.2-tokenizergpt2-tokenizer-corpusskeletoken-tokenizernepali-tokenizer-corpusRULER-131072-llama-3.2-tokenizerRULER-65536-llama-3.1-tokenizer-chat-templatepopular-tokenizersRULER-131072-llama-3.1-tokenizer-chat-templateRULER-16384-Qwen2.5-3B-tokenizerRULER-32768-llama-3.1-tokenizer-chat-templateRULER-32768-llama-3.2-tokenizerfarsi_tokenizer_robustness
TokSuite Benchmark (Farsi Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team
Language(s): Farsi/Persian (fa)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.RULER-65536-llama-3.2-tokenizerRULER-4096-llama-3.1-tokenizer-chat-templatetokenizer-tax
owóorí: the tokenizer tax, measured
Token-count premiums for the 204 languages of FLORES-200 under 12
tokenizers, from the identical 1,012 professionally translated sentences.
The premium is tokens(language) / tokens(English) on the same content, so
it reads directly as a price multiplier for API cost, latency and context
shrinkage.
Interactive explorer: https://kenny0bi.github.io/owoori/
Method, figures, code: https://github.com/Kenny0bi/owoori
Files… See the full description on the dataset page: https://huggingface.co/datasets/kenny0bi/tokenizer-tax.fr-wiki-popular-200-tokenizer-pruning
French Wikipedia corpus for tokenizer pruning
Prepared by Daniil Koblov. Contains 200 complete plain-text article extracts:
180 training articles and 20 held-out articles, split with Python's random seed 1337.
Candidates come from the 2025 monthly French Wikipedia top-1000 pageview lists.
Articles must appear in at least three months. Ranking uses month recurrence,
then the sum of reciprocal monthly ranks. The first 200 qualifying articles are
shuffled and split. Non-article… See the full description on the dataset page: https://huggingface.co/datasets/dakoblov/fr-wiki-popular-200-tokenizer-pruning.RULER-8192-llama-3.1-tokenizer-chat-templatenepali-tokenizer-corpusRULER-131072-Qwen2.5-3B-tokenizerRULER-16384-llama-3.2-tokenizertokenizer-leaderboard
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.oasst2_orpo_mix_tokenizer_phi_3_v1
https://huggingface.co/datasets/NickyNicky/orpo-dpo-mix-54k
