Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes1.2k downloads2y agoHugging Face02SaylorTwift /RULER-8192-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes351 downloads1y agoHugging Face03SaylorTwift /RULER-32768-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes343 downloads1y agoHugging Face04christopher /all-tokenizerstabular100K<n<1M0 likes317 downloads9mo agoHugging Face05OpenLLM-France /RULER-luciole_tokenizer_128k-arab-regional_v2tabular10K<n<100K0 likes234 downloads10mo agoHugging Face06SlayerLab /tokenizers SlayerLab Tokenizers Normalized tokenizer artifacts collected from the contributor directories in slayerlabs/tokenizer, pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1. The dataset contains one row per tokenizer: the 38 workshop submissions plus the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset Viewer to sort, filter, and compare tokenizers without navigating folders. Columns author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.tabularn<1K0 likes202 downloads16d agoHugging Face07eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes172 downloads1y agoHugging Face08SakethVemula /fixed-tokenizer-morphscore-segmentstabular10M<n<100M0 likes164 downloads7mo agoHugging Face09SaylorTwift /RULER-4096-llama-3.2-tokenizertabular1K<n<10K0 likes116 downloads1y agoHugging Face10himalaya-ai /gpt2-tokenizer-corpustabular1M<n<10M0 likes104 downloads6mo agoHugging Face11christopher /skeletoken-tokenizertabular1K<n<10K0 likes86 downloads9mo agoHugging Face12himalaya-ai /nepali-tokenizer-corpustabular1M<n<10M0 likes70 downloads7mo agoHugging Face13SaylorTwift /RULER-131072-llama-3.2-tokenizertabular1K<n<10K0 likes69 downloads1y agoHugging Face14SaylorTwift /RULER-65536-llama-3.1-tokenizer-chat-templatetabular1K<n<10K0 likes67 downloads1y agoHugging Face15christopher /popular-tokenizerstabular1K<n<10K1 likes64 downloads9mo agoHugging Face16SaylorTwift /RULER-131072-llama-3.1-tokenizer-chat-templatetabular1K<n<10K0 likes62 downloads1y agoHugging Face17SaylorTwift /RULER-16384-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes53 downloads1y agoHugging Face18SaylorTwift /RULER-32768-llama-3.1-tokenizer-chat-templatetabular1K<n<10K0 likes48 downloads1y agoHugging Face19SaylorTwift /RULER-32768-llama-3.2-tokenizertabular1K<n<10K0 likes44 downloads1y agoHugging Face20r-three /farsi_tokenizer_robustness TokSuite Benchmark (Farsi Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team Language(s): Farsi/Persian (fa) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.tabularmultiple-choicen<1K1 likes43 downloads1y agoHugging Face21SaylorTwift /RULER-65536-llama-3.2-tokenizertabular1K<n<10K0 likes41 downloads1y agoHugging Face22SaylorTwift /RULER-4096-llama-3.1-tokenizer-chat-templatetabular1K<n<10K0 likes39 downloads1y agoHugging Face23kenny0bi /tokenizer-tax owóorí: the tokenizer tax, measured Token-count premiums for the 204 languages of FLORES-200 under 12 tokenizers, from the identical 1,012 professionally translated sentences. The premium is tokens(language) / tokens(English) on the same content, so it reads directly as a price multiplier for API cost, latency and context shrinkage. Interactive explorer: https://kenny0bi.github.io/owoori/ Method, figures, code: https://github.com/Kenny0bi/owoori Files… See the full description on the dataset page: https://huggingface.co/datasets/kenny0bi/tokenizer-tax.tabularn<1K0 likes39 downloads1mo agoHugging Face24dakoblov /fr-wiki-popular-200-tokenizer-pruning French Wikipedia corpus for tokenizer pruning Prepared by Daniil Koblov. Contains 200 complete plain-text article extracts: 180 training articles and 20 held-out articles, split with Python's random seed 1337. Candidates come from the 2025 monthly French Wikipedia top-1000 pageview lists. Articles must appear in at least three months. Ranking uses month recurrence, then the sum of reciprocal monthly ranks. The first 200 qualifying articles are shuffled and split. Non-article… See the full description on the dataset page: https://huggingface.co/datasets/dakoblov/fr-wiki-popular-200-tokenizer-pruning.tabularn<1K0 likes39 downloads7d agoHugging Face25SaylorTwift /RULER-8192-llama-3.1-tokenizer-chat-templatetabular1K<n<10K0 likes38 downloads1y agoHugging Face26dineshkarki /nepali-tokenizer-corpustabular100K<n<1M0 likes38 downloads7mo agoHugging Face27SaylorTwift /RULER-131072-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes36 downloads1y agoHugging Face28SaylorTwift /RULER-16384-llama-3.2-tokenizertabular1K<n<10K0 likes34 downloads1y agoHugging Face29Lyte /tokenizer-leaderboard Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): en License: mit Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.tabularn<1K0 likes33 downloads5mo agoHugging Face30NickyNicky /oasst2_orpo_mix_tokenizer_phi_3_v1 https://huggingface.co/datasets/NickyNicky/orpo-dpo-mix-54k tabular10K<n<100K1 likes31 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.