datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tokenizer-conformance
Tokenizer conformance fixtures
Reference inputs and Python fast-tokenizer outputs for tokenizer implementations.
The initial corpus contains 83 inputs in 30 categories, with 498 reference encodings across six tokenizers.
This is a regression dataset, not a model-quality benchmark.
Provenance and attribution
The input corpus and reference entries come from apocryphx's swift-transformers PR #360,
at commit ce847085784bacd8c3c15180c976b17c8ce73e31.
The corpus… See the full description on the dataset page: https://huggingface.co/datasets/pcuenq/tokenizer-conformance.sr-tokenizer-test
Sr Tokenizer test
This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models.
It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fields.
Dataset Structure
Metadata has been stripped; Each record is a JSON object with:
id: unique identifier
text: raw Serbian text
Source coprora
Znanje(sr) corpus: ~6.6 GB… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/sr-tokenizer-test.speakleash-tokenizer-5gb-sample
SpeakLeash tokenizer 42GB quality sample
Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10.
Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup.
Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind.
This is intended for tokenizer/BPE training convergence tests.
tokenizer-tax
owóorí: the tokenizer tax, measured
Token-count premiums for the 204 languages of FLORES-200 under 12
tokenizers, from the identical 1,012 professionally translated sentences.
The premium is tokens(language) / tokens(English) on the same content, so
it reads directly as a price multiplier for API cost, latency and context
shrinkage.
Interactive explorer: https://kenny0bi.github.io/owoori/
Method, figures, code: https://github.com/Kenny0bi/owoori
Files… See the full description on the dataset page: https://huggingface.co/datasets/kenny0bi/tokenizer-tax.fr-wiki-popular-200-tokenizer-pruning
French Wikipedia corpus for tokenizer pruning
Prepared by Daniil Koblov. Contains 200 complete plain-text article extracts:
180 training articles and 20 held-out articles, split with Python's random seed 1337.
Candidates come from the 2025 monthly French Wikipedia top-1000 pageview lists.
Articles must appear in at least three months. Ranking uses month recurrence,
then the sum of reciprocal monthly ranks. The first 200 qualifying articles are
shuffled and split. Non-article… See the full description on the dataset page: https://huggingface.co/datasets/dakoblov/fr-wiki-popular-200-tokenizer-pruning.Vietnamese-Tokenizer-Corpus
Vietnamese Tokenizer Training Corpus v1
This dataset is the training corpus used to build PaxiAI/Vietnamese-Tokenizer.
It was created by sampling and combining Vietnamese, English, and source-code datasets into an approximately 8 GiB corpus intended specifically for tokenizer training.
Its purpose is to provide a diverse and representative sample from which a Vietnamese-focused Byte-level BPE vocabulary can be learned.
Dataset Summary
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/PaxiAI/Vietnamese-Tokenizer-Corpus.co-funer
CO-Fun: Tokenized Sentences
This datasets hosts a sentence-tokenized version of the CO-Fun: A German Dataset on Company Outsourcing in Fund Prospectuses for Named Entity Recognition and Relation Extraction dataset.
Creation
The following script can be used to reproduce the creation of the dataset:
import flair
import json
from flair.datasets.sequence_labeling import ColumnCorpus
from flair.file_utils import cached_path
from pathlib import Path
from typing import… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/co-funer.german-ler
German LER: Tokenized Sentences
This datasets hosts a sentence-tokenized version of the German LER dataset.
Creation
The following script can be used to reproduce the creation of the dataset:
import json
from flair.datasets import NER_GERMAN_LEGAL
corpus = NER_GERMAN_LEGAL()
with open("./train.jsonl", "wt") as f_out:
for sentence in corpus.train:
current_example = {
"text": sentence.to_tokenized_string()
}… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/german-ler.tiny_tokenizer_tokenstokenizerfw_edu_tokenizer_training_dataThis dataset is a small subset sampled from the 10bt subset of FineWeb-Edu dataset
germeval14
GermEval 2014: Tokenized Sentences
This datasets hosts a sentence-tokenized version of the GermEval 2014 NER dataset.
Creation
The following script can be used to reproduce the creation of the dataset:
import json
from flair.datasets import NER_GERMAN_GERMEVAL
corpus = NER_GERMAN_GERMEVAL()
with open("./germeval14/train.jsonl", "wt") as f_out:
for sentence in germeval_corpus.train:
current_example = {
"text": sentence.to_tokenized_string()… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/germeval14.biofid
BIOfid: Tokenized Sentences
This datasets hosts a sentence-tokenized version of the BIOfid dataset.
Creation
The following script can be used to reproduce the creation of the dataset:
import json
from flair.datasets import NER_GERMAN_BIOFID
corpus = NER_GERMAN_BIOFID()
with open("./train.jsonl", "wt") as f_out:
for sentence in corpus.train:
current_example = {
"text": sentence.to_tokenized_string()
}… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/biofid.ud-hdt
UD German-HDT: Tokenized Sentences
This datasets hosts a sentence-tokenized version of the Universal Dependencies German-HDT dataset.
Creation
The following script can be used to reproduce the creation of the dataset:
import json
from flair.datasets import UD_GERMAN_HDT
corpus = UD_GERMAN_HDT()
with open("./train.jsonl", "wt") as f_out:
for sentence in corpus.train:
current_example = {
"text": sentence.to_tokenized_string()
}… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/ud-hdt.jp-datasets-for-tokenizerstokenizer-benchmark
kamoo tokenizer-benchmark
Dezelfde Nederlandse alinea door meerdere tokenizers. Minder tokens = meer
context in hetzelfde venster en lagere kosten per antwoord. Reken het na.
Dit is het meetlog achter de tokenizer-claims van kamoo.nl.
Alles in deze repo is genoeg om de meting zelf te herhalen: de testalinea,
het script en de uitkomsten.
Meting (2026-07-06)
Testalinea: de eerste alinea van het Nederlandse Wikipedia-artikel
Nederland (CC-BY-SA, opgehaald
2026-07-06)… See the full description on the dataset page: https://huggingface.co/datasets/kamoo-ai/tokenizer-benchmark.Qwen_coder_tokenizer-PrimeVul_splitssmarter-llm-tokenizer-datamobie
DFKI MobIE: Tokenized Sentences
This datasets hosts a sentence-tokenized version of the DFKI MobIE dataset.
Creation
The following script can be used to reproduce the creation of the dataset:
import json
from flair.datasets import NER_GERMAN_MOBIE
corpus = NER_GERMAN_MOBIE()
with open("./train.jsonl", "wt") as f_out:
for sentence in corpus.train:
current_example = {
"text": sentence.to_tokenized_string()
}… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/mobie.deepwriting-stroke-tokenizer-256
