Team Ai
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes172 downloads1y agoHugging Face02pcuenq /tokenizer-conformance Tokenizer conformance fixtures Reference inputs and Python fast-tokenizer outputs for tokenizer implementations. The initial corpus contains 83 inputs in 30 categories, with 498 reference encodings across six tokenizers. This is a regression dataset, not a model-quality benchmark. Provenance and attribution The input corpus and reference entries come from apocryphx's swift-transformers PR #360, at commit ce847085784bacd8c3c15180c976b17c8ce73e31. The corpus… See the full description on the dataset page: https://huggingface.co/datasets/pcuenq/tokenizer-conformance.textothern<1K0 likes94 downloads20d agoHugging Face03procesaur /sr-tokenizer-test Sr Tokenizer test This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models. It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fields. Dataset Structure Metadata has been stripped; Each record is a JSON object with: id: unique identifier text: raw Serbian text Source coprora Znanje(sr) corpus: ~6.6 GB… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/sr-tokenizer-test.text100K<n<1M0 likes76 downloads5mo agoHugging Face04kacperwikiel /speakleash-tokenizer-5gb-sample SpeakLeash tokenizer 42GB quality sample Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10. Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind. This is intended for tokenizer/BPE training convergence tests. texttext-generation1M<n<10M0 likes67 downloads3mo agoHugging Face05kenny0bi /tokenizer-tax owóorí: the tokenizer tax, measured Token-count premiums for the 204 languages of FLORES-200 under 12 tokenizers, from the identical 1,012 professionally translated sentences. The premium is tokens(language) / tokens(English) on the same content, so it reads directly as a price multiplier for API cost, latency and context shrinkage. Interactive explorer: https://kenny0bi.github.io/owoori/ Method, figures, code: https://github.com/Kenny0bi/owoori Files… See the full description on the dataset page: https://huggingface.co/datasets/kenny0bi/tokenizer-tax.tabularn<1K0 likes39 downloads1mo agoHugging Face06dakoblov /fr-wiki-popular-200-tokenizer-pruning French Wikipedia corpus for tokenizer pruning Prepared by Daniil Koblov. Contains 200 complete plain-text article extracts: 180 training articles and 20 held-out articles, split with Python's random seed 1337. Candidates come from the 2025 monthly French Wikipedia top-1000 pageview lists. Articles must appear in at least three months. Ranking uses month recurrence, then the sum of reciprocal monthly ranks. The first 200 qualifying articles are shuffled and split. Non-article… See the full description on the dataset page: https://huggingface.co/datasets/dakoblov/fr-wiki-popular-200-tokenizer-pruning.tabularn<1K0 likes39 downloads7d agoHugging Face07PaxiAI /Vietnamese-Tokenizer-Corpusgated Vietnamese Tokenizer Training Corpus v1 This dataset is the training corpus used to build PaxiAI/Vietnamese-Tokenizer. It was created by sampling and combining Vietnamese, English, and source-code datasets into an approximately 8 GiB corpus intended specifically for tokenizer training. Its purpose is to provide a diverse and representative sample from which a Vietnamese-focused Byte-level BPE vocabulary can be learned. Dataset Summary The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/PaxiAI/Vietnamese-Tokenizer-Corpus.text1M<n<10M0 likes28 downloads24d agoHugging Face08german-tokenizer-benchmark /co-funer CO-Fun: Tokenized Sentences This datasets hosts a sentence-tokenized version of the CO-Fun: A German Dataset on Company Outsourcing in Fund Prospectuses for Named Entity Recognition and Relation Extraction dataset. Creation The following script can be used to reproduce the creation of the dataset: import flair import json from flair.datasets.sequence_labeling import ColumnCorpus from flair.file_utils import cached_path from pathlib import Path from typing import… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/co-funer.textn<1K0 likes24 downloads11mo agoHugging Face09german-tokenizer-benchmark /german-ler German LER: Tokenized Sentences This datasets hosts a sentence-tokenized version of the German LER dataset. Creation The following script can be used to reproduce the creation of the dataset: import json from flair.datasets import NER_GERMAN_LEGAL corpus = NER_GERMAN_LEGAL() with open("./train.jsonl", "wt") as f_out: for sentence in corpus.train: current_example = { "text": sentence.to_tokenized_string() }… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/german-ler.text10K<n<100K0 likes21 downloads11mo agoHugging Face10noanabeshima /tiny_tokenizer_tokens1M<n<10M0 likes15 downloads2y agoHugging Face11healthyfat /tokenizertext1K<n<10K0 likes15 downloads5mo agoHugging Face12toklens /fw_edu_tokenizer_training_dataThis dataset is a small subset sampled from the 10bt subset of FineWeb-Edu dataset text1M<n<10M0 likes14 downloads2mo agoHugging Face13german-tokenizer-benchmark /germeval14 GermEval 2014: Tokenized Sentences This datasets hosts a sentence-tokenized version of the GermEval 2014 NER dataset. Creation The following script can be used to reproduce the creation of the dataset: import json from flair.datasets import NER_GERMAN_GERMEVAL corpus = NER_GERMAN_GERMEVAL() with open("./germeval14/train.jsonl", "wt") as f_out: for sentence in germeval_corpus.train: current_example = { "text": sentence.to_tokenized_string()… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/germeval14.text10K<n<100K0 likes12 downloads11mo agoHugging Face14german-tokenizer-benchmark /biofid BIOfid: Tokenized Sentences This datasets hosts a sentence-tokenized version of the BIOfid dataset. Creation The following script can be used to reproduce the creation of the dataset: import json from flair.datasets import NER_GERMAN_BIOFID corpus = NER_GERMAN_BIOFID() with open("./train.jsonl", "wt") as f_out: for sentence in corpus.train: current_example = { "text": sentence.to_tokenized_string() }… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/biofid.text10K<n<100K0 likes11 downloads11mo agoHugging Face15german-tokenizer-benchmark /ud-hdt UD German-HDT: Tokenized Sentences This datasets hosts a sentence-tokenized version of the Universal Dependencies German-HDT dataset. Creation The following script can be used to reproduce the creation of the dataset: import json from flair.datasets import UD_GERMAN_HDT corpus = UD_GERMAN_HDT() with open("./train.jsonl", "wt") as f_out: for sentence in corpus.train: current_example = { "text": sentence.to_tokenized_string() }… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/ud-hdt.text100K<n<1M0 likes9 downloads11mo agoHugging Face16aoUTlum /jp-datasets-for-tokenizerstext100K<n<1M0 likes9 downloads6mo agoHugging Face17kamoo-ai /tokenizer-benchmark kamoo tokenizer-benchmark Dezelfde Nederlandse alinea door meerdere tokenizers. Minder tokens = meer context in hetzelfde venster en lagere kosten per antwoord. Reken het na. Dit is het meetlog achter de tokenizer-claims van kamoo.nl. Alles in deze repo is genoeg om de meting zelf te herhalen: de testalinea, het script en de uitkomsten. Meting (2026-07-06) Testalinea: de eerste alinea van het Nederlandse Wikipedia-artikel Nederland (CC-BY-SA, opgehaald 2026-07-06)… See the full description on the dataset page: https://huggingface.co/datasets/kamoo-ai/tokenizer-benchmark.textn<1K0 likes9 downloads3mo agoHugging Face18starsofchance /Qwen_coder_tokenizer-PrimeVul_splitstext10K<n<100K0 likes8 downloads1y agoHugging Face19imsheriff /smarter-llm-tokenizer-datagatedtext100M<n<1B0 likes8 downloads5mo agoHugging Face20german-tokenizer-benchmark /mobie DFKI MobIE: Tokenized Sentences This datasets hosts a sentence-tokenized version of the DFKI MobIE dataset. Creation The following script can be used to reproduce the creation of the dataset: import json from flair.datasets import NER_GERMAN_MOBIE corpus = NER_GERMAN_MOBIE() with open("./train.jsonl", "wt") as f_out: for sentence in corpus.train: current_example = { "text": sentence.to_tokenized_string() }… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/mobie.text1K<n<10K0 likes7 downloads11mo agoHugging Face21simulaXrm /deepwriting-stroke-tokenizer-256gatedtext10K<n<100K0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.