pcuenq/tokenizer-conformance
Tokenizer conformance fixtures Reference inputs and Python fast-tokenizer outputs for tokenizer implementations. The initial corpus contains 83 inputs in 30 categories, with 498 reference encodings across six tokenizers. This is a regression dataset, not a model-quality benchmark. Provenance and attribution The input corpus and reference entries come from apocryphx's swift-transformers PR #360, at commit ce847085784bacd8c3c15180c976b17c8ce73e31. The corpus… See the full description on the dataset page: https://huggingface.co/datasets/pcuenq/tokenizer-conformance.
Tokenizer conformance fixtures
Reference inputs and Python fast-tokenizer outputs for tokenizer implementations. The initial corpus contains 83 inputs in 30 categories, with 498 reference encodings across six tokenizers. This is a regression dataset, not a model-quality benchmark.
Provenance and attribution
The input corpus and reference entries come from apocryphx's swift-transformers PR #360, at commit `ce847085784bacd8c3c15180c976b17c8ce73e31`. The corpus originated in ObjCTokenizer and diagnosed the Unicode tokenization bugs documented in swift-transformers #352. Daisuke Majima (john-rocky) contributed test-design ideas in PR #357, including stored decoded forms and the ungated TinyLlama reference model. The original contributor documented AI assistance in the linked issue and PR. The Apache 2.0 license from the source repository is included in LICENSE.
The 83 input records are unchanged. All six original baseline entry arrays were reproduced exactly with the versions and model commits listed here before publication. Baseline metadata was expanded to record those commits and the Rust tokenizers version. Tools/generate_tokenizer_baselines.py is adapted from the contributor's generator.
Layout
manifest.json: corpus IDs, input paths, and baseline paths with model IDs and immutable model revisions.multilingual/inputs.json: records{id, category, text}; IDs are stable within the corpus. Preserve text exactly, including combining marks and whitespace.multilingual/baselines/*.json:{metadata, entries}for each tokenizer.Tools/generate_tokenizer_baselines.pyandTools/requirements.txt: regeneration and verification.
Each baseline entry has id, input_ids, tokens, decoded_with_special, and decoded_skip_special. Metadata records model_id, model_revision, transformers_version, tokenizers_version, generated_at, and input_count. Generation uses transformers.AutoTokenizer with use_fast=True and add_special_tokens=True. The Swift conformance test currently compares token IDs; tokens aid diagnostics and decoded forms are retained for future decoder tests.
Reproduce or verify
From a checkout or downloaded snapshot of this dataset, using Python 3.12 and uv:
uv run --python 3.12 --with-requirements Tools/requirements.txt python Tools/generate_tokenizer_baselines.py --check
uv run --python 3.12 --with-requirements Tools/requirements.txt python Tools/generate_tokenizer_baselines.py--check compares every generated entry (IDs, tokens and both decoded forms), returns nonzero on a mismatch, and writes nothing. It ignores metadata such as the generation timestamp. Omit --check to regenerate all baselines with fresh provenance metadata. Models are loaded at the immutable revisions in manifest.json, never implicitly at main. The first run downloads tokenizer files; model weights are not needed.
Swift consumption and expansion
Download a pinned dataset commit with HubApi.snapshot(from: Hub.Repo(id: "pcuenq/tokenizer-conformance", type: .datasets), revision: ..., matching: "*.json"). The Swift suite reads the manifest and validates unique IDs, exact corpus coverage, model revisions, and reference lengths before comparing output. Hub caching avoids downloading unchanged files on each run. Network or malformed-data errors fail tests rather than silently skipping coverage.
To add cases, append stable IDs to a corpus and regenerate all its baselines. To add a model or a separate corpus, extend the manifest and run the generator. Review the reference diff, publish a new dataset commit, and explicitly update the consumer's pinned revision.
Do not replace Python expectations with Swift output. Implementation-specific known divergences belong in the consuming test suite. At initial verification, Swift matched 490/498 encodings; the eight already tracked in PR #360 remain (three Qwen Thai cases and five TinyLlama whitespace cases). All entries remain in this dataset, including those eight, with the original Python expectations.
