hf-internal-testing/tokenizers-test-data
tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.
Add ModernBERT-base tokenizer fixture (#13)
Add siglip-base-patch16-224 and mistral-7b-v0.1 oracle fixtures (#12)
Add reproducible tokbench v1 model inputs
Add sentence-transformers all-MiniLM-L6-v2 and all-mpnet-base-v2 tokenizer.json fixtures (#11)
Add meta-llama/Llama-3.2-1B tokenizer.json fixture (#10)
Make bench tokenizer configs byte-identical to their source models (#8)
xet-track llama-2.json and gpt2.json (#7)
Upload fixtures/modalities/added_normalized_dense.txt with huggingface_hub
Upload fixtures/modalities/added_normalized_sparse.txt with huggingface_hub
Upload fixtures/modalities/added_special_dense.txt with huggingface_hub
Upload fixtures/modalities/added_special_sparse.txt with huggingface_hub
Add llama-2-7b-chat-hf tokenizer.json (canonical, array-safe merges) for pipeline oracle (#6)
Xet-track full tokenizer configs (pipeline bench) (#5)
Add README (repo layout) and xet/LFS rule for fixtures/ (#4)
Add slim byte-level tokenizer fixtures for tokenizers byte_level_fast tests (#3)
feat: upload `xnli.txt` to assert validity of the whitespace split (#2)
Add all tokenizers test artifacts
initial commit
