Team Ai
Datasetpublic

pcuenq/tokenizer-conformance

Tokenizer conformance fixtures Reference inputs and Python fast-tokenizer outputs for tokenizer implementations. The initial corpus contains 83 inputs in 30 categories, with 498 reference encodings across six tokenizers. This is a regression dataset, not a model-quality benchmark. Provenance and attribution The input corpus and reference entries come from apocryphx's swift-transformers PR #360, at commit ce847085784bacd8c3c15180c976b17c8ce73e31. The corpus… See the full description on the dataset page: https://huggingface.co/datasets/pcuenq/tokenizer-conformance.

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
0likes94downloads

pcuenq/tokenizer-conformance · main · files are served by the source, never re-hosted here