cymatt/AutoModel-Autotokenizer-App
0
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Tokenizer Learning Notes
Welcome to my Tokenizer learning notes. Here, I will record the principles, usage, examples, and related resources about Hugging Face Tokenizers.
Table of Contents
- Introduction to Tokenizers
- Common Tokenizer Types
- Key Concepts
- Example Code
- Inference Performance between BPE and wordpiece
Introduction to Tokenizers
A tokenizer is a tool that splits text into smaller units called tokens. It is an important preprocessing step in natural language processing, converting raw text into token sequences understandable by models.
Common Tokenizer Types
- Word-level Tokenizer: splits text by spaces or punctuation
- Character-level Tokenizer: splits text into individual characters
- Subword Tokenizer:
- BPE (Byte Pair Encoding)
- WordPiece
- Unigram
Key Concepts
- Vocabulary: the collection of all tokens
- Token ID: the index of a token in the vocabulary
- Encoding: converting text into token IDs
- Decoding: converting token IDs back to text
Example Code
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
text = "Hello, Hugging Face!"
tokens = tokenizer.tokenize(text)
token_ids = tokenizer.encode(text)
print("Tokens:", tokens)
print("Token IDs:", token_ids)Differences Between BPE and WordPiece Tokenizers in Model Inference
This note explains the main differences between Byte Pair Encoding (BPE) and WordPiece tokenizers during model inference and how they can impact performance and output.
Comparison Table
Detailed Explanation
1. Inference Speed Differences
- BPE typically produces longer tokens and fewer total tokens, resulting in slightly faster inference due to shorter input sequences.
- WordPiece produces more fine-grained tokens, increasing token count and slightly increasing computation time during inference.
However, these speed differences are usually minimal on modern hardware.
2. Inference Output Differences
- Both tokenizers yield very similar inference accuracy and semantic understanding since both aim for effective vocabulary coverage.
- WordPiece’s finer tokenization often handles rare words and typos better, potentially producing more stable token sequences at inference.
- BPE is simpler and may underperform slightly on extremely rare or novel words.
3. Practical Recommendations
- Use the tokenizer associated with your pretrained model: GPT-2 uses BPE, BERT uses WordPiece.
- For custom tokenizers or training new models, BPE is simpler to implement, while WordPiece offers more flexibility but is more complex.
- Newer tokenization methods like Unigram and SentencePiece also exist, each with unique inference characteristics.
