RNDRandoM/gemma-4-e2b-pruning-exp
Gemma 4 E2B-it, vocabulary pruned to English (m = 1)
Homework 1 (tokenizer pruning), Modern Methods and Algorithms of Generative AI, Skoltech, Fall 2026.
google/gemma-4-E2B-it with its vocabulary pruned from 262,144 to 103,351 tokens (padded to 103,360 rows): every token that occurs at least once when 12 MB of English Wikipedia (wikimedia/wikipedia, 20231101.en) is tokenized, plus all special/added tokens, the tokenizer's base alphabet, and the merge closure (route closure: every piece formed while BPE builds a kept token). No training: the input embedding (tied with the output head) and the per-layer embedding table are sliced with the id map (pruning.json: new id -> old id); all token-id fields in the configs are remapped.
Loads with transformers 5.14 (AutoModelForImageTextToText, AutoTokenizer). Rows 103,351-103,359 are zero padding: suppress them when generating. Text outside English Wikipedia's vocabulary (other scripts, some markdown, ALL-CAPS words) is spelled with smaller pieces or bytes and is generated less reliably.
