Team Ai
Modelpublic

E6E831728/fixed-minimal-binary-code

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
0likes281downloads
README.md199 linesDownload Raw Back to root
1---2license: apache-2.03library_name: transformers4tags:5- causal-lm6- text-generation7- transformer8- decoder-only9- fixed-embeddings10- binary-token-codes11- research12language:13  - en14---15 16# Fixed Minimal Binary Code Model17 18This is an anonymized research checkpoint for the paper:19 20**Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes**21 22## Model variant23 24This repository contains the **fixed minimal binary token-code model**.25 26Instead of a trainable input embedding table, each token ID is represented by its exact minimal binary code.27 28For vocabulary size:29 30```text31V = 65,53632```33 34the minimal injective binary code width is:35 36```text37K = ceil(log2(V)) = 1638```39 40The 16-dimensional binary code is tiled to model width 1024.41 42The model therefore uses:43 44```text450 trainable input-embedding parameters46```47 48The output projection remains standard and trainable.49 50## Architecture51 52- decoder-only Transformer53- vocabulary size: 65,53654- model width: 102455- number of layers: 3256- number of attention heads: 3257- context length: 102458- rotary positional embeddings59- GELU activations60- untied trainable output projection61 62## Loading example63 64```python65import torch66from transformers import AutoTokenizer, AutoModelForCausalLM67 68repo_id = "E6E831728/fixed-minimal-binary-code"69 70tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)71model = AutoModelForCausalLM.from_pretrained(repo_id, trust_remote_code=True)72model.eval()73 74prompt = "Question: What is the capital of France?\nAnswer:"75input_ids = torch.tensor([tokenizer.encode(prompt)], dtype=torch.long)76 77with torch.no_grad():78    output_ids = model.generate(input_ids, max_new_tokens=3, do_sample=False)79 80print(tokenizer.decode(output_ids[0].tolist()))81```82 83## Standardized base-model evaluation84 85The checkpoint was evaluated as a base causal language model with86EleutherAI LM Evaluation Harness `v0.4.10`.87 88Evaluation protocol:89 90- Hugging Face backend: `hf`91- maximum context length: 1,02492- `add_bos_token=False`93- no chat template94- deterministic likelihood-based evaluation95- harness seeds: `0,1234,1234,1234`96- base checkpoints only; no SFT or instruction checkpoints97 98| Metric | Learned input table | Fixed Binary-16 | Affine GF(2), table-free | SmolLM2-135M | SmolLM2-360M |99|---|---:|---:|---:|---:|---:|100| HellaSwag acc | 28.49 ± 0.45 | 29.04 ± 0.45 | 29.04 ± 0.45 | 35.36 ± 0.48 | 43.05 ± 0.49 |101| HellaSwag acc_norm | 31.32 ± 0.46 | 32.32 ± 0.47 | 31.80 ± 0.46 | 43.02 ± 0.49 | 56.28 ± 0.50 |102| ARC-Easy acc | 46.38 ± 1.02 | 47.90 ± 1.03 | 47.64 ± 1.02 | 64.44 ± 0.98 | 70.24 ± 0.94 |103| ARC-Easy acc_norm | 40.70 ± 1.01 | 40.87 ± 1.01 | 41.20 ± 1.01 | 58.75 ± 1.01 | 68.18 ± 0.96 |104| ARC-Challenge acc | 20.39 ± 1.18 | 19.62 ± 1.16 | 21.33 ± 1.20 | 28.07 ± 1.31 | 36.26 ± 1.40 |105| ARC-Challenge acc_norm | 25.85 ± 1.28 | 26.19 ± 1.28 | 24.83 ± 1.26 | 29.61 ± 1.33 | 38.05 ± 1.42 |106| PIQA acc | 62.35 ± 1.13 | 62.57 ± 1.13 | 62.68 ± 1.13 | 68.44 ± 1.08 | 71.38 ± 1.05 |107| PIQA acc_norm | 60.61 ± 1.14 | 62.08 ± 1.13 | 60.94 ± 1.14 | 68.39 ± 1.08 | 71.82 ± 1.05 |108| WinoGrande acc | 50.20 ± 1.41 | 50.12 ± 1.41 | 50.43 ± 1.41 | 52.57 ± 1.40 | 59.35 ± 1.38 |109| OpenBookQA acc | 18.40 ± 1.73 | 17.20 ± 1.69 | 17.60 ± 1.70 | 22.00 ± 1.85 | 24.80 ± 1.93 |110| OpenBookQA acc_norm | 29.20 ± 2.04 | 31.00 ± 2.07 | 29.40 ± 2.04 | 32.60 ± 2.10 | 37.80 ± 2.17 |111| CommonsenseQA acc | 20.31 ± 1.15 | 19.90 ± 1.14 | 20.23 ± 1.15 | 19.90 ± 1.14 | 21.05 ± 1.17 |112| MMLU 0-shot | 24.13 ± 0.36 | 23.86 ± 0.36 | 24.11 ± 0.36 | 24.24 ± 0.36 | 25.47 ± 0.37 |113| MMLU 5-shot | 25.68 ± 0.37 | 25.60 ± 0.37 | 25.66 ± 0.37 | 25.39 ± 0.37 | 25.05 ± 0.37 |114| LAMBADA accuracy | 22.38 ± 0.58 | 21.23 ± 0.57 | 21.99 ± 0.58 | 42.97 ± 0.69 | 53.31 ± 0.70 |115| LAMBADA perplexity | 95.14 ± 4.01 | 101.74 ± 4.27 | 100.61 ± 4.17 | 19.06 ± 0.63 | 9.38 ± 0.27 |116| WikiText word perplexity | 81.04 | 74.87 | 76.17 | 25.53 | 18.84 |117| WikiText byte perplexity | 2.27 | 2.24 | 2.25 | 1.83 | 1.73 |118| WikiText bits/byte | 1.19 | 1.16 | 1.17 | 0.87 | 0.79 |119 120The three paper checkpoints form the controlled architectural comparison.121SmolLM2-135M and SmolLM2-360M are external reference models, not matched122baselines: they use different architectures, tokenizers, training mixtures,123and much larger pretraining budgets. SmolLM2-135M was trained on approximately1242T tokens and SmolLM2-360M on approximately 4T tokens, whereas the paper125checkpoints saw approximately 16–17B tokens. Their scores therefore provide126context for absolute capability and must not be interpreted as isolating the127effect of the input parameterization.128 129Perplexity values should be interpreted especially cautiously across different130tokenizers. The primary controlled comparison is among the three paper models,131which share the same tokenizer, data pipeline, and architecture.132 133 134## Input-interface audit135 136This checkpoint stores a deterministic `65,536 × 16` binary codebook as a137frozen `nn.Embedding` for computational convenience. During training, the138table was initialized from the fixed token codes, marked with139`requires_grad=False`, and excluded from the optimizer. It therefore contained1401,048,576 stored but non-trainable values and contributed zero trainable input141parameters.142 143The released checkpoint can be audited directly:144 145```python146import torch147from transformers import AutoModelForCausalLM148 149repo_id = "E6E831728/fixed-minimal-binary-code"150 151model = AutoModelForCausalLM.from_pretrained(152    repo_id,153    trust_remote_code=True,154    torch_dtype=torch.float32,155).cpu().eval()156 157embedding = model.get_input_embeddings()158weight = embedding.weight.detach()159 160vocab_size, code_bits = weight.shape161ids = torch.arange(vocab_size, dtype=torch.long)162positions = torch.arange(code_bits, dtype=torch.long)163 164expected = ((ids[:, None] >> positions[None, :]) & 1).float()165expected[model.config.pad_token_id].zero_()166 167print("shape:", tuple(weight.shape))168print("unique values:", torch.unique(weight).tolist())169print("all entries binary:", bool(torch.all((weight == 0) | (weight == 1))))170print("exact canonical-code match:", bool(torch.equal(weight, expected)))171print("mismatching entries:", int((weight != expected).sum().item()))172 173assert tuple(weight.shape) == (65536, 16)174assert torch.all((weight == 0) | (weight == 1))175assert torch.equal(weight, expected)176```177 178Expected audit properties:179 180```text181shape: (65536, 16)182unique values: [0.0, 1.0]183all entries binary: True184exact canonical-code match: True185mismatching entries: 0186```187 188 189## Intended use190 191This checkpoint is provided for anonymous review and reproducibility of the paper's main claim: a trainable input embedding table is not necessary for useful language modeling in the studied regime.192 193## Limitations194 195This model is a research checkpoint. It is not intended for deployment. It may produce incorrect, biased, unsafe, or nonsensical outputs.196 197## Training data198 199The model was trained on the same FineWeb-Edu + Cosmopedia mixture used for the matched comparisons in the paper. Dataset terms and licenses are those of the original datasets.