Team Ai
Modelpublic

Bochkov/ab_ext_binary16

sourceHugging Faceupdated 3d agoView on Hugging Face
0likes568downloads
README.md339 linesDownload Raw Back to root
1---2library_name: transformers3pipeline_tag: text-generation4language:5  - en6tags:7  - causal-lm8  - base-model9  - custom-code10  - safetensors11  - research12  - fixed-token-codes13  - frozen-input-representations14---15 16# AB-EXT Binary16 — 1.711B parameters, 100B-token target17 18A **base pretrained decoder-only causal language model** released for19research on trainable input embedding tables and fixed token identities.20 21It is not an instruction-tuned or preference-optimized assistant.22 23## Research question24 25Can a shared contextual network learn useful language-modeling behavior26without independently trainable token-specific input vectors?27 28The controlled family contains a learned-input model, a canonical2916-bit-code model, and an invertibly recoded GF2 model. They share the30contextual backbone and output-head architecture, but not the same total31trainable parameter count.32 33**The results support viability, not performance equivalence.**34The fixed-code models retain substantial capability while the learned35model performs better on several informative evaluations.36 37## Model specification38 39| Property | Value |40|---|---|41| Input mode | `binary16` |42| Trainable parameters | 1,711,376,384 |43| Trainable input parameters | 0 |44| Trainable body parameters, excluding input/output | 1,610,713,088 |45| Trainable untied output-head parameters | 100,663,296 |46| Persistent input-buffer values | 786,432 |47| Hidden width | 2048 |48| Decoder blocks | 24 |49| Attention heads | 32 |50| FFN intermediate width | 8192 |51| Training context | 2048 tokens |52| Position encoding | RoPE |53| Normalization / activation | RMSNorm / SwiGLU |54| Tokenizer source | `HuggingFaceTB/SmolLM2-1.7B` |55| Exported tokenizer revision | effd688a12921b4cc83e3312b6feb579f70f9c71 |56| Evaluated training runs for this interface | One |57 58The stored tensor-value count includes buffers and must not be reported59as the trainable parameter count.60 61## Input representation62 63 64Each token ID is represented by its canonical little-endian binary code:65 66$$67c(t)_j =68\left\lfloor \frac{t}{2^j} \right\rfloor \bmod 2,69\qquad j=0,\ldots,15.70$$71 72Since the vocabulary contains 49,152 entries, an injective fixed-length73binary code requires 16 bits.74 75The code is repeated 128 times to width 2048:76 77$$78x(t)=79\underbrace{c(t)\Vert\cdots\Vert c(t)}_{128\text{ copies}}.80$$81 82There are **zero trainable input-interface parameters** and no additional83trainable input projection before the standard backbone.84The backbone and the untied output vocabulary projection remain trainable.85 86The evaluated implementation stores the 49,152-by-16 codebook as a87persistent non-trainable buffer. It is not literally lookup-free.88“Minimal” describes fixed-length binary identity width, not the storage89of the complete model or an entropy-optimal token code.90 91 92 93## Training94 95- **Training tokens:** Approximately 100B according to the run report; target budget 100,000,000,000 prediction targets. 96- **Training precision:** FP32 parameters with BF16 autocast in the supplied trainer.97- **Reported recipe:** AdamW; peak learning rate 0.00015; minimum98  scheduled learning rate 0.00001; 2000 warmup steps; cosine decay;99  weight decay 0.01; betas 0.9 and 0.95; gradient clipping 1.0.100- **Reported launch geometry:** two GPUs per run, microbatch eight per101  GPU, eight accumulation steps, sequence length 2048.102 103The launch geometry corresponds to 262,144 prediction targets per104optimizer step. Exact final counts must come from the checkpoint,105not from the requested budget.106 107The supplied sampler selects within-document windows from eligible108documents of at least 2049 tokens. Sampling can repeat or overlap109windows; 100B processed targets does not imply 100B unique corpus tokens.110The original trainer does not fully restore per-rank sampling state111on resume. A shared recipe alone does not establish identical realized112sample order across interrupted runs.113 114The model weights were NOT initialized from SmolLM2.115SmolLM2 supplies tokenizer artifacts, not pretrained model weights.116 117## Evaluation results118 119These scores are transcribed from the supplied completed evaluation120summary; the generator does not rerun benchmarks. Raw, unrounded121harness outputs remain authoritative.122 123Accuracy entries are percentages. Their reported `±` values are124evaluation standard errors, **not variation across training seeds**.125Perplexities and bits per byte are not percentages.126 127| Metric | Shots | Result |128|---|---:|---:|129| HellaSwag acc (%) | 0 | 40.70 ± 0.49 |130| HellaSwag acc_norm (%) | 0 | 52.40 ± 0.50 |131| ARC-Easy acc (%) | 0 | 67.59 ± 0.96 |132| ARC-Easy acc_norm (%) | 0 | 61.53 ± 1.00 |133| ARC-Challenge acc (%) | 0 | 32.34 ± 1.37 |134| ARC-Challenge acc_norm (%) | 0 | 34.04 ± 1.38 |135| PIQA acc (%) | 0 | 70.51 ± 1.06 |136| PIQA acc_norm (%) | 0 | 71.11 ± 1.06 |137| WinoGrande acc (%) | 0 | 55.33 ± 1.40 |138| OpenBookQA acc (%) | 0 | 28.20 ± 2.01 |139| OpenBookQA acc_norm (%) | 0 | 38.00 ± 2.17 |140| CommonsenseQA acc (%) | 0 | 20.56 ± 1.16 |141| MMLU acc (%) | 0 | 25.88 ± 0.37 |142| MMLU acc (%; some prompts truncated) | 5 | 25.57 ± 0.37 |143| LAMBADA accuracy (%) | 0 | 42.75 ± 0.69 |144| LAMBADA perplexity ↓ | 0 | 17.91 ± 0.62 |145| WikiText word perplexity ↓ | — | 18.58 |146| WikiText byte perplexity ↓ | — | 1.73 |147| WikiText bits/byte ↓ | — | 0.79 |148 149### Audit and coverage limitations150 151The supplied audit reports identical sample/prompt multisets across152all six models in each completed task group.153 154For MMLU 5-shot, **1,508 / 56,168 candidate log-likelihood requests**155were marked as truncated for each model, approximately 2.68%.156These are candidate requests, not necessarily distinct questions.157The displayed MMLU 5-shot score therefore includes truncated prompts.158 159No truncations were reported for the other groups by that audit.160For WikiText rolling likelihood, this does not mean that whole documents161fit into one model context: rolling windowing is part of scoring.162 163<details>164<summary>Full six-model comparison</summary>165 166| Metric | AB-EXT Learned | AB-EXT Binary16 | AB-EXT GF2 | SmolLM2-135M | SmolLM2-360M | SmolLM2-1.7B |167|---|---:|---:|---:|---:|---:|---:|168| HellaSwag acc (%); shots=0 | 44.21 ± 0.50 | 40.70 ± 0.49 | 40.24 ± 0.49 | 35.36 ± 0.48 | 43.05 ± 0.49 | 53.38 ± 0.50 |169| HellaSwag acc_norm (%); shots=0 | 57.79 ± 0.49 | 52.40 ± 0.50 | 51.44 ± 0.50 | 43.02 ± 0.49 | 56.28 ± 0.50 | 71.43 ± 0.45 |170| ARC-Easy acc (%); shots=0 | 71.63 ± 0.92 | 67.59 ± 0.96 | 66.84 ± 0.97 | 64.44 ± 0.98 | 70.24 ± 0.94 | 77.86 ± 0.85 |171| ARC-Easy acc_norm (%); shots=0 | 66.04 ± 0.97 | 61.53 ± 1.00 | 60.73 ± 1.00 | 58.75 ± 1.01 | 68.18 ± 0.96 | 73.36 ± 0.91 |172| ARC-Challenge acc (%); shots=0 | 35.92 ± 1.40 | 32.34 ± 1.37 | 30.55 ± 1.35 | 28.07 ± 1.31 | 36.26 ± 1.40 | 44.37 ± 1.45 |173| ARC-Challenge acc_norm (%); shots=0 | 37.63 ± 1.42 | 34.04 ± 1.38 | 34.22 ± 1.39 | 29.61 ± 1.33 | 38.05 ± 1.42 | 47.27 ± 1.46 |174| PIQA acc (%); shots=0 | 72.69 ± 1.04 | 70.51 ± 1.06 | 71.16 ± 1.06 | 68.44 ± 1.08 | 71.38 ± 1.05 | 76.99 ± 0.98 |175| PIQA acc_norm (%); shots=0 | 72.14 ± 1.05 | 71.11 ± 1.06 | 72.14 ± 1.05 | 68.39 ± 1.08 | 71.82 ± 1.05 | 77.20 ± 0.98 |176| WinoGrande acc (%); shots=0 | 58.56 ± 1.38 | 55.33 ± 1.40 | 55.01 ± 1.40 | 52.57 ± 1.40 | 59.35 ± 1.38 | 65.98 ± 1.33 |177| OpenBookQA acc (%); shots=0 | 27.60 ± 2.00 | 28.20 ± 2.01 | 25.20 ± 1.94 | 22.00 ± 1.85 | 24.80 ± 1.93 | 32.20 ± 2.09 |178| OpenBookQA acc_norm (%); shots=0 | 37.80 ± 2.17 | 38.00 ± 2.17 | 36.80 ± 2.16 | 32.60 ± 2.10 | 37.80 ± 2.17 | 44.40 ± 2.22 |179| CommonsenseQA acc (%); shots=0 | 19.82 ± 1.14 | 20.56 ± 1.16 | 19.74 ± 1.14 | 19.90 ± 1.14 | 21.05 ± 1.17 | 41.69 ± 1.41 |180| MMLU acc (%); shots=0 | 25.32 ± 0.37 | 25.88 ± 0.37 | 26.11 ± 0.37 | 24.25 ± 0.36 | 25.47 ± 0.37 | 48.40 ± 0.41 |181| MMLU acc (%; some prompts truncated); shots=5 | 25.48 ± 0.37 | 25.57 ± 0.37 | 24.66 ± 0.36 | 25.15 ± 0.36 | 25.03 ± 0.37 | 50.06 ± 0.41 |182| LAMBADA accuracy (%); shots=0 | 47.72 ± 0.70 | 42.75 ± 0.69 | 42.29 ± 0.69 | 42.97 ± 0.69 | 53.31 ± 0.70 | 67.51 ± 0.65 |183| LAMBADA perplexity ↓; shots=0 | 12.88 ± 0.42 | 17.91 ± 0.62 | 18.47 ± 0.63 | 19.06 ± 0.63 | 9.38 ± 0.27 | 4.44 ± 0.10 |184| WikiText word perplexity ↓; shots=— | 16.50 | 18.58 | 19.03 | 23.14 | 17.12 | 11.62 |185| WikiText byte perplexity ↓; shots=— | 1.69 | 1.73 | 1.73 | 1.80 | 1.70 | 1.58 |186| WikiText bits/byte ↓; shots=— | 0.76 | 0.79 | 0.79 | 0.85 | 0.77 | 0.66 |187 188</details>189 190### How to interpret SmolLM2 comparisons191 192All scores above are from the supplied local evaluation summary, not193copied leaderboard scores.194 195The SmolLM2 technical report gives approximate training budgets of:196 197| External reference | Published budget | Relative to 100B |198|---|---:|---:|199| SmolLM2-135M | 2T tokens | 20× |200| SmolLM2-360M | 4T tokens | 40× |201| SmolLM2-1.7B | 11T tokens | 110× |202 203Source: https://arxiv.org/abs/2502.02737204 205These models differ in architecture, size, data, training schedule,206and compute. They are quality references, **not matched controls** and207not proof of a sample-efficiency advantage.208 209## Usage210 211Review the custom Python files before enabling `trust_remote_code=True`.212Use a tested Transformers version and pin the Hub revision for213reproducible deployment.214 215```python216import torch217from transformers import AutoTokenizer, AutoModelForCausalLM218 219model_id = 'Bochkov/ab_ext_binary16'220# For published Hub use, pin revision to a reviewed commit.221revision = None222 223tokenizer = AutoTokenizer.from_pretrained(224    model_id,225    revision=revision,226)227 228model = AutoModelForCausalLM.from_pretrained(229    model_id,230    revision=revision,231    trust_remote_code=True,232    dtype=torch.bfloat16,233).to("cuda").eval()234 235inputs = tokenizer(236    "Gravity is",237    return_tensors="pt",238    add_special_tokens=False,239    return_attention_mask=True,240).to("cuda")241 242pad_id = tokenizer.pad_token_id243if pad_id is None:244    pad_id = tokenizer.eos_token_id245 246with torch.inference_mode():247    output = model.generate(248        input_ids=inputs["input_ids"],249        attention_mask=inputs["attention_mask"],250        max_new_tokens=32,251        do_sample=False,252        use_cache=False,253        eos_token_id=tokenizer.eos_token_id,254        pad_token_id=pad_id,255    )256 257print(tokenizer.decode(output[0], skip_special_tokens=True))258```259 260The implementation does not provide a KV cache.261The trained context is 2048 tokens; the supplied generation adapter262uses a sliding window when the context grows beyond its limit.263This is not evidence of trained long-context capability.264 265### Loss API266 267The original training model consumes already-shifted targets.268The HF runtime is intended to expose the usual causal-LM convention269with an internal label shift. Do not pass already-shifted labels to270such a runtime.271 272Before fine-tuning, verify the actual runtime's loss implementation.273Forward-logit equivalence does not by itself test label conventions.274 275## Verification and integrity276 277The supplied verification logs report:278 279- successful BF16 loading and generation for all three releases;280- exactly matching original/exported forward logits on four short281  prompts for each model, in the tested verification configuration.282 283These are smoke and implementation-parity checks, not an exhaustive284test across padding, context lengths, dtypes, or generation modes.285 286- Weight file: `model.safetensors`287- Weight SHA-256: `c3e78f24f03dff36b6174985dc5189e8ce37a52685d022ade36fa68c3239801b`288- Stored tensor values: 1,712,162,816289- Stored values by dtype: `{"F32": 1712162816}`290- Trainable parameter count: 1,711,376,384291- Persistent input-buffer values: 786,432292 293This card update does not modify the weights, tokenizer, model code,294or configuration.295 296## Limitations and intended use297 298- Research use and text completion; not a validated high-stakes assistant.299- One evaluated training run per input interface at this scale.300- Fixed-code and learned-input models are backbone-matched, not301  total-parameter-matched.302- No measured runtime or energy advantage is established by parameter303  counts alone.304- One GF2 recoding does not establish invariance to arbitrary codes.305- The output vocabulary matrix remains trainable and token-specific.306- Input-code structure is not fitted to the pretraining objective, but307  the tokenizer and its ID assignment can contain corpus-derived structure.308- Benchmark contamination has not been independently certified absent.309- Generated text can be false, biased, or harmful.310 311- Zero/one coding maps token ID zero to a zero input vector; the312  zero-offset GF2 transform preserves it. In the supplied bias-free313  architecture, a context made entirely of zero-code tokens gives314  uniform logits. This does not apply to arbitrary contexts ending315  in that token.316 317 318## Attribution and licensing319 320Tokenizer artifacts are sourced from `HuggingFaceTB/SmolLM2-1.7B`.321 322---323 324## 🧑‍🔬 Citation & Concept325 326If you use this model or the underlying concepts in your research, please cite our work:327 328```329@misc{bochkov2026languagemodelsneedtrainable,330      title={Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale}, 331      author={A. Bochkov},332      year={2026},333      eprint={2610.04002},334      archivePrefix={arXiv},335      primaryClass={cs.CL},336      url={https://arxiv.org/abs/2610.04002}, 337}338```339