Team Ai
Modelpublic

sajalmadan09/bert-tiny-native-cpp

sourceHugging Facemitupdated 25d agoView on Hugging Face
2likes147downloads
README.md144 linesDownload Raw Back to root
1---2license: mit3language:4- en5tags:6- bert7- native-inference8- sajal-labs9- transformer10- feature-extraction11base_model: prajjwal1/bert-tiny12base_model_relation: finetune13pipeline_tag: feature-extraction14---15 16# bert-tiny — Native C++ Port (Sajal Labs)17 18**This is not a new model.** The weights, architecture, and pretraining are19entirely [`prajjwal1/bert-tiny`](https://huggingface.co/prajjwal1/bert-tiny) by Prajjwal20Bhargava (MIT license) — a compact pretrained BERT encoder introduced in21Turc et al. 2019 ("Well-Read Students Learn Better") and ported to22Hugging Face for Bhargava et al. 2021 ("Generalization in NLI"). **Please23cite both papers if you use this model** (citations below).24 25**What Sajal Labs added**: a from-scratch native C++ port of the encoder26and the WordPiece tokenizer — no PyTorch, no `transformers`, no Python at27inference time — with rigorous equivalence and benchmark validation against28the original. See [Sajal Labs](https://github.com/Sajalmadan09/sajal-labs),29experiments30[exp11](https://github.com/Sajalmadan09/sajal-labs/tree/main/research/experiments/exp11-real-pretrained-transformer)31through32[exp14](https://github.com/Sajalmadan09/sajal-labs/tree/main/research/experiments/exp14-wordpiece-native),33for full methodology.34 35## Model details (unchanged from the original)36 37- **Architecture**: BERT encoder, 2 layers, hidden=128, heads=2, intermediate=51238- **Vocabulary**: 30522 WordPiece tokens (bert-base-uncased vocab)39- **Parameters**: 4,385,92040- **Precision**: fp3241- **Base model license**: MIT (prajjwal1/bert-tiny)42- **Port license**: MIT (Sajal Labs' C++ code)43 44## What was verified (Sajal Labs' contribution)45 46**Tokenizer**: 17/17 real test sentences — including contractions ("don't"),47hyphenation ("COVID-19"), an out-of-vocabulary word forcing an 11-piece48subword split, and an email address — produced **byte-identical token IDs**49to the original `BertTokenizerFast`. This is an exact-match bar, not a50tolerance: tokenization is deterministic.51 52**Encoder**: max absolute error 9.54e-06 (hidden states),532.19e-06 (pooled `[CLS]` output), cosine similarity54~1.0, across 10 real sentences of varying length (4-25 tokens) — fp32-scale55agreement, consistent with floating-point non-associativity between two56independent implementations (not a bug; see the repo's `research/papers.md`).57 58## Benchmark (single request, Apple M4 CPU — full data in `benchmark_results.json`)59 60True end-to-end cold invocation (process spawn → raw text in → prediction out61→ process exit, external wall-clock). **Primary comparison: native vs. ONNX62Runtime paired with the lean, standalone `tokenizers` library — the63best-case Python deployment, not the easiest target to beat:**64 65| Implementation | Cold invocation p50 |66|---|---:|67| **Native C++ (Sajal runtime)** | 11.37ms |68| **ONNX Runtime + lean tokenizer** | **95.72ms (8.4x slower)** |69| ONNX Runtime + 🤗 transformers tokenizer | 2495.61ms (219.6x slower) |70| PyTorch + 🤗 transformers | 4973.91ms (437.6x slower) |71 72The last two rows are real and worth knowing, but they're the easy targets73(heavier Python stacks) — the 8.4x number above is74the one that holds up against someone who already optimized their Python75deployment correctly. Worth knowing before you read too much into "ONNX76Runtime" as a single number: its own cold-start77number depends heavily on which tokenizer library it's paired with — using78`transformers` for convenience costs ~25x more than using the lean, standalone79`tokenizers` library for the exact same token IDs. Native sidesteps that80whole dependency-choice question by construction. Full discussion in81[exp14](https://github.com/Sajalmadan09/sajal-labs/tree/main/research/experiments/exp14-wordpiece-native).82 83**Honest scope note on the warm-loop numbers**: native's advantage is84*not* unconditional the way cold-invocation is — Sajal Labs found it85depends on model width (`hidden_size`), with a measured crossover around86`hidden≈250` on this hardware87([exp12](https://github.com/Sajalmadan09/sajal-labs/tree/main/research/experiments/exp12-warm-latency-width-depth)/[exp13](https://github.com/Sajalmadan09/sajal-labs/tree/main/research/experiments/exp13-width-threshold)).88This model's `hidden=128`89sits comfortably below that, so native keeps a real warm-loop edge too —90but that's a property of this model's size, not a general claim.91 92## How to use93 94### Native (Sajal runtime, zero Python)95 96```bash97sajal run <this_directory> "The quick brown fox jumps over the lazy dog."98```99 100### PyTorch / transformers (the original)101 102```python103from transformers import BertModel, BertTokenizerFast104model = BertModel.from_pretrained("prajjwal1/bert-tiny")105tokenizer = BertTokenizerFast.from_pretrained("prajjwal1/bert-tiny")106```107 108### ONNX Runtime109 110```python111import onnxruntime as ort112session = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])113# feed input_ids/attention_mask/token_type_ids from any WordPiece tokenizer114```115 116## Citations (required if you use this model)117 118```bibtex119@article{turc2019distillation,120  title={Well-Read Students Learn Better: On the Importance of Pre-training Compact Models},121  author={Turc, Iulia and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},122  journal={arXiv preprint arXiv:1908.08962v2},123  year={2019}124}125 126@misc{bhargava2021generalization,127  title={Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics},128  author={Bhargava, Prajjwal and Drozd, Aleksandr and Rogers, Anna},129  year={2021},130  eprint={2110.01518},131  archivePrefix={arXiv},132  primaryClass={cs.CL}133}134```135 136## Files137 138- `model.safetensors` — the original weights, HF/PyTorch-ecosystem format139- `model.onnx` + `model.onnx.data` — ONNX Runtime-compatible export (weights externalized to the `.data` file; both are required together)140- `word_embeddings.bin`, `position_embeddings.bin`, `token_type_embeddings.bin`, `emb_ln_*.bin`, `layer{i}_*.bin`, `pooler_*.bin` — raw native Sajal runtime format141- `vocab.txt`, `tokenizer_config.txt` — WordPiece vocabulary + config for the native tokenizer port142- `config.json` — architecture metadata143- `benchmark_results.json` — full machine-readable benchmark/equivalence data behind the numbers above144