Tamkimd/tamev-nano-flash-tinybert
TAMEV-Nano-Flash-TinyBERT — Tiny Decision Model & LLM Router
<img src="banner.png" alt="TAMEV-Nano-Flash-TinyBERT — Tiny decision model: 77.5% top-1, 14,390,185 released parameters, 4.00 ms p50 on a single CPU thread" width="100%">
Scores your candidate options instead of generating text. Send a state and a list of options — a tool menu, a queue list, an intent catalogue, a rubric — and get back calibrated probabilities over those options in one forward pass, on CPU. choice (one of K), noul (yes/no) and score (ordinal) go through the same call.

top-1 `77.45%` · ECE `0.0191` · `4.00 ms` p50 on a single CPU thread · `0.0 (bit-exact)` drift under option permutation.
⚡ At a glance
This tier is self-contained: the backbone and pointer head are already fused in model.safetensors.🚀 Quick start
from transformers import AutoModel, AutoTokenizer
model_id = "Tamkimd/tamev-nano-flash-tinybert"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True)
# options are an argument of the call -- the class list is not baked into the weights
result = model.predict_decision( # `model.decide(...)` is an alias
state="My transaction on a Visa card was rejected while I was travelling in Tokyo.",
question="Which service queue should handle this incident?",
options=["verify_travel_unblock", "file_fraud_dispute", "replace_damaged_card", "branch_appointment"],
tokenizer=tokenizer,
)
print(result["best_option"], f"{result['confidence']:.0%}") # -> the queue to route to
print(result["probabilities"]) # calibrated, sums to 1Requirements: torch>=2.4, transformers>=4.40, Python ≥ 3.10.
Typed questions in one call — choice / noul / score
The tamev package is not on PyPI yet; install it from the repository:
from tamev import Choice, Noul, Score, TypeSafeDirectClient
with TypeSafeDirectClient(model_name="TAMEV-Nano-Flash-TinyBERT") as client:
res = client.system_one(
state="Customer reports a debit card block during an overseas ATM withdrawal.",
questions={
"action": Choice(
instructions="Select the incident resolution playbook",
criteria={
"travel_unblock": "Verify identity and lift the travel restriction",
"dispute_charge": "Open an unauthorized-transaction fraud case",
"branch_visit": "Direct the customer to the nearest branch",
},
),
"is_emergency": Noul(instructions="Is the customer stranded and in urgent need of cash?"),
"urgency_score": Score(instructions="Rate the incident urgency",
criteria=["Routine", "Elevated", "Critical"]),
},
)
print(res.answers["action"].choice, res.answers["action"].confidence)
print(res.answers["is_emergency"].noul, res.answers["urgency_score"].score)# From a TAMEV source checkout
uv pip install -e ".[serve]"Over HTTP
uv run tamev serve --checkpoint Tamkimd/tamev-nano-flash-tinybert --port 8008 --device cpu
curl -X POST http://127.0.0.1:8008/v1/systemone \
-H "Content-Type: application/json" \
-d '{"state": "Payment processor latency spiked to 4s", "questions": {"playbook": {"type": "choice", "instructions": "Pick a remediation", "criteria": {"throttle": "Throttle traffic", "scale": "Scale replicas", "restart": "Restart pods"}}}}'🎯 Use cases
LLM routing · tool-call gating · agent routing · intent routing (77-option catalogues are scored in full) · safety and policy triage · multiple-choice scoring · ordinal scoring. The candidate list is an argument of the call, so the same checkpoint serves a tool menu, a queue list or a rubric without retraining.
🧭 Why not just a classifier — or an LLM router?
- A fixed-label classifier freezes one output per class: add an option, retrain. TAMEV-Nano-Flash-TinyBERT scores whatever list you pass in.
- A generative router writes text your code then parses — prompt tokens, parse failures, and a number that is not a probability.
- TAMEV-Nano-Flash-TinyBERT returns a calibrated distribution over your options in one forward pass.
🔒 Permutation invariance
Guaranteed by construction, measured on these weights. In this checkpoint the state and each option are encoded separately — no option ever attends to another option, and one shared function scores every option slot. Reordering your candidate list therefore reorders the score vector — it cannot change it. That is an architectural property, not a trained behaviour, and it holds on weights nobody has trained at all.
- Measured: 0.0 (bit-exact) drift over 8 random permutations x 19,999 evaluation items = 159,992 reordered pairs, with 0 decision flips, through the documented
predict_decisionhelper. - Bounded, too: each option gets its own 48-token window and options are encoded in chunks, so a 77-option question is scored as 77 options instead of being truncated.
- Not covered: order is free, wording is not — rephrasing an option changes its score, and nothing here makes the model right about your domain.
📊 Evaluation
The set. 19,999 balanced held-out typed-decision items, scored with the weights in this repository through the documented predict_decision helper — one call per item, one CPU thread, safetensors FP32. The split is held out against training, validation and calibration by exact-id, normalised state-text, passage-group and question-text checks, all zero. It shares the same question cells as training by construction, which is why the external rows below exist.
Top-1 / top-3 are exact match; ECE (95% CI 0.0154–0.0243), Brier and NLL are temperature-scaled proper scores (lower is better); permutation drift (159,992 permutation pairs, 0 flips on the shipped weights) is the largest change in the score vector when the candidate list is reordered.
Independently replicated. A second, independent full-set run of the same weights returned identical top-1, top-3, ECE, Brier and NLL rows.
The same rows with content-free option text. The rows above carry each score cell's real option text, which is what a caller sends. On the identical rows with identical labels and only that text replaced by the content-freeRating itemplateDecisionItem.from_dictsynthesises, this checkpoint scores top-1 0.7436 (-3.09pp vs the row above), top-3 0.9476, ECE 0.0381 — read the row above as the deployment number. This row is instrument history and is deliberately kept out of the headline.
Calibration
- ECE 0.0191 (95% CI 0.0154–0.0243) · Brier 0.2959 · NLL 0.5880 over 19,999 items, temperature-scaled with
temperature = 1.32plus the per-K map inconfig.json. - Project criterion is ECE ≤ 0.05; this tier meets it.
- The per-K
calibration_by_kmap served with these weights is inherited unchanged from the previous generation of this tier and was not refit on this head; only the scalartemperatureis this checkpoint's own value.
⚡ Performance
📦 Export formats
Sizes are the files in this repository, in MiB. ONNX FP16 is a cast of the FP32 graph (not a re-export) at half the download, and picks the same answer on 512 sampled gate items. The Hub marks this file Suspicious (Protect AIPAIT-TCHST-301): a TorchScript archive is a ZIP of pickles, so the container alone puts every one of them in scope for a code-execution finding, whatever is inside — the same scan passes this repository'smodel_mps_fp16.pt, which is a real pickle. This one was audited statically rather than loaded: all 101 callables across 3 pickle members aretorchtypes, and the 80 TorchScript source files (77,598 characters) contain no file, process or network primitive and no custom operators. The audit disassembles the pickles rather than loading them, so running it cannot execute what it inspects, and it ships with the source distribution. The file it covers is sha2566d469ef4cb7e.
hf download Tamkimd/tamev-nano-flash-tinybert --local-dir ./tamev-nano-flash-tinybertCompatibility note.TypeSafeDirectClientspeaks this project's own/v1/systemonerequest/response shape. It is not a wrapper around TypeSafe's models and carries no TypeSafe quality guarantee.
🏗 Architecture & training
- Encoder only, no decoder: 312-d backbone (`huawei-noah/TinyBERT_General_4L_312D`),
clspooling, a 64-d projection, and one pointer head shared across option slots (option_fusion: bilinear). 128 tokens of state, 48 per option. - Trained supervised on (state, typed questions, candidate options, label) rows — not on generated text.
choice,noulandscorequestions share the weights. - Teacher / distillation: not re-verified for this release.
- Calibrated after training (temperature scaling, plus the per-K map served in
config.json); the training config and the evaluation harness ship with the TAMEV source distribution.
🎯 Intended use
Built for: the use cases above, through predict_decision / decide or the local HTTP service. It is a decision component inside your system, not an autonomous agent, and it does not know your policy — it scores the options you hand it.
Not for: open-ended generation, factual question answering, or use as an unaudited safety control.
⚠️ Limitations
- Accuracy is in-distribution. The held-out set is drawn from the same source/question cells and templates as training, and the exact-match leakage checks are blind to a paraphrase of a training passage. Pointed at any suite this tier was not trained on — a public QA set, your own ticket queue — top-1 is near chance: fine-tune on your distribution before shipping.
- Calibration. Set-specific. ECE 0.0191 and Brier 0.2959 on the held-out set (n=19,999), 95% CI 0.0154–0.0243 — read the interval, not the point estimate. Target is ≤ 0.05; this tier is inside it.
- Latency is one host, one catalogue shape. 4.00 ms p50 was measured on a single CPU thread on Apple M4, warm option cache; your numbers will differ.
- This tier trades accuracy for latency, on purpose. It shares nano's backbone and parameter count exactly, and differs in one place: the options are encoded independently of each other (
option_fusion: bilinear) instead of jointly with the state (cross). That is what makes the warm option cache above sound, and it is worth several times the wall-clock. The cost is top-1 — 77.45% against nano's 80.93% on the same rows. If top-1 is your binding constraint, use nano; if a per-request latency budget is, use this. - This is a decision model, not a general one. It scores the options you pass it; it does not answer questions or write text.
🥊 TAMEV model zoo
Medium and Large: in progress — larger than the encoder tiers and graded on a different instrument, so no number is published for them yet.
Encoder tiers ship as one fused model. Tiers were evaluated with different option budgets — compare within a column.
🧱 Model details
- Repository: `Tamkimd/tamev-nano-flash-tinybert` · License: Apache-2.0 (LICENSE)
- Call:
model.predict_decision(state, question, options, tokenizer=...)—model.decide(...)is an alias. The localtamevservice exposes the same shape over HTTP.
📄 License & citation
Apache License 2.0 — LICENSE.
@misc{tamev2026,
title={TAMEV: System One Decision Models for Edge AI, LLM Routing and Tool-Call Gating},
author={TAMEV Contributors},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/Tamkimd/tamev-nano-flash-tinybert}
}🔗 References
- Base model: `huawei-noah/TinyBERT_General_4L_312D`
- Training suite: ten public classification sources (listed under
datasets:above), re-formed into typed-decision records with train/validation/calibration/test splits held out by question cell.
