Team Ai
Modelpublic

LNTTushar/trynmini-v2-static-7m-v2

sourceHugging Faceapache-2.0updated 6d agoView on Hugging Face
0likes26downloads
Model Card

TrynMini v2 — a tiny static sentence embedder (7.7M params · 7.8 MB · no GPU)

Semantic similarity, search, clustering and RAG retrieval at ~8,000 sentences/sec on a single CPU core — from a 7.8 MB file, with only `numpy` + `tokenizers`. No PyTorch. No GPU. No transformer at inference.

TrynMini v2 is a static embedding model: instead of running a transformer at inference, each token is a row in a learned table. Encoding a sentence is just:

tokenize → look up rows → SIF/Zipf-weighted mean pool → small residual matmul → L2-normalize

That makes it ~10× smaller and dramatically faster than encoder models like all-MiniLM-L6-v2, while keeping most of their quality. It's trained — not hashed — by distilling a strong teacher (BAAI/bge-small-en-v1.5) and fine-tuning with a ranking objective, so the vectors are genuinely useful, not just fast.


Why use this?

**TrynMini v2 (this model)**all-MiniLM-L6-v2potion-base-8M
Size on disk7.8 MB (int8)~90 MB~30 MB
Runtime depsnumpy, tokenizersPyTorch / ONNXnumpy
GPU neededNo (0 VRAM)OptionalNo
Speed (1 CPU core)~8,300 sent/sec~hundreds/secfast
STSB-dev Spearman0.715~0.82~0.75

Use it when you want good-enough semantic vectors that are trivial to ship — serverless functions, edge devices, browsers (via Pyodide), CI, or anywhere a 90 MB PyTorch dependency is too heavy. You get ~87% of all-MiniLM's STS quality at under 1/10th the size and no GPU.

Reach for a full transformer encoder instead when you need the last few points of accuracy and can afford the size/latency.


Performance (measured)

Single CPU core, numpy only, batched encode of short sentences:

MetricValue
Throughput (batched)~8,300 sentences/sec
Latency (one at a time)0.23 ms / sentence
Model load time~0.4 s
RAM resident~75 MB
VRAM0

Benchmarks — STSBenchmark dev (Spearman)

Matryoshka training means you can truncate the vector to trade a little accuracy for a lot of memory/speed:

DimSpearmanNotes
640.700smallest / fastest; already beats v1 at 256
1280.710great default
2560.715full quality

Reference points (reported by their authors, same STSB task): all-MiniLM-L6-v2 ≈ 0.82, Model2Vec potion-base-8M ≈ 0.75.


Install

bash
pip install -U huggingface_hub numpy tokenizers safetensors

Quickstart (5 lines)

python
import sys; from huggingface_hub import snapshot_download
d = snapshot_download("LNTTushar/trynmini-v2-static-7m-v2"); sys.path.insert(0, d)
from modeling_trynmini import TrynMiniV2
m = TrynMiniV2.from_pretrained(d)
emb = m.encode(["a man plays guitar", "someone plays a guitar"], dim=256)   # [2, 256], L2-normalized
print("cosine similarity:", float(emb[0] @ emb[1]))

m.encode(texts, dim=64|128|256) returns L2-normalized vectors, so cosine similarity is just a dot product. The repo also ships example.py (a runnable demo + a small built-in STS sanity check).


How it was trained

  1. 1.Distillation — encode a large sentence corpus with the teacher (BAAI/bge-small-en-v1.5, 384-d), reduce to 256-d, and fit a static token table weighted by Zipf/SIF frequency.
  2. 2.Fine-tuning — a post-pool residual is trained with Multiple-Negatives Ranking Loss (MNRL) + Matryoshka on ~156k NLI / STS / paraphrase pairs (SNLI/MNLI, Quora, STS-B).
  3. 3.Quantization — the table is stored as per-row int8 (≈4× smaller, lossless on STSB here).

Specs

  • —Params: ~7.7M (30,000 × 256 table + 256² residual)
  • —Dims: 256, Matryoshka-truncatable to 128 / 64
  • —Tokenizer: 30k WordPiece (tokenizer.json)
  • —Files: model.safetensors (int8 table), residual.npy, sif.npy, vocab.json, tokenizer.json, modeling_trynmini.py

Limitations

A static model has no attention at inference, so it can't model word order or long-range context the way a transformer can — it trades those last accuracy points for size and speed. It's English, trained for sentence-level similarity (not a drop-in for long-document or multilingual retrieval). Further gains are possible with a stronger teacher (bge-base / e5) and more training.

License

Apache-2.0.