shreyanbr/nev-lite-systemone
nev-lite
nev-lite is the static, no-attention decision head of the nev / System One engine. It answers a typed question over a known, fixed option set with no transformer forward pass at inference: the state is tokenised, its rows are gathered from a learned embedding table and pooled, options are pooled the same way and cached once, and a small interaction MLP scores state against each option. A decision is a table lookup, a pool, and a handful of dot products.
Because options are static and cached, the per-request cost depends only on the state, not on the number of options. This is what nev-lite is built to win on: cost per decision.
- Repository / source code: https://github.com/Shreyan1/nev-system-one
- Backbone (embedding seed only):
sentence-transformers/all-MiniLM-L6-v2 - Model size: ~11.8M parameters, dominated by the 30522×384 embedding table (the MLP and projection are tiny)
- Weights:
nev_lite.pt(PyTorch),option_cache.pt(precomputed option vectors),nev_lite_config.json, tokenizer files
What it does
nev-lite scores three typed primitives, all as a single scalar margin per option:
The margin is not yet a probability; a temperature or Platt scaler is fitted separately (see systemone/calibration.py).
Architecture
state text ──tokenise──► embedding table (static, learned) ──frequency-weighted mean pool──► project(128) ─┐
├─► interaction MLP ─► margin per option
option text ──tokenise──► same pool ──► project(128) ──► CACHED once per question ────────────────────────┘The embedding table is seeded from the backbone's input embeddings (the Model2Vec recipe, minus the offline PCA) and then trained on each benchmark's own gold labels. None of the transformer layers are used at inference — only the embedding matrix comes along, which is what makes nev-lite a lookup rather than a forward pass. Contrast the cross-encoder tier (nev-high), which re-encodes the state once per option, so a K-option question costs K encoder passes.
How to use
nev-lite is a custom head, so load it through the repository code (there is no AutoModel entry point).
import json
import torch
from huggingface_hub import snapshot_download
from transformers import AutoTokenizer
# the StaticModel class and helpers live in the repo:
# pip install git+https://github.com/Shreyan1/nev-system-one.git
from systemone.lite import MAX_STATE_TOKENS, StaticModel, render_state
path = snapshot_download('shreyanbr/nev-lite-systemone')
cfg = json.load(open(f'{path}/nev_lite_config.json'))
tokenizer = AutoTokenizer.from_pretrained(path)
model = StaticModel(cfg['vocab_size'], cfg['embedding_dim'], cfg['projection_dim']).eval()
model.load_state_dict(torch.load(f'{path}/nev_lite.pt', map_location='cpu'))
# OptionCache is a dataclass, so the pickle is not weights-only
option_cache = torch.load(f'{path}/option_cache.pt', map_location='cpu', weights_only=False)
def decide(task_qid: str, instructions: str, state_text: str) -> str:
"""Return the winning option key for one cached question."""
bank = option_cache[task_qid]
encoded = tokenizer(
[render_state(instructions, {'text': state_text})],
truncation=True, max_length=MAX_STATE_TOKENS, return_tensors='pt',
)
with torch.no_grad():
state_vector = model.pool(encoded['input_ids'], encoded['attention_mask'])[0]
logits = model.score(state_vector, bank.vectors)
return bank.options[int(logits.argmax())]
print(decide('ag_news:ag_news', 'Which topic is this news article about?',
'The central bank raised interest rates by half a point.'))option_cache keys are f'{task}:{qid}' (for example ag_news:ag_news, clinc150:clinc150). To score your own option set instead of a cached benchmark, pool your option texts with model.pool(...) once and reuse the vectors.
Calibrated probabilities
Raw margins are not probabilities. LiteDecider applies the fitted calibration (a convex temperature law for choice/score, a Platt scaler for noul) and returns a probability distribution plus an entropy confidence:
from systemone.lite import LiteDecider
decider = LiteDecider('shreyanbr/nev-lite-systemone') # or a local path
decision = decider.decide('ag_news:ag_news',
{'type': 'choice', 'instructions': 'Which topic is this news article about?',
'criteria': {'World': None, 'Sports': None, 'Business': None, 'Sci/Tech': None}},
{'text': 'The central bank raised interest rates by half a point.'})
print(decision.key, decision.confidence, decision.probabilities)The calibration is fitted on the held-out calib split (train.calibrate_lite), cross-fitted so its quality is measured on items the fit never saw, and stored as calibration.json. A well-calibrated model is confident when it is right and unsure when it is not, so confidence is a usable signal for routing or deferral.
Backbone
The embedding table is seeded from a configurable backbone (only the token embeddings are used, never attention layers). Supported seeds: any transformer encoder (sentence-transformers/all-MiniLM-L6-v2, all-mpnet-base-v2, BGE, etc.) and model2vec static-embedding models such as minishlab/potion-base-32M, which are purpose-built for pooling. The checkpoint's actual backbone is recorded in nev_lite_config.json.
Training
- Backbone:
sentence-transformers/all-MiniLM-L6-v2(embedding table only, seeded then trained) - Objective: cross-entropy over the option set (
choice), cross-entropy over ordered levels (score), binary cross-entropy on the truth logit (noul) - Hyperparameters: 4 epochs, batch 128, LR 1e-3, AdamW, weight decay 0.01, seed 42
- Hardware: single NVIDIA Tesla T4 (15.6 GB), a few minutes end to end
- Data: each dataset's official splits; supervision is each dataset's own gold labels
Reproduce with `notebooks/nev_lite_colab.ipynb`.
Evaluation
Measured on each benchmark's official test split (validation where the test labels are withheld), on a Tesla T4. floor is the majority-class baseline (always answer the most frequent train label). tok/dec is the input tokens the model actually processes per decision — state only, because options are cached. dec/sec is single-item (batch 1) throughput, a floor that batching raises substantially.
¹ The GLUE SST-2 test split ships unlabelled (all labels −1). The evaluation now scores SST-2 on its labelled validation split; the previous run's SST-2 number was measured against the unlabelled split and is therefore omitted. Re-run the training/eval notebook to populate it.
² This run predated a fix to the calibration hold-out: a flat 500-row calib split left only 46 of deepset/prompt-injections' 546 train rows for training. With the fraction-capped hold-out the task now trains on ~437 rows and reaches ~0.90 in local runs; re-run the notebook to populate the joint number.
The choice tasks are strong and clear the majority floor by a wide margin. Among the noul (single-logit) tasks, sentiment and prompt-injection are learnable once given enough data, but BoolQ is an architectural limit: a no-attention bag-of-embeddings cannot do reading comprehension over a question and passage, so it sits near its floor no matter how it is trained. Route entailment/reading-comprehension questions to the attention tiers (nev-med/high) rather than nev-lite.
Cost model
nev-lite's value proposition is cost per decision, so here is how to compute it rather than a single headline number (the number depends on your hardware, your batch size, and what you compare against).
Per-decision compute cost on a given machine:
cost_per_decision = hourly_rate / (3600 * decisions_per_sec)Worked example, T4 at an assumed $0.40/GPU-hour, batch 1 (substitute your own rate and your batched throughput):
- at 1000 dec/sec → ~$0.11 per million decisions
- at 640 dec/sec (boolq, long states) → ~$0.17 per million decisions
So order-of-magnitude ~$0.10–0.20 per million decisions at batch 1 on a T4, and lower with batching.
Why it is cheap, structurally:
- Options are encoded once and cached, so a K-option question costs the state only. The cross-encoder tier (nev-high) re-encodes the state per option (measured ~2438 tokens/decision on Banking77 with K=77), and an LLM classifier must send the state, the instructions, and every option description in-context on every call.
- At, say, $0.20 per 1M input tokens, an in-context classifier sending ~2400 tokens/decision costs ~$480 per million decisions in input alone — before output tokens — versus nev-lite's ~$0.10–0.20. That is a rough illustration with an assumed rate, not a measured head-to-head.
To get a real per-use-case cost, you need (1) batched throughput on your target hardware, and (2) a concrete baseline to compare against. See "Next steps" in the repository.
Limitations and biases
- Fixed option sets. nev-lite answers questions over a known option set. It does not generate text and cannot answer open-ended questions.
- Weak on binary/entailment (`noul`) tasks in this release, as the evaluation shows.
- English only, and only as good as the eight public benchmarks it was trained on; expect domain shift on data unlike those datasets.
- Prompt-injection detection is a soft classifier, not a security control. ~0.61 accuracy is far from reliable; do not use it as a standalone guardrail.
- Static embeddings carry the biases of the backbone's token embeddings and the training corpora. No debiasing was applied.
- Margins are not calibrated probabilities out of the box; fit a scaler (
systemone/calibration.py) before thresholding.
License
Apache-2.0. See the repository LICENSE.
Citation
@software{nev_system_one,
title = {nev / System One: typed decisions over a known option set, one forward pass, zero generated tokens},
author = {Basu Ray, Shreyan},
year = {2026},
url = {https://github.com/Shreyan1/nev-system-one}
}