inferenceprince/laya-onnx-int8
Laya — ONNX (int8 weight-only)
The same Laya model, 236 MB smaller, with identical output. Quantised with MatMulNBits (weight-only, block size 64, symmetric) and exported to ONNX for ONNX Runtime.
Prefer this build when download size or load time matters. Prefer `inferenceprince/laya-onnx` if you need WebGPU without testing — see the note at the bottom.
pip install onnxruntime huggingface_hub tokenizers numpy
hf download inferenceprince/laya-onnx-int8 --local-dir laya-onnx-int8How to use it
Identical to the fp16 build — same five inputs, same output names, same prompt format. Only the directory changes:
import json
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer
MODEL_DIR = "laya-onnx-int8" # the only line that differs from the fp16 build
config = json.load(open(MODEL_DIR + "/rl_agent_config.json"))
tokenizer = Tokenizer.from_file(MODEL_DIR + "/tokenizer/tokenizer.json")
session = ort.InferenceSession(MODEL_DIR + "/model.onnx", providers=["CPUExecutionProvider"])
CLS = tokenizer.token_to_id("[CLS]")
SEP = tokenizer.token_to_id("[SEP]")
MASK = tokenizer.token_to_id("[MASK]")
def tokenize(text):
"""Text to a list of token ids."""
return tokenizer.encode(text, add_special_tokens=False).ids
def temperature_for(question_type, option_count):
"""The fitted temperature for this kind of question."""
if option_count <= 2:
size = "2"
elif option_count <= 5:
size = "3-5"
elif option_count <= 10:
size = "6-10"
else:
size = "11+"
key = question_type + ":" + size
if key in config["temperature_by_options"]:
return config["temperature_by_options"][key]
return config["temperature"][0]
state = "I was billed twice. Please refund the duplicate."
question = "Which team should handle this?"
options = {
"billing": "invoices, payments, refunds",
"technical": "bugs, outages",
"sales": "pricing",
}
# Build the prompt. The model reads:
# [CLS] choice question: <question> [SEP]
# [MASK] billing: ... [MASK] technical: ... [MASK] sales: ... [SEP]
# <the text to analyse> [SEP]
# Every option gets its own [MASK] token, and the model scores that position.
token_ids = [CLS] + tokenize("choice question: " + question) + [SEP]
option_positions = []
for label, description in options.items():
option_positions.append(len(token_ids))
token_ids.append(MASK)
token_ids += tokenize(" " + label + ": " + description)[:48]
token_ids.append(SEP)
room = config["max_len"] - len(token_ids) - 1
token_ids += tokenize(state)[:room]
token_ids.append(SEP)
option_count = len(option_positions)
outputs = session.run(None, {
"input_ids": np.array([token_ids], dtype=np.int64),
"attention_mask": np.ones((1, len(token_ids)), dtype=np.int64),
"marker_pos": np.array([option_positions], dtype=np.int64),
"marker_mask": np.ones((1, option_count), dtype=bool),
"qtype": np.array([0], dtype=np.int64), # 0=choice, 1=score, 2=noul
})
logits = outputs[0]
temperature = temperature_for("choice", option_count)
scores = logits[0, :option_count] / temperature
scores = scores - scores.max()
probabilities = np.exp(scores)
probabilities = probabilities / probabilities.sum()
labels = list(options.keys())
print("answer:", labels[int(probabilities.argmax())])
for label, probability in zip(labels, probabilities):
print(" %-12s %.4f" % (label, probability))Real output from running the above against this build — note it lands a hair under the fp16 build's 0.9696, which is the quantisation showing up at the fourth decimal:
answer: billing
billing 0.9670
technical 0.0165
sales 0.0165The quantised weights are dequantised inside the MatMulNBits kernel, so this needs an ONNX Runtime recent enough to ship that op — 1.18 or later on CPU. If your runtime is older, use the fp16 build. Everything else in the notes for the fp16 build applies here too: options must fit the 192-token head budget, marker_pos is int64, and the temperature must be applied before softmax.
Why this works when plain int8 does not
Ordinary dynamic int8 destroys this model. It quantises activations as well as weights, and ModernBERT has outlier activation channels that a per-tensor dynamic scale cannot represent. Measured on 8 validation cases, dynamic int8 changed 3 answers outright.
MatMulNBits quantises weights only and dequantises inside the kernel, so activations never leave fp32. The ablation, on the same validation set:
The head and scorer are additionally held in fp32 — about 6% of parameters, but they produce the logits directly, and the runtime divides those by a temperature as low as 0.1, which amplifies any error tenfold.
What it costs you
Nothing measurable. On 116 synthetic items across 15 categories — labels generated by a language model rather than by human annotators, so read the table below as a behavioural comparison between builds rather than as an accuracy benchmark:
Calibration is unchanged too: noul ECE 0.054, choice 0.369. Not a single answer moved.
Startup on an Intel i5-14400F, cold process: 1–2 s, against 3–5 s for fp16 and 25–35 s for the PyTorch original. Inference latency is comparable to fp16. This machine is shared and load swung identical work by more than 10x between runs, so take latency as indicative rather than as a specification.
Limits
These are properties of the base model, not of the quantisation:
- English only, and it collapses on non-Latin scripts while staying confident.
- Not a zero-shot decision engine — the authors say so directly; fine-tune for real accuracy.
- Weak on fine procedural distinctions (invoice action 25%) and abstract severity scales (16.7%).
scoreis coarse: never off by more than one level, only half the time exact. - Recalibrate `choice` confidence before automating on it.
- Keep
choiceunder ~20 options.
One thing to test before relying on it
ONNX Runtime Web's WebGPU MatMulNBits kernel supports 4- and 8-bit, so this build should run on WebGPU — but issue #25231 reports 8-bit at accuracy_level=4 failing on the dp4 path, and this build uses that setting. It is untested in a browser. If WebGPU rejects it, rebuild with a lower --accuracy-level, or use the fp16 build.
Attribution
Model, training and weights by Nandakishor M and Convai Innovations, Apache-2.0. Quantisation and conversion by this project — an independent work, not an official Convai Innovations release.
