Team Ai
Modelpublic

inferenceprince/laya-onnx-int8

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
1likes265downloads
Model Card

Laya — ONNX (int8 weight-only)

The same Laya model, 236 MB smaller, with identical output. Quantised with MatMulNBits (weight-only, block size 64, symmetric) and exported to ONNX for ONNX Runtime.

Prefer this build when download size or load time matters. Prefer `inferenceprince/laya-onnx` if you need WebGPU without testing — see the note at the bottom.

bash
pip install onnxruntime huggingface_hub tokenizers numpy
hf download inferenceprince/laya-onnx-int8 --local-dir laya-onnx-int8
FileSizeWhat
model.onnx3.1 MBthe computation graph
model.onnx.data606.3 MBint8 weights
tokenizer/3.6 MBunchanged from the base model
rl_agent_config.json—max_len, head_max_len, fitted temperatures

How to use it

Identical to the fp16 build — same five inputs, same output names, same prompt format. Only the directory changes:

python
import json
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer

MODEL_DIR = "laya-onnx-int8"        # the only line that differs from the fp16 build

config = json.load(open(MODEL_DIR + "/rl_agent_config.json"))
tokenizer = Tokenizer.from_file(MODEL_DIR + "/tokenizer/tokenizer.json")
session = ort.InferenceSession(MODEL_DIR + "/model.onnx", providers=["CPUExecutionProvider"])

CLS = tokenizer.token_to_id("[CLS]")
SEP = tokenizer.token_to_id("[SEP]")
MASK = tokenizer.token_to_id("[MASK]")


def tokenize(text):
    """Text to a list of token ids."""
    return tokenizer.encode(text, add_special_tokens=False).ids


def temperature_for(question_type, option_count):
    """The fitted temperature for this kind of question."""
    if option_count <= 2:
        size = "2"
    elif option_count <= 5:
        size = "3-5"
    elif option_count <= 10:
        size = "6-10"
    else:
        size = "11+"

    key = question_type + ":" + size
    if key in config["temperature_by_options"]:
        return config["temperature_by_options"][key]
    return config["temperature"][0]


state = "I was billed twice. Please refund the duplicate."
question = "Which team should handle this?"
options = {
    "billing": "invoices, payments, refunds",
    "technical": "bugs, outages",
    "sales": "pricing",
}

# Build the prompt. The model reads:
#   [CLS] choice question: <question> [SEP]
#   [MASK] billing: ...  [MASK] technical: ...  [MASK] sales: ... [SEP]
#   <the text to analyse> [SEP]
# Every option gets its own [MASK] token, and the model scores that position.
token_ids = [CLS] + tokenize("choice question: " + question) + [SEP]

option_positions = []
for label, description in options.items():
    option_positions.append(len(token_ids))
    token_ids.append(MASK)
    token_ids += tokenize(" " + label + ": " + description)[:48]

token_ids.append(SEP)

room = config["max_len"] - len(token_ids) - 1
token_ids += tokenize(state)[:room]
token_ids.append(SEP)

option_count = len(option_positions)
outputs = session.run(None, {
    "input_ids":      np.array([token_ids], dtype=np.int64),
    "attention_mask": np.ones((1, len(token_ids)), dtype=np.int64),
    "marker_pos":     np.array([option_positions], dtype=np.int64),
    "marker_mask":    np.ones((1, option_count), dtype=bool),
    "qtype":          np.array([0], dtype=np.int64),   # 0=choice, 1=score, 2=noul
})
logits = outputs[0]

temperature = temperature_for("choice", option_count)
scores = logits[0, :option_count] / temperature

scores = scores - scores.max()
probabilities = np.exp(scores)
probabilities = probabilities / probabilities.sum()

labels = list(options.keys())
print("answer:", labels[int(probabilities.argmax())])
for label, probability in zip(labels, probabilities):
    print("  %-12s %.4f" % (label, probability))

Real output from running the above against this build — note it lands a hair under the fp16 build's 0.9696, which is the quantisation showing up at the fourth decimal:

answer: billing
  billing      0.9670
  technical    0.0165
  sales        0.0165

The quantised weights are dequantised inside the MatMulNBits kernel, so this needs an ONNX Runtime recent enough to ship that op — 1.18 or later on CPU. If your runtime is older, use the fp16 build. Everything else in the notes for the fp16 build applies here too: options must fit the 192-token head budget, marker_pos is int64, and the temperature must be applied before softmax.

Why this works when plain int8 does not

Ordinary dynamic int8 destroys this model. It quantises activations as well as weights, and ModernBERT has outlier activation channels that a per-tensor dynamic scale cannot represent. Measured on 8 validation cases, dynamic int8 changed 3 answers outright.

MatMulNBits quantises weights only and dequantises inside the kernel, so activations never leave fp32. The ablation, on the same validation set:

Variantargmax agreement
dynamic int8, per-tensor69.2%
dynamic int8, per-channel76.9%
dynamic int8, MatMuls only65.4%
weight-only int8 (this build)100%

The head and scorer are additionally held in fp32 — about 6% of parameters, but they produce the logits directly, and the runtime divides those by a temperature as low as 0.1, which amplifies any error tenfold.

What it costs you

Nothing measurable. On 116 synthetic items across 15 categories — labels generated by a language model rather than by human annotators, so read the table below as a behavioural comparison between builds rather than as an accuracy benchmark:

BuildSizeOverall`choice``score``noul`Decisions changed
fp16849 MB76.7%77.8%50.0%86.3%—
int8 (this build)613 MB76.7%77.8%50.0%86.3%0 / 116

Calibration is unchanged too: noul ECE 0.054, choice 0.369. Not a single answer moved.

Startup on an Intel i5-14400F, cold process: 1–2 s, against 3–5 s for fp16 and 25–35 s for the PyTorch original. Inference latency is comparable to fp16. This machine is shared and load swung identical work by more than 10x between runs, so take latency as indicative rather than as a specification.

Limits

These are properties of the base model, not of the quantisation:

  • —English only, and it collapses on non-Latin scripts while staying confident.
  • —Not a zero-shot decision engine — the authors say so directly; fine-tune for real accuracy.
  • —Weak on fine procedural distinctions (invoice action 25%) and abstract severity scales (16.7%). score is coarse: never off by more than one level, only half the time exact.
  • —Recalibrate `choice` confidence before automating on it.
  • —Keep choice under ~20 options.

One thing to test before relying on it

ONNX Runtime Web's WebGPU MatMulNBits kernel supports 4- and 8-bit, so this build should run on WebGPU — but issue #25231 reports 8-bit at accuracy_level=4 failing on the dp4 path, and this build uses that setting. It is untested in a browser. If WebGPU rejects it, rebuild with a lower --accuracy-level, or use the fp16 build.

Attribution

Model, training and weights by Nandakishor M and Convai Innovations, Apache-2.0. Quantisation and conversion by this project — an independent work, not an official Convai Innovations release.