Team Ai
Modelpublic

leli14/laya_cpu_onnx

sourceHugging Faceapache-2.0updated 18d agoView on Hugging Face
0likes
Model Card

Laya CPU ONNX (Multilingual Intent Classifier)

Laya CPU ONNX is a high-performance, lightweight in-process intent classification engine optimized for Real-Time Conversational AI & Digital Human Agents (e.g., AI Interviewer, Virtual Agent).

Based on the ModernBERT architecture with multi-choice intent heads, this model provides millisecond-level decision making on standard x86_64 CPU environments with zero external network RTT.


Model Highlights

  • —Pure CPU Optimized: Quantized to Dynamic INT8; no GPU / CUDA dependencies required.
  • —Ultra-low Latency: ~150ms – 180ms per decision on 4 standard CPU cores.
  • —In-Process Inference: Fully compatible with Java (via onnxruntime-java & DJL tokenizers) and Python (onnxruntime).
  • —Multilingual Tokenizer: Pre-trained on 166,000+ vocabulary tokens covering Chinese, English, and multi-turn conversational semantics.

Model Architecture & Artifacts

FileSizeDescription
model_quantized.onnx1.23 GBDynamic INT8 quantized ONNX graph (~322M parameters)
tokenizer.json32.8 MBHuggingFace FastTokenizer vocabulary (166k tokens)
tokenizer_config.json524 BSpecial token mappings (<bos>, <eos>, <mask=4>, <pad=0>)
special_tokens_map.json731 BToken definitions

Graph I/O Specification

1. Inputs

NameTypeShapeDynamic AxesDescription
input_idsint64[batch, seq]batch, seqToken IDs (Max sequence: 512)
attention_maskint64[batch, seq]batch, seqAttention mask (1 for valid, 0 for pad)
marker_posint64[batch, options]batch, optionsPosition index for <mask=4> markers
marker_maskbool[batch, options]batch, optionsBoolean mask for candidates
qtypeint64[batch]batchTask type: 0 for Choice Question

2. Outputs

NameTypeShapeDescription
logitsfloat32[batch, options]Unnormalized logits across the candidate intents
act_probsfloat32[batch, 2]Binary action probabilities

Decision States & Business Scenarios

The model is designed for turn-level dialogue arbitration, supporting two primary stages:

1. Initial Answer Stage (INITIAL_ANSWER)

  • —FINISHED: The user has completed answering the question.
  • —CONTINUE: The user is still elaborating their thoughts.
  • —THINKING: The user explicitly requests time to organize their thoughts.
  • —CLARIFICATION: The user is asking for question clarification or scope definition.
  • —UNCERTAIN: Ambiguous intent, recommended for fallback to an LLM.

2. Supplement Stage (SUPPLEMENT)

  • —NO_SUPPLEMENT: The user has no further comments.
  • —HAS_SUPPLEMENT: The user provides new facts or follow-up details.
  • —THINKING: The user needs time to ponder.
  • —UNCERTAIN: Ambiguous state.

Quickstart (Python)

python
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer

# 1. Load Tokenizer & Session
tokenizer = Tokenizer.from_file("tokenizer.json")
sess_opts = ort.SessionOptions()
sess_opts.intra_op_num_threads = 4
session = ort.InferenceSession("model_quantized.onnx", sess_options=sess_opts, providers=['CPUExecutionProvider'])

# 2. Prepare Candidates & Inputs
CRITERIA = [
    "FINISHED: 候选人已回答完毕,核心内容表达完整或明确表示回答完毕",
    "CONTINUE: 候选人正在回答中尚未结束,还在继续说明",
    "THINKING: 候选人需要思考或稍等",
    "CLARIFICATION: 候选人正在向面试官确认题意或概念范围",
    "UNCERTAIN: 无法可靠判断是否结束"
]

prompt = "choice question: 针对面试问题【请介绍你的项目经历】,判断候选人的回答状态"
state = "我主要负责微服务架构和核心接口开发,我说完了。"

# 3. Build sequence with <bos>=2, <eos>=1, <mask=4>
head_ids = tokenizer.encode(prompt).ids
seq = [2] + head_ids + [1]
markers = []
for opt in CRITERIA:
    markers.append(len(seq))
    seq.append(4)  # <mask=4>
    seq.extend(tokenizer.encode(" " + opt).ids[:48])
seq.append(1)
seq.extend(tokenizer.encode(state).ids[: 511 - len(seq)])
seq.append(1)

# 4. Run Inference
inputs = {
    "input_ids": np.array([seq], dtype=np.int64),
    "attention_mask": np.ones((1, len(seq)), dtype=np.int64),
    "marker_pos": np.array([markers], dtype=np.int64),
    "marker_mask": np.ones((1, len(markers)), dtype=bool),
    "qtype": np.array([0], dtype=np.int64)
}

outputs = session.run(None, inputs)
logits = outputs[0][0]
probs = np.exp(logits) / np.sum(np.exp(logits))
best_idx = np.argmax(probs)
print(f"Predicted Decision: {CRITERIA[best_idx].split(':')[0]} (Confidence: {probs[best_idx]:.3f})")

Quickstart (Java)

Add the following dependencies in your pom.xml:

xml
<dependency>
    <groupId>com.microsoft.onnxruntime</groupId>
    <artifactId>onnxruntime</artifactId>
    <version>1.18.0</version>
</dependency>
<dependency>
    <groupId>ai.djl.huggingface</groupId>
    <artifactId>tokenizers</artifactId>
    <version>0.31.1</version>
</dependency>

In your Java service:

java
OrtEnvironment env = OrtEnvironment.getEnvironment();
OrtSession.SessionOptions opts = new OrtSession.SessionOptions();
opts.setIntraOpNumThreads(4);
OrtSession session = env.createSession("model_quantized.onnx", opts);
HuggingFaceTokenizer tokenizer = HuggingFaceTokenizer.newInstance(Paths.get("tokenizer.json"));