Team Ai
Modelpublic

iapp/OpenThai-SystemOne-NVFP4

sourceHugging Faceapache-2.0updated 13d agoView on Hugging Face
0likes15downloads
Model Card

OpenThai-SystemOne — NVFP4

OpenThai-SystemOne is an open Thai + English System One decision model: one forward pass answers typed questions (choice over up to 255 options, ordinal score, yes/no noul) about a text / JSON state with calibrated probabilities, no text generation. It is a Qwen3.5-0.8B text tower (Thai continued pre-training) plus a 256-slot decision head. This repo is a quantization of v0.3 (commit f3709948).

What is quantized: the Linear layers of the tower. The token embeddings, the 256-slot decision head and the per-type temperatures stay in bf16. Quantization therefore only perturbs the hidden state the head reads.

Format: NVFP4 W4A4 (FP4 weights and activations, FP8 scales). compressed-tensors (llm-compressor) checkpoint. Loaded through transformers the weights are decompressed to bf16 at load time (same speed as bf16, smaller download); native FP8 / INT8 / FP4 kernels need a runtime with this architecture (the decision head is custom, so vLLM does not serve it out of the box). NVFP4 kernels need NVIDIA Blackwell.

Size: 790 MB (bf16 original: 1,509 MB).

Usage

bash
pip install torch transformers safetensors pydantic && pip install compressed-tensors
python
from transformers import AutoModel, AutoTokenizer        # trust_remote_code files are in this repo
model = AutoModel.from_pretrained("iapp/OpenThai-SystemOne-NVFP4", trust_remote_code=True)
# or with the pip package (git+https://github.com/iapp-technology/openthai-systemone):
from openthai_systemone import SystemOneClient
c = SystemOneClient("iapp/OpenThai-SystemOne-NVFP4")
r = c.system_one("ร้านนี้อาหารอร่อยมาก แต่รอนานเกือบชั่วโมง", {"sentiment": {"type": "choice", "instructions": "ความรู้สึก",
                 "criteria": {"บวก": None, "ลบ": None, "กลาง": None}}})

Accuracy vs the bf16 original (same records, single option order, first 800 per set)

Macro: public 73.7 (original 74.3), Thai 78.7 (original 80.1).

subsetbf16 originalthisΔ
public 13-subset bench
aegis2 (noul)83.284.0+0.8
boolq (noul)79.779.0-0.7
civil_comments (noul)79.075.7-3.3
helpsteer2 (score)41.641.6+0.0
massive-de-DE (choice)88.385.1-3.1
massive-en-US (choice)88.388.6+0.3
multinli (choice)89.084.9-4.0
paws (noul)94.092.8-1.2
pubmedqa (choice)64.062.0-2.0
squad2 (noul)89.386.6-2.7
summeval-consistency (score)75.079.9+4.9
summeval-relevance (score)21.726.7+5.0
vitaminc-dev (choice)72.571.3-1.2
macro, public 13-subset bench74.373.7-0.6
Thai held-out / eval sets
banking77 (choice)59.157.0-2.1
contrastive_th (choice)80.779.4-1.4
contrastive_th (noul)83.581.9-1.6
contrastive_th (score)78.675.0-3.6
massive_th (choice)90.688.0-2.6
prachathai (choice)98.398.5+0.2
prachathai (noul)93.493.6+0.2
sib200_th (choice)77.975.0-2.9
wisesight (choice)48.948.4-0.5
wongnai (score)64.563.5-1.0
xlam_tools (choice)99.499.5+0.1
xnli_th (choice)79.877.2-2.5
xnli_th (noul)86.886.0-0.8
macro, Thai held-out / eval sets80.178.7-1.4

Notes

  • —Scores are single-option-order accuracy on the first 800 records of each set (scripts/06_eval.py --limit 800), the same records for the original and the quantization. score subsets report exact level accuracy.
  • —Base model, data, training and the full benchmark tables: iapp/OpenThai-SystemOne.
  • —License Apache-2.0 (same as the base). Built by iApp Technology / OpenThaiGPT.