Team Ai
Modelpublic

iapp/OpenThai-SystemOne-MLX-4bit

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
0likes23downloads
Model Card

OpenThai-SystemOne — mlx-4bit

OpenThai-SystemOne is an open Thai + English System One decision model: one forward pass answers typed questions (choice over up to 255 options, ordinal score, yes/no noul) about a text / JSON state with calibrated probabilities, no text generation. It is a Qwen3.5-0.8B text tower (Thai continued pre-training) plus a 256-slot decision head. This repo is a quantization of v0.3 (commit f3709948).

What is quantized: the tower including the token embeddings (MLX quantizes the embedding table too). The 256-slot decision head and the per-type temperatures stay in fp32 (head.safetensors). Quantization therefore only perturbs the hidden state the head reads.

Format: MLX 4-bit affine (group 64) for Apple Silicon (mlx-lm). mlx-lm runs the tower; the included client applies the decision head on the final hidden states. Size: 424 MB.

Measured on a MacBook Pro M3 Max: 4-bit ≈ 19 ms per 3-question Thai decision (the PyTorch model on MPS: ~150 ms).

Usage

bash
pip install mlx-lm torch transformers safetensors pydantic
huggingface-cli download iapp/OpenThai-SystemOne-MLX-4bit --local-dir openthai-mlx
python
import sys; sys.path.insert(0, "openthai-mlx")
from openthai_systemone.mlx_client import MLXSystemOneClient
c = MLXSystemOneClient("openthai-mlx")
r = c.system_one("ร้านนี้อาหารอร่อยมาก แต่รอนานเกือบชั่วโมง", {"sentiment": {"type": "choice", "instructions": "ความรู้สึก",
                 "criteria": {"บวก": None, "ลบ": None, "กลาง": None}}})
print(r.answers["sentiment"].choice, r.answers["sentiment"].probabilities)

Accuracy vs the bf16 original (same records, single option order, first 800 per set)

Macro: public 72.9 (original 74.3), Thai 79.0 (original 80.1).

subsetbf16 originalthisΔ
public 13-subset bench
aegis2 (noul)83.282.0-1.2
boolq (noul)79.777.7-2.0
civil_comments (noul)79.080.3+1.3
helpsteer2 (score)41.641.6+0.0
massive-de-DE (choice)88.379.4-8.9
massive-en-US (choice)88.379.7-8.6
multinli (choice)89.088.0-1.0
paws (noul)94.092.4-1.6
pubmedqa (choice)64.064.0+0.0
squad2 (noul)89.387.0-2.3
summeval-consistency (score)75.080.6+5.6
summeval-relevance (score)21.725.0+3.3
vitaminc-dev (choice)72.570.6-1.8
macro, public 13-subset bench74.372.9-1.3
Thai held-out / eval sets
banking77 (choice)59.152.1-7.0
contrastive_th (choice)80.780.4-0.3
contrastive_th (noul)83.583.9+0.4
contrastive_th (score)78.678.6+0.0
massive_th (choice)90.687.5-3.1
prachathai (choice)98.398.8+0.5
prachathai (noul)93.492.7-0.8
sib200_th (choice)77.977.0-1.0
wisesight (choice)48.946.5-2.4
wongnai (score)64.565.2+0.8
xlam_tools (choice)99.499.4+0.0
xnli_th (choice)79.879.0-0.8
xnli_th (noul)86.886.0-0.8
macro, Thai held-out / eval sets80.179.0-1.1

Notes

  • —Scores are single-option-order accuracy on the first 800 records of each set (scripts/06_eval.py --limit 800), the same records for the original and the quantization. score subsets report exact level accuracy.
  • —Base model, data, training and the full benchmark tables: iapp/OpenThai-SystemOne.
  • —License Apache-2.0 (same as the base). Built by iApp Technology / OpenThaiGPT.