iapp/OpenThai-SystemOne-GGUF
OpenThai-SystemOne — GGUF
OpenThai-SystemOne is an open Thai + English System One decision model: one forward pass answers typed questions (choice over up to 255 options, ordinal score, yes/no noul) about a text / JSON state with calibrated probabilities, no text generation. It is a Qwen3.5-0.8B text tower (Thai continued pre-training) plus a 256-slot decision head. This repo is a quantization of v0.3 (commit f3709948).
What is quantized: the tower (all Linear layers and the token embeddings, at the GGUF file's level). The 256-slot decision head and the per-type temperatures stay in fp32 (head.safetensors). Quantization therefore only perturbs the hidden state the head reads.
llama.cpp runs the tower and returns its final hidden states (embedding=True, pooling_type=NONE); the included Python client applies the decision head on top, so answers are identical in shape to the PyTorch model.
Files
Usage
pip install llama-cpp-python torch transformers safetensors pydantic # CMAKE_ARGS="-DGGML_CUDA=on" or "-DGGML_METAL=on" for GPU
huggingface-cli download iapp/OpenThai-SystemOne-GGUF --local-dir openthai-gguf \
--include "*Q4_K_M.gguf" "head*" "tokenizer*" "openthai_systemone/*"import sys; sys.path.insert(0, "openthai-gguf")
from openthai_systemone.gguf import GGUFSystemOneClient
c = GGUFSystemOneClient("openthai-gguf/OpenThai-SystemOne-v0.3-Q4_K_M.gguf") # n_gpu_layers=-1 by default
r = c.system_one("ร้านนี้อาหารอร่อยมาก แต่รอนานเกือบชั่วโมง พนักงานไม่สนใจลูกค้าเลย", {
"sentiment": {"type": "choice", "instructions": "ความรู้สึกของข้อความ", "criteria": {"บวก": None, "ลบ": None, "กลาง": None}},
"urgent": {"type": "noul", "instructions": "ต้องรีบแก้ไขหรือไม่"},
"stars": {"type": "score", "instructions": "ให้ดาว", "criteria": ["1", "2", "3", "4", "5"]}})
print(r.answers["sentiment"].choice, r.answers["sentiment"].probabilities)The GGUF alone in llama-cli / llama-server is only the tower: its LM head is the tied input embedding, not the decision head, so generated text is meaningless. Use the client (or read hidden states with --embeddings --pooling none and apply head.safetensors yourself: softmax((h @ W.T + b) / exp(log_temperature[qtype])) over the first k slots + slot 255).
Measured on an H100 (llama-cpp-python 0.3.35, CUDA): ~35 ms per 3-question Thai decision for every level.
Accuracy of Q4KM vs the bf16 original (same records, single option order, first 800 per set)
Notes
- Scores are single-option-order accuracy on the first 800 records of each set (
scripts/06_eval.py --limit 800), the same records for the original and the quantization.scoresubsets report exact level accuracy. - Base model, data, training and the full benchmark tables: iapp/OpenThai-SystemOne.
- License Apache-2.0 (same as the base). Built by iApp Technology / OpenThaiGPT.
