mlboydaisuke/OpenThai-SystemOne-CoreAI
041
1---2license: apache-2.03base_model: iapp/OpenThai-SystemOne4language:5- th6- en7tags:8- coreai9- decision-model10- system-one11- zero-shot-classification12- thai13---14 15Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's `coreai-torch` (LLMs: `coreai.llm.export`) into `.aimodel` bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol ([apple-silicon-llm-bench](https://github.com/john-rocky/apple-silicon-llm-bench), macOS 27 beta 26A5353q, 2026-06-11).16 17# OpenThai-SystemOne — Core AI18 19[🤗 mlboydaisuke/OpenThai-SystemOne-CoreAI](https://huggingface.co/mlboydaisuke/OpenThai-SystemOne-CoreAI) · Apache-2.0 · source [iapp/OpenThai-SystemOne](https://huggingface.co/iapp/OpenThai-SystemOne/tree/f3709948b5e3cc9606a57e74ba62b7a639d17dd3) (revision `f3709948`) · base Qwen/Qwen3.5-0.8B-Base20 21A **Thai + English System One decision model**: it reads a state and typed questions — Choice,22Score or Noul — and returns option probabilities. The Qwen3.5 text tower was continued-pretrained23on Thai; its language-model head was replaced by a **256-way biased slot head**. The readout is24at `<|ts_answer|>`. It never generates text.25 26This is the zoo's second decision model. The graph uses the Qwen3.5 decode-only, loop-free S=127recipe with three changes: the `model.*` weight prefix, a 248,339-row embedding, and the slot28head in the `lm_head` position. **int8lin is the ship bundle; fp16 is published beside it as29the reference.** Both have a 4,096-token context and emit `logits` with shape `[1, 1, 256]`.30 31## Readout contract32 33Each question is one independent row in the author's layout, with no chat template or BOS:34 35```text36<|ts_state|> <state>37<|ts_q|><|ts_choice|> <instructions>38<|ts_opt_0|> <name>: <description>39<|ts_opt_1|> <name>40<|ts_answer|>41```42 43The author's encoded sequence has a newline after `<|ts_answer|>`; the answer slot is44`len(ids_full) - 2`. The bundle consumes the prefix through the answer token, id **248082**,45and reads the last call's 256 logits. Removing the trailing newline changes the fp32 oracle46logits by at most **0.000018597** across the fixtures (tolerance 0.0001).47 48- Choice options retain request order. Score uses `<|ts_score|>` and options `i: <level>`,49 with 2–10 levels. Noul uses `<|ts_noul|>` and slots `0 = no`, `1 = yes`, with the supplied50 false/true descriptions when present.51- For `k` options, divide all slot logits by the question type's temperature, mask slots52 `k..254` to negative infinity, and softmax over all 256 slots. Return53 `p_options = p_full[:k] / sum(p_full[:k])`; retain `p_full[255]` as abstain. Slot indices54 are **not vocabulary token ids**. Choice supports up to 255 options.55- Temperatures are read from `exp(log_temperature)` in the checkpoint: **choice 1.058534,56 score 1.043141, noul 1.006767**. They differ from the author's v0.3 card, which quotes57 **choice 1.055, score 1.008, noul 1.047**. Metadata carries the tensor-derived values at58 full precision.59- The author's API exposes abstain for **Choice only**. Score exposes probabilities and60 confidence; Noul exposes the probability of yes. Confidence is one minus normalized61 entropy. The pinned client does not round; its Choice/Score assembly performs a second62 fp32 renormalization. All fixtures use `permutations=1`.63 64The [fixture](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/fixtures-openthai-systemone.json)65contains 18 requests: 48 kit rows (24 Choice, 14 Noul, 10 Score) and two zoo rows with 40 and66255 options. It covers Thai, English, mixed text, dict states and list states. The bundle and67the kit answer one question per row. The independent-row fp32 oracle assembly equals the68author's single-question API exactly for **18/18 requests** (50/50 question calls). The author's69one-pass API places several questions in one causal sequence; its answers to later questions70can differ from independent rows: max **|Δp| 0.375453** on these requests. That comparison is71recorded as API behavior, not a conversion gate.72 73## Measured (Apple M4 Max GPU, macOS 27.0 26A428, 2026-09-23)74 75| | fp16 (reference) | int8lin (ship) |76|---|---:|---:|77| option argmax = author's fp32 oracle | 50/50 | 50/50 |78| argmax on oracle margin ≥ 0.02 | 49/49 | 49/49 |79| max \|Δp\| over option probabilities | 0.005059 | 0.020813 |80| mean of per-row mean \|Δp\| | 0.000225 | 0.000659 |81| max \|Δabstain\| | 0.016488 | 0.018837 |82| Swift pipelined first token = decoded raw-slot argmax | 50/50 | 50/50 |83| Swift sequential first token = decoded raw-slot argmax | 50/50 | 50/50 |84| state reset, row 1 logits bit-identical | yes | yes |85 86The gate requires option-argmax agreement on every row with oracle margin ≥ 0.02, finite87logits and the state-reset proof. Probability and abstain errors are recorded. The only row88below that margin is `r18-slot` (0.009739); it agrees on both bundles. The largest int8lin89probability difference is `r05-dry`, a two-option Noul row with oracle margin 0.061681.90 91Probabilities come from an AOT h16c GPU asset loaded through the Core AI Python runtime with92`SpecializationOptions.default()`, fresh zero states per row and full `position_ids` at each93S=1 step. The engine check uses Release `llm-runner`, raw ids, one greedy token, and94`COREAI_CHUNK_THRESHOLD=1`. It compares `tokenizer.decode([raw256_argmax])` against the same95bundle's unmasked Python readout; that diagnostic string is not a decision answer. Transcripts:96[fp16 readout](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/gate-openthai-systemone-readout-fp16.json),97[int8lin readout](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/gate-openthai-systemone-readout-int8lin.json),98[fp16 engines](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/gate-openthai-systemone-engine-fp16.json),99[int8lin engines](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/gate-openthai-systemone-engine-int8lin.json).100 101**Throughput**, int8lin, Release `llm-benchmark`, p=128 / g=256, two launches × three trials102per engine, `COREAI_CHUNK_THRESHOLD=1`; median (range):103 104| Engine | prefill proxy, tok/s | decode, tok/s | load per launch, s |105|---|---:|---:|---:|106| coreai-pipelined | 252.7 (243.8–258.6) | 250.8 (244.5–253.8) | 1.415 / 0.167 |107| coreai-sequential | 197.1 (196.1–201.3) | 194.8 (192.6–197.7) | 0.172 / 0.166 |108 109This is a **prefill-rate proxy**: the graph is S=1, and synthetic generation throughput is110not decision latency. The benchmark samples ids from the metadata's 256-wide output range.111No other Core AI, Python or Swift engine job appeared in the before/after process snapshots112(`contended: false`). Load is measured per launch, excluding warmup. The frozen Swift tag's113benchmark needed a local CLI option to select `EngineOptions.variant`; the trial loop was114unchanged. [Trials, load times and environment](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/llm-benchmark.json).115 116### JevBench public 231 (Mac, 2026-09-24)117 118| easy 48 | standard 72 | hard 111 | ECE hard | p50 | p95 | hard max |119|---:|---:|---:|---:|---:|---:|---:|120| 1.000 | 0.819 | 0.324 | 0.449 | 0.62 s | 25.93 s | 39.0 s |121 122The benchmark's own harness123([fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench) `2fa63fa`, v1.4.0,124`typesafe` adapter) ran the 231 public items against coreai-kit `adbc755` `decide-cli serve`125with the ship bundle, one question per request. Accuracy per tier and the hard tier's ECE are126JevBench's own scoring (argmax of the returned probabilities); p50 and p95 are per-request127latency over all 231 requests, hard max the maximum over the hard tier. Latency was measured128without an exclusive GPU window (contended), with this model's server running alone; the129graph's prefill is S=1, and at that kit commit the state was prefilled again for every130question. JevBench's published scores (Intelligence and the rest) are chance-corrected over 534131items, sealed ones included, and are not comparable to these accuracies.132 133**iPhone 17 Pro (2026-09-24)**134 135| | easy 48 | standard 72 | hard 20 |136|---|---:|---:|---:|137| accuracy | 1.000 | 0.819 | 0.200 |138| p50 | 0.88 s | 0.99 s | 5.77 s |139| p95 | 1.29 s | 1.28 s | 7.28 s |140| p50 / p95 over | 48 rows, hot | 24 of 72 rows, nominal | 20 rows, nominal |141 142The phone (iOS 27.0 24A437) received the same request bodies as the Mac run, one question143per request. A headless harness app answered each with coreai-kit 0.7.1, through the call the144kit's System One server makes. Every bundle file on the phone matched the Hub revision by145hash. Every answer's argmax equals the Mac run's. The hard column is the middle 20 of the 111146hard items by state length. The Mac run scored 0.200 on the same 20. p50 and p95 are the kit's147time per request: the state's prefill plus the decision. A nominal row started and ended with148the phone on its battery at thermal state nominal. A hot row started or ended at fair or worse.149The easy tier and the first 48 standard rows ran on the charger.150 151## Through the kit152 153**Measured through coreai-kit**, using its sequential engine and tokenizer: `decide-cli parity`154matched tokens 50/50, answer slots 50/50 and option argmax 50/50 on both bundles, including the15540- and 255-option rows. int8lin max |Δp| was **0.0226** (`r05-dry`), mean **0.0009**, and max156|Δabstain| **0.0222**; fp16 was **0.0051** (`r05-task`), **0.0003**, and **0.0165** respectively.157These are kit measurements supplied by the supervisor, separate from the Python-runtime158table above. Median int8lin wall time per fixture question was **354 ms** over the 50 rows, two to159three questions per state (a question on a new state pays for the whole state); the 255-option row (1,449 tokens, S=1 prefill) took **7.2 s**. A160three-question Thai ticket took **351 / 316 / 429 ms** for its 57-, 63- and 83-token rows; the161recurrent hybrid cannot rewind mid-sequence, so every row is prefilled from its first token (0162tokens reused).163 164On SemIf's authored144 — 144 English rows with three options, SemIf's gold labels and unchanged165`benchmarks/evaluate.py` — int8lin on the Mac GPU, measured through coreai-kit, scored166**109/144** raw and **0.7249 mean family balanced accuracy**. The kit README reports **0.681**167for MiniCPM5-2B int8 and **0.821** for Qwen3.5-4B int8 on the same rows and evaluator.168[Kit measurement record](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/measurements-coreai-kit.json).169**iPhone 17 Pro** (iOS 27.0 24A437, the same int8lin bundle sideloaded into the kit's ModelStore, sha256 equal to170the Hub revision, 2026-09-23, a headless harness that runs the kit's own `decide-cli parity` / `oracle` inside an171app; thermal state "serious" throughout): `parity` matched tokens **50/50**, answer slots **50/50** and option172argmax **50/50**, the 40- and 255-option rows included (the 1,449-token row fits this bundle's context); max |Δp|173**0.0210**, mean **0.0009**, max |Δabstain| **0.0213**. Median wall time per fixture question **1,893 ms** (Mac174354 ms); the 255-option row **40.8 s** (Mac 7.2 s). On SemIf's authored144 through `oracle` on the phone: **109/144**175raw and **0.7249** mean family balanced accuracy — the Mac's figures exactly — at **2,074 ms** median per decision.176Cooled to thermal state "nominal" (2026-09-24, `decide-cli bench --repeat 3`, a 111-token state and eight questions):177**1,527 ms per decision** with the state shared, 1,713 from scratch, 13.5 s for the state and its eight (0 tokens178reused). Load 6.6 s. Records: the SemIf rows in the standup record, and179[`gate-openthai-systemone-iphone-parity.json`](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/gate-openthai-systemone-iphone-parity.json).180 181## Bundle182 183[mlboydaisuke/OpenThai-SystemOne-CoreAI](https://huggingface.co/mlboydaisuke/OpenThai-SystemOne-CoreAI)184contains both LanguageBundles, each with `.aimodel`, `metadata.json` and `tokenizer/`:185 186| Path under `gpu-pipelined/` | role | bundle bytes | `main.mlirb` bytes |187|---|---|---:|---:|188| `openthai_systemone_decode_int8lin/` | ship | 1,068,353,811 | 1,039,655,099 |189| `openthai_systemone_decode_fp16/` | reference | 1,534,777,239 | 1,506,078,533 |190 191int8lin quantizes the linears per block of 32; the biased slot head, embeddings, conv1d and192norms stay fp16. `language.vocab_size = 256` describes the logits width because the sequential193engine allocates its output buffer from it. The input tokenizer still contains **248,339194tokens, including all 295 added tokens**. The `decision` metadata carries the slot count,195abstain slot, answer token, temperatures and layout. The source `config.json` is retained for196provenance.197 198Use the zoo's [extra-states runtime patch](https://github.com/john-rocky/coreai-model-zoo/blob/main/apps/coreai-pipelined-extra-states.patch)199for the hybrid's KV, conv and recurrent states, and `COREAI_CHUNK_THRESHOLD=1`. Both pipelined200and sequential engines were checked with Release tools from fork tag `0.2.4-zoo` (`f7a75ec`).201The [recipe](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/openthai-systemone/recipe.toml)202records the source revision and each graph's SHA-256.203 204## Reproduce205 206Run from the zoo checkout with the overlay environment; the oracle uses its own uv-managed207environment. `DEVELOPER_DIR` must select Xcode 27 for the Core AI tools.208 209```bash210python3 conversion/zoo_convert.py run openthai-systemone211python3 conversion/zoo_convert.py run openthai-systemone-fp16212 213uv run conversion/slot/oracle_slot.py \214 --out models/openthai-systemone/fixtures-openthai-systemone.json215 216python3 conversion/slot/readout_gate_slot.py \217 exports/openthai_systemone_decode_int8lin \218 models/openthai-systemone/fixtures-openthai-systemone.json \219 --transcript models/openthai-systemone/gate-openthai-systemone-readout-int8lin.json220 221python3 conversion/slot/engine_argmax_slot.py \222 exports/openthai_systemone_decode_int8lin \223 models/openthai-systemone/fixtures-openthai-systemone.json \224 --readout models/openthai-systemone/gate-openthai-systemone-readout-int8lin.json \225 --runner <fork>/.build/release/llm-runner \226 --engine pipelined --engine sequential \227 --transcript models/openthai-systemone/gate-openthai-systemone-engine-int8lin.json228```229 230Repeat the readout and engine commands with `fp16` paths for the reference. The231[exporter](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/export_openthai_systemone_decode_pipelined.py)232downloads the pinned snapshot itself. [Gate instructions](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/slot/README.md)233and [port notes](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/openthai-systemone-port.md)234record the oracle dependencies and runtime contract.235 236## License237 238Source Apache-2.0 (`iapp/OpenThai-SystemOne`); the bundles inherit it. The pinned source239snapshot has no license file, so `LICENSE` contains the canonical240[Apache License 2.0 text](https://www.apache.org/licenses/LICENSE-2.0.txt). The author's241inference files are downloaded by the oracle at gate time and are not included in the bundles.242 