Team Ai
Modelpublic

FluidInference/intern-decision-0.8b-coreml

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
1likes74downloads
Model Card

Intern-Decision-0.8B for Core ML

Core ML export of internlm/Intern-Decision-0.8B (Shanghai AI Laboratory, Apache-2.0, snapshot 85a0cc5a): a typed-decision model fine-tuned from Qwen3.5-0.8B that answers a set of choice / score / noul (yes/no) questions about a JSON state in one prefill pass. The text path is exported here; the vision tower is not (text-only requests).

Files

PathWhat
L320_F8/DecisionRow_fp16.mlpackageRequests up to 320 tokens and 8 fields (the model card's three-field shape is 319 tokens).
L512_F8/DecisionRow_fp16.mlpackageUp to 512 tokens, 8 fields.
L512_F8/DecisionRow_w8.mlpackageSame, int8 per-channel weights (480 MB; 2 of 240 answers differ from the reference).
L1024_F16/DecisionRow_fp16.mlpackageUp to 1,024 tokens, 16 fields.
*/config.jsonBucket dimensions, marker / pad / answer-symbol ids, temperature, system prompt.
embeddings.f16Token embeddings (fp16, 248,320 × 1,024), gathered on the host.
tokenizer.jsonThe checkpoint's tokenizer (Qwen3.5 plus the <decision> token).

Inputs: hidden [1, L, 1024] (embedding rows of the right-padded prompt), cos / sin [L, 64] (RoPE tables for positions 0…L−1), field_onehot [F, L] (row i selects the token before field i's <decision>). Output: logits [F, 62] over the answer symbols; take the first n entries for a field with n options, softmax, divide log-probs by the temperature. fp16, GPU (cpuAndGPU), iOS 17 / macOS 14. Swift runtime: InternDecisionManager in FluidUse; conversion scripts in mobius (models/computer-use/intern-decision-0.8b/coreml).

What the package computes

Intern-Decision renders the request as one chat prompt: a fixed system prompt, a user turn with the state as JSON and the decision schema (one line per field, its options mapped to the answer symbols A–Z, a–z, 0–9), and an assistant turn that is a JSON skeleton with one <decision> token per field. The logits at the position immediately before each marker, restricted to that field's first n symbols, are its answer; the published temperature (2.7478 for 0.8B) rescales the restricted softmax without changing the argmax. Nothing is generated.

DecisionRow (decision_export.py) is one fixed-length request: the Qwen3.5 decoder from qwen35_export.py (the Kev-0.8B / Cua-S1-4B export), the final norm at the pre-marker positions selected with a one-hot map per field, and the 62 tied-embedding rows of the answer symbols. Token embeddings are gathered on the host (embeddings.f16); the host applies the restricted softmax and temperature. Prompt compilation and the chat template come from the checkpoint's own inference.py and tokenizer, so the token stream is the reference's.

PackageTokensFieldsFits
L256_F42564Jevbench easy/original (~220 tokens), AG News (p95 299)
L320_F83208the model card's three-field request shape
L384_F8, L512_F8384 / 5128WildJailBreak (p95 536), 62% of ToolACE
L768_F16, L1024_F16768 / 1,02416Typed Decision (5 fields, p50 852, max 1,223), ToolACE (max 886)

The fixed prompt (system prompt, headings, skeleton) is about 250 tokens, so the smallest three-field request is 319 tokens; 40% of Jevbench-Hard exceeds 1,024 tokens (max 4,074).

Fidelity

Reference: the checkpoint's DecisionEngine in fp32 on the Apple GPU (MPS), probabilities after temperature scaling, on the bundled accuracy suites from the Intern-Decision repo (benchmarks/accuracy-v1, shuffled with seed 0). The fp32 PyTorch wrapper matches the reference to 3.6e-6 (50 Typed Decision fields).

PackageSuitesDecisionsTop answer differsMax \Δp\
L512_F8 fp16Jevbench (3), ToolACE, AG News, WildJailBreak24000.006
L1024_F16 fp16Typed Decision (5 fields), Jevbench-Hard, ToolACE28000.007
L320_F8 fp16Jevbench easy/original, AG News, WildJailBreak10000.006
L256_F4 fp16Jevbench easy/original, AG News, WildJailBreak10000.006
L512_F8 int8 (per-channel)same as L512_F8 fp1624020.047

Reference and Core ML accuracy against the suite labels are identical on every subset (reports in the mobius directory). int4 per-block compression needs an iOS 18 deployment target and was not built.

Latency

One request, 319 input tokens, three fields (choice, yes/no, score), the model card's RTX 4090 shape. M5 Pro (24 GB), warmed, bench.py; Core ML on CPU_AND_GPU, times include the host embedding gather and RoPE tables.

Runtimep50p95
Core ML fp16 L320_F857 ms59 ms
Core ML fp16 L384_F866 ms67 ms
Core ML fp16 L512_F888 ms96 ms
Core ML fp16 L768_F16133 ms145 ms
Core ML fp16 L1024_F16182 ms190 ms
PyTorch MPS bf16 (checkpoint inference.py)150 ms170 ms
PyTorch MPS fp32177 ms191 ms

The pass is compute-bound and scales with the bucket, not the request, so ship the smallest bucket that fits. int8 weights leave GPU time unchanged (88 ms at L512_F8) and halve the package (955 → 480 MB). ComputeUnit.ALL matches CPU_AND_GPU: the Gated DeltaNet backbone does not run on the Neural Engine (see the Kev-0.8B notes). The model card reports 34 ms for this request on an RTX 4090.