Team Ai
Modelpublic

FluidInference/kev-0.8b-coreml

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
4likes153downloads
Model Card

Kev-0.8B for Core ML

Core ML export of jaredpalmer/kev-0.8b (Jared Palmer, Apache-2.0): Kev's LoRA folded into Qwen/Qwen3.5-0.8B-Base in fp32 with Kev's own loader, plus Kev's pointer head. Kev answers questions about a piece of text (a state): multiple choice, yes/no, or a score, with calibrated probabilities.

Files

PathWhat
fused/KevFused.mlpackageMultifunction package, weight-shared: fused_S<state>_P<packed>_B16_K16 for states of 32–384 tokens and up to 192 packed question tokens. One call reads the state and answers all of a request's questions (up to 16 answers, 16 options each).
fused/config.jsonDimensions, special tokens, temperature.
L512_K16/KevRow_fp16.mlpackageOne question per call (state + question ≤ 512 tokens, ≤ 16 options), for questions that do not fit the fused path.
L1024_K80/KevRow_fp16.mlpackageOne question per call up to 1,024 tokens and 80 options (e.g. 77-way intent questions).
L512_K16/embeddings.f16Token embeddings (fp16, 248,320 × 1,024), gathered on the host.
tokenizer.jsonQwen3.5 tokenizer.

fp16, GPU (cpuAndGPU), iOS 18 / macOS 15. Swift: KevFastManager in FluidUse (own Qwen BPE tokenizer, exact against Hugging Face tokenizers).

How the fused call works

The state's tokens come first, then every question packed end to end. A segment mask keeps each question from seeing the others: it is the attention mask, the local decay of the Gated DeltaNet layers, and the triangle of their WY solve, so every question restarts from the state exactly as if it were its own row. Everything except the delta-rule core runs once over all positions.

Fidelity

Kev's own benchmark, unchanged, scored the row packages against Kev's fp32 PyTorch model on the development splits:

SuiteQuestionsTop answer differsKev fp32PublishedCore ML fp16
transfer-v476410.64790.6480.6463
decision-v71,46820.82670.8270.8252

The fused pass matches the row form in fp32 (0 flips, max |Δp| 3e-6 with every question packed three times), and the fp16 fused functions agree with the fp16 row packages on 191 questions (0 flips, max |Δp| 0.004).

Speed and memory against the original model

MacBook Pro M5 Pro (24 GB). Guess Who over 80 Wikipedia people (DBpedia-14 test split): each bio is one request with 12 yes/no questions, 960 decisions. The original is Kev's own checkpoint loader and serving path in PyTorch on the GPU (MPS).

Core ML (this repo, fp16, GPU)Original, PyTorch bf16Original, PyTorch fp32 (default)
80 bios × 12 questions2.98 s98.6 s123.9 s
per bio (12 answers, one call)36.6 ms1,070 ms1,287 ms
peak memory footprint0.67 GB (+1.45 GB weights mapped from disk)6.6 GB9.5 GB
weights1.45 GB (fused package + embeddings)1.79 GB (bf16 base + LoRA/head)same
answers vs original fp320 of 960 differ3 of 960 differreference

PyTorch on a Mac runs Qwen3.5's Gated DeltaNet and causal conv through transformers' reference implementations (the flash-linear-attention / causal_conv1d kernels are not available there). Peak memory footprint is Activity Monitor's Memory; Core ML maps its weights from disk, so count them as resident for a conservative ~2.1 GB. This backbone does not suit the Neural Engine (one call: GPU 30.8 ms, CPU 186 ms, CPU + ANE 740 ms). A function left idle pays a 0.3–0.8 s re-setup on its next call; warm it before latency-sensitive work.

Conversion code and reports: FluidInference/mobius, models/computer-use/kev-0.8b/coreml.