FluidInference/kev-0.8b-coreml
Kev-0.8B for Core ML
Core ML export of jaredpalmer/kev-0.8b (Jared Palmer, Apache-2.0): Kev's LoRA folded into Qwen/Qwen3.5-0.8B-Base in fp32 with Kev's own loader, plus Kev's pointer head. Kev answers questions about a piece of text (a state): multiple choice, yes/no, or a score, with calibrated probabilities.
Files
fp16, GPU (cpuAndGPU), iOS 18 / macOS 15. Swift: KevFastManager in FluidUse (own Qwen BPE tokenizer, exact against Hugging Face tokenizers).
How the fused call works
The state's tokens come first, then every question packed end to end. A segment mask keeps each question from seeing the others: it is the attention mask, the local decay of the Gated DeltaNet layers, and the triangle of their WY solve, so every question restarts from the state exactly as if it were its own row. Everything except the delta-rule core runs once over all positions.
Fidelity
Kev's own benchmark, unchanged, scored the row packages against Kev's fp32 PyTorch model on the development splits:
The fused pass matches the row form in fp32 (0 flips, max |Δp| 3e-6 with every question packed three times), and the fp16 fused functions agree with the fp16 row packages on 191 questions (0 flips, max |Δp| 0.004).
Speed and memory against the original model
MacBook Pro M5 Pro (24 GB). Guess Who over 80 Wikipedia people (DBpedia-14 test split): each bio is one request with 12 yes/no questions, 960 decisions. The original is Kev's own checkpoint loader and serving path in PyTorch on the GPU (MPS).
PyTorch on a Mac runs Qwen3.5's Gated DeltaNet and causal conv through transformers' reference implementations (the flash-linear-attention / causal_conv1d kernels are not available there). Peak memory footprint is Activity Monitor's Memory; Core ML maps its weights from disk, so count them as resident for a conservative ~2.1 GB. This backbone does not suit the Neural Engine (one call: GPU 30.8 ms, CPU 186 ms, CPU + ANE 740 ms). A function left idle pays a 0.3–0.8 s re-setup on its next call; warm it before latency-sensitive work.
Conversion code and reports: FluidInference/mobius, models/computer-use/kev-0.8b/coreml.
