Team Ai
Modelpublic

mlboydaisuke/Granite-Embedding-97M-Multilingual-R2-CoreAI

sourceHugging Faceapache-2.0updated 21h agoView on Hugging Face
3likes285downloads
README.md221 linesDownload Raw Back to root
1---2library_name: coreai3license: apache-2.04base_model: ibm-granite/granite-embedding-97m-multilingual-r25tags:6  - coreai7  - sentence-similarity8  - feature-extraction9  - apple-silicon10  - on-device11  - modernbert12language:13  - multilingual14  - ja15  - en16pipeline_tag: sentence-similarity17---18 19Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's `coreai-torch` (LLMs: `coreai.llm.export`) into `.aimodel` bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol ([apple-silicon-llm-bench](https://github.com/john-rocky/apple-silicon-llm-bench), macOS 27 beta, 2026-06).20 21<!-- gen-cards:devicemark begin (managed by scripts/gen-cards + tools/devicemark_row.py — edit cards.json, not this block) -->22This model has no row on [DeviceMark](https://devicemark.github.io/), the on-device LLM leaderboard.23<!-- gen-cards:devicemark end -->24 25# Granite-Embedding-97M-Multilingual-R2 — Core AI export26 27Zoo card, recipe and gate transcript: [coreai-model-zoo/models/granite-embedding-97m](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/granite-embedding-97m/README.md).28 29Sample app: [CoreAISearchiOS](https://github.com/john-rocky/coreai-samples/tree/main/CoreAISearchiOS) downloads `ios/fp32-s128` from this repo and searches your own notes by meaning, 16.5 ms per query on an iPhone 18 Pro (iOS 27.2, 2026-10-09).30 31IBM's 97M-parameter **multilingual text embedder** — a ModernBERT encoder, 384-d CLS-pooled32unit vectors, Japanese and English among its languages — as a static `.aimodel` for macOS 2733and iOS 27, with bundles compiled ahead of time for the iPhone 17 Pro beside it.34[`ibm-granite/granite-embedding-97m-multilingual-r2`](https://huggingface.co/ibm-granite/granite-embedding-97m-multilingual-r2)35(Apache-2.0, revision `835ad1408…`) is the **smallest embedder in this catalog** (390 MB fp32,36against 1.2 GB for EmbeddingGemma-300m and 1.1 GB for Qwen3-Embedding-0.6B) and its **first37encoder-architecture one** — every other embedder here is a causal decoder run as an encoder.38Its retrieval quality relative to those three was **not** measured here: the fixture set below39is a parity instrument (35 texts, 4 queries, 12 documents), not a benchmark.40 41**This is an encoder, not a generator** — one forward over the right-padded grid returns one42unit vector. No autoregressive loop, no KV cache, no LM head. It runs as a plain `.aimodel`43through raw `AIModel.run` (like the vision encoders), not the pipelined generate engine.44 45Architecture (`model_type: modernbert`): 12 layers, hidden 384, 12 heads × 32, GLU MLP 153646(SiLU), vocabulary 180,000, biasless everything (attention, MLP, LayerNorm ε 1e-5). Global47attention at layers **0, 3, 6, 9** (RoPE θ 150,000); the other eight are **local**, a sliding48window of inclusive radius 64 (129 keys per interior query, RoPE θ 160,000). Layer 0 has no49attention pre-norm (the embedding LayerNorm serves). Pooling is CLS → L2 normalize, both in the50graph.51 52## Graph contract53 54```55input  "input_ids"       [1, S]    int32   right-padded to the grid S with 17993556input  "attention_mask"  [1, S]    int32   1 over real tokens, 0 over padding57output "embedding"       [1, 384]  fp32    CLS-pooled, L2-normalized58S = 128 or 512 (export-time choice); batch = 159```60 61**Host recipe** — the tokenizer is the whole contract, and the stock one is not enough:62- **No prefix, no stripping, no normalization.** Query and document prompts are both empty in63  the checkpoint. Raw whitespace is kept: sentence-transformers strips text before tokenizing,64  the upstream README's `AutoTokenizer` path does not, and the two disagree on `"  東京駅から…\n"`.65  The reference is the raw path.66- Tokenize with the pinned `tokenizer.json`: regex `Split(Isolated)` → `ByteLevel` (no prefix67  space) → byte BPE with **`ignore_merges = true`** (a whole pre-token that is in the vocabulary68  wins; ` ક` is token 2999, not three). A BPE that ignores the flag tokenizes differently.69- Truncate the **body to S−2**, then wrap: `[CLS 179934] body… [SEP 179938]`, right-pad with70  **PAD 179935** and mask 0. Truncating after adding the specials loses SEP; padding with 0 is a71  different token. Both are silent.72- Similarity = dot product (unit vectors). Dimension truncation is not a property of this model.73 74[`conversion/granite_embedding/_granite_tokenizer.py`](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/granite_embedding/_granite_tokenizer.py) is that recipe with no HF import, and75`host/GraniteTokenizer.swift` in this repo the same recipe in Foundation-only Swift; the gate76holds both to `AutoTokenizer` exactly (ids and masks) over **681 texts × 2 grids = 1,362 cases**77including every added token in five boundary contexts, and proves four mutations are caught78(pad 0 / lose SEP / strip / `ignore_merges=false`).79 80## Measured81 82**iPhone 17 Pro** (iPhone18,1), iOS 27.0 build **24A437**, the compiled `h18p` bundles loaded by83the native `AIModel` loader, GPU-preferred (MPSGraph/Metal plan). Every row: 35 HF texts, the84gate below, 105 warm samples, thermal state fair before and after, caches retained (so "first"85is process-first, not cache-cold). Peak footprint is the whole app process, tokenizer and file86hashing included. Measured 2026-09-19.87 88| Variant | S | Gate | Min cosine vs HF | Max \|err\| | Load | First after load | **Warm median** | Peak footprint |89|---|---:|---|---:|---:|---:|---:|---:|---:|90| fp32 | 128 | 35/35 | 0.999999999999407 | 1.97e-7 | 81 ms | 23.1 ms | **5.54 ms** | 640 MB |91| fp32 | 512 | 35/35 | 0.999999999999486 | 2.38e-7 | 586 ms | 39.9 ms | **20.99 ms** | 640 MB |92| w8 / fp32 table | 128 | 35/35 | 0.999410 | 5.67e-3 | 61 ms | 25.3 ms | 6.64 ms | 555 MB |93| w8 / fp32 table | 512 | 35/35 | 0.999410 | 5.67e-3 | 447 ms | 138.7 ms | 23.07 ms | 553 MB |94 95Each row matched 4/4 retrieval top-1s with 0 clear-pair flips and 0 repeat drift.96 97**Mac** (M4 Max, Mac16,9), macOS 27.0 build 26A428, the JIT `.aimodel`, GPU-preferred, fp32.98The driver refused to run while any foreign accelerator job was present; 105 warm samples.99 100| S | Gate | Min cosine vs HF | Max \|err\| | Load | First after load | **Warm median** |101|---:|---|---:|---:|---:|---:|---:|102| 128 | 35/35 | 0.99999999999967 | 2.98e-7 | 481 ms | 642 ms | **4.14 ms** |103| 512 | 35/35 | 0.99999999999887 | 2.98e-7 | 470 ms | 264 ms | **4.89 ms** |104 105The Mac h16c AOT twin also passed 70/70 (same numerics), but its timings were taken with another106lane's GPU evaluation running and are not reported. The w8 variant on Mac is gated on **CPU107only** (min cosine 0.999410, max |err| 5.67e-3, ranking exact); Mac GPU for w8 was not run.108 109**iPhone 18 Pro** (iPhone19,2, h19p), iOS 27.0 build **24A437**, the JIT `.aimodel` bundles in `ios/`110(the same files as `macos/`), loaded 2026-09-26 by the zoo's DecideGate app in its load-only mode,111GPU-preferred, without the increased-memory entitlement. Each first load was the first after a fresh112install of the app. The call is one run on all-zero inputs. Each first load wrote a specialization of113about the bundle's size into the app container, and the load after a relaunch reused it. One measurement114per graph115([knowledge/jit-distribution.md](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/jit-distribution.md)).116 117| JIT bundle in `ios/` | MB | first load | first call | load after relaunch |118|---|---:|---:|---:|---:|119| `fp32-s128/granite97m_fp32_s128_bound.aimodel` | 390 | 1.01 s | 619 ms | 0.26 s |120| `fp32-s512/granite97m_fp32_s512_bound.aimodel` | 390 | 0.40 s | 88 ms | 0.26 s |121| `w8-fp32table-s128/granite97m_w8_fp32table_s128.aimodel` | 305 | 0.63 s | 132 ms | 0.21 s |122| `w8-fp32table-s512/granite97m_w8_fp32table_s512.aimodel` | 306 | 0.33 s | 93 ms | 0.21 s |123 124The iPhone gate above (the iPhone 17 Pro rows) ran the `h18p` export, now in `ios-h18p/`. The JIT IR in125`ios/` was only loaded and called once on the iPhone 18 Pro. That call returned a 384-value float32126embedding with no non-finite values. On 2026-10-09 the sample app above embedded 9,859 sentences on an127iPhone 18 Pro (iOS 27.2, 24B5099f) with `ios/fp32-s128`, GPU preferred: the 9,859 × 384 float32 matrix has the128same SHA-256 as the one the `h18p` bundle produced on the iPhone 17 Pro on 2026-09-20 (bit-identical), so the129five example queries return the same top 5 with the same scores. First load 0.44 s, first call 612 ms, load130after a relaunch 0.054 s, 4.91 ms per sentence over the 9,859, one run.131 132The fixed grid computes every position, so pick the smallest grid that covers the text: S=128133for queries and short notes, S=512 for passages. **fp32 is the default.** w8 is a storage134option only — 22% smaller, not faster here — because the 180,000×384 fp32 vocabulary table is135276 MB of the bundle and palettization touches the 48 linear weights alone.136 137## Numerics gate138 139One gate at every stage, the oracle being official HF eager CPU fp32 (transformers 4.57.6):140per text cosine ≥ 0.999, max element error ≤ 0.02, L2-norm error ≤ 0.002; per query exact141top-1 over the 12 documents, retrieval-score error ≤ 0.01, and no inversion of any document142pair the oracle separates by ≥ 0.001; repeat drift ≤ 1e-6. A wrong-pairing control (every vector143matched to the wrong text) must FAIL.144 145- **Authoring** (`gate_granite_authoring.py`): the re-authored graph against every one of the 13146  saved hidden states, max |err| ≤ **1e-4** at fp32, both grids. Five mutations must trip it:147  all-global, all-local, ignore-padding and mean-pooling fail the embedding gate; a local radius of148  **63 instead of 64** passes the embedding gate (cos 0.99995) and fails only the layer gate —149  which is why the layer gate exists. Whole-model **fp16 fails** this layer gate on both grids.150- **Export**: the torch-exported, decomposed graph is gated before conversion, on both grids.151- **Runtime**: Mac CPU and GPU (JIT), Mac h16c AOT, iPhone h18p AOT — the tables above; iPhone 18 Pro152  JIT: load and one zero-input call only.153- **w8**: the same gate at prepared, finalized and decomposed stages, 48 `lut_to_dense` ops154  counted, palettes hashed; the iOS w8 export reuses the Mac palettes byte for byte.155 156[`models/granite-embedding-97m/gate-granite-embedding-97m.json`](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/granite-embedding-97m/gate-granite-embedding-97m.json) in the zoo is the transcript: the eight runtime rows157(min cosine, max error, retrieval, timings, device/OS build), the tokenizer gate and the158authoring gate, each with the sha256 of the full record it summarizes.159 160## ⬇️ Bundle161 162This repo — one folder per variant, each self-contained: the bundle, `tokenizer/`, `reference.json` (the16335 HF fixtures with ids, masks and embeddings — the parity test) and `provenance/` (export164manifest with per-file sha256, the runtime gate record). `coreai-kit.json` at the root maps165platform → folder.166 167| Folder | Platform | Format | Bundle | Bytes |168|---|---|---|---|---:|169| `macos/fp32-s512/` **(default)** | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |170| `macos/fp32-s128/` | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |171| `ios/fp32-s512/` **(default)** | iOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |172| `ios/fp32-s128/` | iOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |173| `macos/w8-fp32table-s512/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |174| `macos/w8-fp32table-s128/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |175| `ios/w8-fp32table-s512/` | iOS 27 | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |176| `ios/w8-fp32table-s128/` | iOS 27 | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |177| `ios-h18p/fp32-s512/` | iOS 27, **h18p only** | AOT `.aimodelc` | `granite97m_fp32_s512_bound.h18p.aimodelc` | 390,308,788 |178| `ios-h18p/fp32-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_fp32_s128_bound.h18p.aimodelc` | 390,081,410 |179| `ios-h18p/w8-fp32table-s512/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s512_r02.h18p.aimodelc` | 305,479,184 |180| `ios-h18p/w8-fp32table-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s128_r02.h18p.aimodelc` | 305,251,774 |181 182`ios/` holds the same JIT bundles as `macos/`; every iPhone generation specializes them on its first183load. The `ios-h18p/` bundles moved there from `ios/` in revision `a27dc73e` (2026-09-26). They are184compiled for one device architecture (`h18p`, the iPhone 17 Pro) with185`xcrun coreai-build compile --platform iOS --min-deployment-version 27.0 --preferred-compute gpu186--architecture h18p` (coreai-build 3600.83.1), from a separate iOS export whose IR is reproducible187from the recipe but not shipped. The iPhone 18 Pro refuses an h18p bundle with188`incompatibleCompiledAssetArchitecture`. **Never load an `ios-h18p/` bundle on a Mac.**189 190Convert yourself: [`conversion/granite_embedding/`](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/granite_embedding/README.md)191— five staged scripts; [`recipe.toml`](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/granite-embedding-97m/recipe.toml) names the commands.192 193## CoreAIKit (Swift)194 195**Not enrolled** in the kit catalog. The kit's `TextEmbedder` pads with 0, truncates after adding196the special tokens (losing SEP), applies its own BPE without `ignore_merges`, discovers a single197`.aimodel`, and has no grid / architecture selection — every one of those is wrong for this198model. Running it today means: the Swift tokenizer from this repo's `host/` folder, a fixed199grid, `AIModel` on the platform's folder. Enrolling it needs a `textEmbedding` driver that takes200the pad id, a SEP-preserving truncation, a per-platform variant path and an AOT-aware loader —201tracked as maintainer work, not a blocker on the bundle.202 203## The port in one lesson: gate the layers, not just the vector204 205ModernBERT's alternating local/global attention is the whole risk. The config says206`local_attention: 128`; the executed window is inclusive `|i − j| ≤ 64` — 129 keys — and a207window of 63 reproduces the final embedding to cos 0.99995 while every hidden state past layer 1208is wrong. Only a per-layer oracle catches it. Three more things the raw checkpoint settles that209the modeling file hides: layer 0 has no attention norm (adding one loads a missing weight),210the two RoPE thetas are per-layer-kind, and the CLS/L2 head needs an explicit `clamp_min`211epsilon because the converter's `F.normalize` decomposition drops it.212 213## License and limits214 215Apache-2.0 at the pinned upstream revision; this repo carries IBM's unmodified card as216`UPSTREAM_README.md` and a `LICENSE-NOTE.md` listing the changes (static graph, in-graph217pooling, optional w8 palettes, h18p compile). Not tested: phones other than the iPhone 17 Pro (the218h18p gate) and the iPhone 18 Pro (JIT load and one call), other OS builds, the JIT bundles' embedding219parity on an iPhone, the Mac GPU with w8, the Neural Engine, dynamic or batched shapes, S > 512, languages beyond the JA/EN220fixtures, retrieval quality on a benchmark, sustained thermals, true cache-cold load.221