OpenInfinity/embeddinggemma-2-coreml-int8
EmbeddingGemma 2 — Core ML int8 (text + image, compiled)
<a id="english"></a>
English
The text and image encoders of google/embeddinggemma-2, based on FluidInference/embeddinggemma-2-coreml. We changed three things:
- int8 weights. Both models have int8 weights, which halves their size. Activations stay fp16.
- int8 token table. The token embedding table is int8 with one fp16 scale per row, which also halves it.
- Precompiled. Both models are compiled to
.mlmodelc, so an app can load them without callingMLModel.compileModel.
On our tests the retrieval quality is the same as the fp32 reference (see Validation). The text model runs entirely on the Apple Neural Engine. The audio encoder is not included.
The output is the same 768-d, L2-normalized embedding as SentenceTransformer("google/embeddinggemma-2").encode(...). Matryoshka truncation to 512, 256 or 128 dimensions works as in the original: keep the first N values and re-normalize.
Files
Text-only use needs config.json and text/ (about 313 MB). Image search also needs vision/ (about 158 MB). Sizes on this page are in MB = 10⁶ bytes (FluidInference's README uses MiB).
What changed compared to FluidInference's package
Functions, inputs and outputs are unchanged, so FluidInference's README and its Swift code (FluidUse, EmbeddingGemma2Manager) describe how to call these models. The only difference in calling them is the table lookup (step 3 below). The file layout, config.json and the table format do differ, so FluidUse cannot load this repo without changes.
How the int8 models were made. The coremltools int8 API (linear_quantize_weights) rejects multifunction packages, so we:
- split each package into one single-function model per function (via the MIL program);
- quantized each model to int8;
- merged them back with
ct.utils.save_multifunction, which stores identical weights only once; - compiled the result with
ct.utils.compile_model.
The compiled models give the same output as the .mlpackage they came from: cosine 1.000000, max difference 0.
Use (text)
- Prefix. Add the task prefix, e.g.
title: none | text:for documents andtask: search result | query:for queries. All prefixes are inconfig.json. - Tokenize. Use
text/tokenizer.json. The sequence must start with<bos>(2) and end with<eos>(1); the tokenizer adds both. If it is longer than 512 tokens, keep the first 511 and end with<eos>. - Look up rows. For each token id, read its row from
embeddings.int8and multiply byscale[id] × √512(√512 ≈ 22.627417). Pass the result as fp16. - Run. Call the smallest
embed_Sthat fits, S ∈ {32, 48, 64, 128, 256, 512}. Inputs:inputs_embeds[1, S, 512] andattention_mask[1, S], 1 for tokens, right padding. Output:embedding[1, 768]. To embed up to eight short texts in one call, usepack_256; see FluidInference's README.
import CoreML
// Load one function of the compiled model (macOS 15 / iOS 18 or later).
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
config.functionName = "embed_128"
let model = try MLModel(contentsOf: textURL, configuration: config) // text/EmbeddingGemma2Text.mlmodelc
// Row lookup for one token: int8 row × per-row scale × √512.
func row(_ id: Int, table: UnsafePointer<Int8>, scales: UnsafePointer<Float16>, into out: UnsafeMutablePointer<Float16>) {
let s = Float(scales[id]) * 22.627417
for j in 0..<512 { out[j] = Float16(Float(table[id * 512 + j]) * s) }
}import numpy as np, coremltools as ct
from tokenizers import Tokenizer
tok = Tokenizer.from_file("text/tokenizer.json")
table = np.memmap("text/embeddings.int8", dtype=np.int8, mode="r").reshape(262144, 512)
scale = np.fromfile("text/embeddings.int8.scale.f16", dtype=np.float16).astype(np.float32)
ids = tok.encode("task: search result | query: 周末去哪里爬山").ids # the tokenizer adds <bos> (2) and <eos> (1)
x = np.zeros((1, 64, 512), np.float16); m = np.zeros((1, 64), np.float16)
x[0, :len(ids)] = table[ids] * scale[ids, None] * 22.627417; m[0, :len(ids)] = 1
model = ct.models.CompiledMLModel("text/EmbeddingGemma2Text.mlmodelc",
compute_units=ct.ComputeUnit.CPU_AND_NE, function_name="embed_64")
v = model.predict({"inputs_embeds": x, "attention_mask": m})["embedding"][0]
v /= np.linalg.norm(v)Use (images)
The steps are the same as in FluidInference's README:
- Resize the image, keeping its aspect ratio.
- Cut it into 16 px patches.
- Run
vision_70,vision_140orvision_280. - Put the resulting tokens into
<bos> <|image> … <image|> <eos>and embed that sequence with the text model.
The table rows for the special tokens come from embeddings.int8, as in step 3 above.
Validation
Tests ran on one MacBook Pro (M3 Max, macOS 26.7). Quality was measured with the Neural Engine (cpuAndNeuralEngine) for both text and images. "int8" means int8 models and the int8 token table, which is what this repo ships.
Similarity to the fp32 reference (sentence-transformers, PyTorch fp32):
Retrieval quality:
About the test sets:
- Simulated personal knowledge base. 2,829 chunks of notes, Wikipedia pages, open-source documentation and code, mixed Chinese and English, with 113 queries written by an LLM.
- MLQA. 200 questions.
- Screenshots. 100 screenshots: 60 synthetic app screens and 40 Wikipedia pages, half Chinese and half English, with 100 queries. OCR is Apple Vision.
- Score fusion.
z(image score) + 0.5 · z(OCR-text score), where z standardizes scores per query.
Speed (M3 Max, one call, warm):
An image needs one embed_S call after the image model. The GPU times are short runs at full GPU clock; under sustained load the GPU slowed down (see Energy).
Neural Engine placement (MLComputePlan): every op runs on the Neural Engine in all seven text functions (3,862 ops each, 3,863 in pack_256) and in vision_70 and vision_140 (2,389 ops each). We did not check vision_280 on the Neural Engine.
int8 has about the same speed as fp16. The gain is size: text model and table 552 → 281 MB, image model 306 → 155 MB.
Energy. Same machine, on AC power in High Power mode. The 100 screenshots ran in a loop for 30 s, each setup twice, and idle power was subtracted. The text model ran on the Neural Engine in every case.
- The Neural Engine uses about 7× less energy per screenshot. Indexing 10,000 screenshots costs about 1.5 Wh on the Neural Engine and about 10 Wh on the GPU; the battery of this 16-inch MacBook Pro holds 100 Wh.
- The GPU is faster only in short runs. It drew about 54 W at first. After about 1.5 minutes of GPU load, the system lowered the GPU clock from about 1.3 GHz to 0.6–0.8 GHz. The image model then took 146 ms per screenshot, slower than the Neural Engine, which stayed at 123 ms.
- Recommendation: run
vision_70andvision_140on the Neural Engine (cpuAndNeuralEngine), especially for background indexing. Use the GPU forvision_280, which we did not try on the Neural Engine. On the GPU the int8 image tokens match FluidInference's fp16 package (cosine ≥ 0.999 for every image function on test inputs).
Limitations
- Tested on one machine only: M3 Max with macOS 26.7. The models need macOS 15 / iOS 18 or later, because they are multifunction models.
- First load is slow. The first load on a device compiles each function for the Neural Engine. FluidInference reports several minutes for the seven text functions. Core ML caches the result. Apps should warm the functions up in the background, because a
.mlmodelccannot be compiled for the Neural Engine ahead of time. - Fixed shapes. Inputs longer than 512 tokens must be truncated.
- `vision_280` was not tried on the Neural Engine. Run it on the GPU.
- Small test sets. The test sets are small (100–200 queries each) and some queries are LLM-written. Treat the numbers as a check that int8 loses nothing, not as a benchmark.
- No audio. Audio is not included. In our tests EmbeddingGemma 2's audio embeddings were weak for Chinese speech and for cross-language search, and speech-to-text followed by text embedding worked better. If you need the audio encoder, use FluidInference's fp16 package.
License and attribution
- License. Apache 2.0, the same as the source models.
- Google. EmbeddingGemma 2 is by Google DeepMind. Deployments must follow the Gemma Prohibited Use Policy.
- FluidInference. The Core ML conversion (fp16-safe RMSNorm, functions, packing, image pipeline) is FluidInference's.
- This repo. We added the int8 models, the int8 token table, the compilation and the tests.
- No affiliation. This repo is not affiliated with Google or FluidInference.
<a id="中文"></a>
中文
本仓库提供 google/embeddinggemma-2 的文本编码器和图像编码器, 基于 FluidInference/embeddinggemma-2-coreml 的 Core ML 版本。 我们做了三处改动:
- 模型权重 int8:两个模型的权重都压成了 int8,体积减半;计算仍用 fp16。
- 词表 int8:词表也压成了 int8,每行配一个 fp16 缩放系数,体积同样减半。
- 预先编译:两个模型都编译成了
.mlmodelc,App 可以直接加载,不用再调用MLModel.compileModel。
在我们的测试中,检索效果和 fp32 原版一样(见下文“验证”)。文本模型完全在 Apple 神经网络引擎(ANE)上运行。 本仓库不含音频编码器。
输出和 SentenceTransformer("google/embeddinggemma-2").encode(...) 相同:768 维、已归一化的向量。 如果想用 512、256 或 128 维,和原版一样取前 N 维,再重新归一化即可。
文件
只用文本:需要 config.json 和 text/,约 313 MB。还要搜图片:再加 vision/,约 158 MB。 本页的 MB 指 10⁶ 字节(FluidInference 的 README 用的是 MiB)。
和 FluidInference 原版的区别
函数名、输入和输出都没变,调用方法可以直接参考 FluidInference 的 README 和它的 Swift 代码 (FluidUse 里的 EmbeddingGemma2Manager)。调用上唯一的区别是查词表(见下面第 3 步)。但文件布局、config.json 和词表格式都不同,FluidUse 不改代码无法直接加载本仓库。
int8 模型是怎么做的:coremltools 的 int8 压缩接口(linear_quantize_weights)不支持多函数模型,所以我们分四步:
- 通过 MIL 程序把每个函数拆成单独的模型;
- 逐个压成 int8;
- 用
ct.utils.save_multifunction合并回去,相同的权重只存一份; - 用
ct.utils.compile_model编译。
编译后的模型和编译前的 .mlpackage 输出完全一致:余弦相似度 1.000000,最大差值 0。
用法(文本)
- 加前缀:加上任务前缀,例如文档用
title: none | text:,查询用task: search result | query:。所有前缀都在config.json里。 - 分词:用
text/tokenizer.json分词。序列必须以<bos>(2)开头、以<eos>(1)结尾,分词器会自动加上这两个。 超过 512 个 token 时,保留前 511 个,最后接<eos>。 - 查词表:对每个 token id,从
embeddings.int8取出对应的行,乘以scale[id] × √512(√512 ≈ 22.627417),转成 fp16。 - 运行:调用能放下这段文本的最小
embed_S,S ∈ {32, 48, 64, 128, 256, 512}。 - 输入是
inputs_embeds[1, S, 512] 和attention_mask[1, S](有 token 的位置为 1,补齐放在右侧); - 输出是
embedding[1, 768]; - 想一次算最多 8 段短文本,用
pack_256,见 FluidInference 的 README。
import CoreML
// 加载编译好的模型中的一个函数(需要 macOS 15 / iOS 18 或更高版本)
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
config.functionName = "embed_128"
let model = try MLModel(contentsOf: textURL, configuration: config) // text/EmbeddingGemma2Text.mlmodelc
// 查一个 token 的行:int8 行 × 该行的缩放系数 × √512
func row(_ id: Int, table: UnsafePointer<Int8>, scales: UnsafePointer<Float16>, into out: UnsafeMutablePointer<Float16>) {
let s = Float(scales[id]) * 22.627417
for j in 0..<512 { out[j] = Float16(Float(table[id * 512 + j]) * s) }
}import numpy as np, coremltools as ct
from tokenizers import Tokenizer
tok = Tokenizer.from_file("text/tokenizer.json")
table = np.memmap("text/embeddings.int8", dtype=np.int8, mode="r").reshape(262144, 512)
scale = np.fromfile("text/embeddings.int8.scale.f16", dtype=np.float16).astype(np.float32)
ids = tok.encode("task: search result | query: 周末去哪里爬山").ids # 分词器会自动加上 <bos>(2)和 <eos>(1)
x = np.zeros((1, 64, 512), np.float16); m = np.zeros((1, 64), np.float16)
x[0, :len(ids)] = table[ids] * scale[ids, None] * 22.627417; m[0, :len(ids)] = 1
model = ct.models.CompiledMLModel("text/EmbeddingGemma2Text.mlmodelc",
compute_units=ct.ComputeUnit.CPU_AND_NE, function_name="embed_64")
v = model.predict({"inputs_embeds": x, "attention_mask": m})["embedding"][0]
v /= np.linalg.norm(v)用法(图片)
和 FluidInference 的 README 相同:
- 按原比例缩放图片;
- 切成 16 px 的图块;
- 运行
vision_70、vision_140或vision_280; - 把得到的 token 放进
<bos> <|image> … <image|> <eos>,再用文本模型算出向量。
特殊 token 的行同样从 embeddings.int8 里查,方法同上面第 3 步。
验证
在一台 MacBook Pro(M3 Max,macOS 26.7)上测试。效果测试时,文本和图像都跑在神经网络引擎上(cpuAndNeuralEngine)。 下面的“int8”指 int8 模型加 int8 词表,也就是本仓库发布的版本。
和 fp32 原版的相似度(参照:sentence-transformers,PyTorch fp32):
检索效果:
测试集说明:
- 模拟个人知识库:2,829 个文本块,内容包括笔记、维基百科页面、开源文档和代码,中英混合;113 条查询由大模型编写。
- MLQA:200 个问题。
- 截图:100 张,其中 60 张合成的软件界面、40 张维基百科页面,中英各半;查询 100 条;文字识别用 Apple Vision。
- 分数融合:
z(图像分数) + 0.5 · z(OCR 文字分数),z 表示在每条查询内把分数标准化。
速度(M3 Max,单次调用,已预热):
每张图在图像模型之后,还要调用一次 embed_S。GPU 的耗时是 GPU 满频时短时间跑出来的;持续运行时 GPU 会变慢(见下面的“能耗”)。
算子分配(MLComputePlan):7 个文本函数(每个 3,862 个算子,pack_256 是 3,863 个)以及 vision_70、vision_140 (每个 2,389 个算子)的所有算子都在神经网络引擎上。vision_280 没有在神经网络引擎上检查。
int8 的速度和 fp16 基本一样,好处在体积:文本模型加词表从 552 MB 降到 281 MB,图像模型从 306 MB 降到 155 MB。
能耗:同一台机器,接电源,高性能模式。100 张截图循环跑 30 秒,每种配置跑两轮,扣除空闲功耗。文本模型都在神经网络引擎上。
- 神经网络引擎每张截图的能耗约为 GPU 的 1/7:给 1 万张截图建索引,在神经网络引擎上约耗 1.5 Wh,在 GPU 上约 10 Wh; 这台 16 英寸 MacBook Pro 的电池是 100 Wh。
- GPU 只在短时间内更快:一开始 GPU 功率约 54 W。GPU 持续工作约 1.5 分钟后,系统把 GPU 频率从约 1.3 GHz 降到 0.6–0.8 GHz, 图像模型每张要 146 ms,比神经网络引擎还慢;神经网络引擎一直稳定在 123 ms。
- 建议:
vision_70和vision_140在神经网络引擎上运行(cpuAndNeuralEngine),后台批量建索引时尤其如此。vision_280没有在神经网络引擎上试过,请用 GPU。在 GPU 上,int8 图像模型输出的 token 和 FluidInference 的 fp16 版本一致 (用测试输入检查,每个图像函数的余弦相似度都 ≥ 0.999)。
局限
- 只在一台机器上测过:M3 Max,macOS 26.7。模型是多函数模型,需要 macOS 15 / iOS 18 或更高版本。
- 第一次加载很慢:在每台设备上第一次加载时,系统要把每个函数编译给神经网络引擎;FluidInference 报告 7 个文本函数共需几分钟。 Core ML 会缓存编译结果。
.mlmodelc没法提前编译到神经网络引擎这一步,所以建议 App 在后台预热各个函数。 - 形状固定:超过 512 个 token 的输入必须截断。
- `vision_280` 没有在神经网络引擎上试过:请在 GPU 上运行。
- 测试集很小:每组 100–200 条查询,部分查询由大模型编写。这些数字只用来确认 int8 没有损失效果,不能当作基准成绩。
- 不含音频:在我们的测试中,EmbeddingGemma 2 的音频向量对中文语音和跨语言搜索效果较弱,先转文字再算文本向量效果更好。 如果需要音频编码器,请使用 FluidInference 的 fp16 版本。
许可证和来源
- 许可证:Apache 2.0,和原模型相同。
- Google:EmbeddingGemma 2 由 Google DeepMind 发布。使用时须遵守 Gemma 禁止用途政策。
- FluidInference:Core ML 转换由 FluidInference 完成, 包括防 fp16 溢出的 RMSNorm、各个函数、拼接功能和图像流程。
- 本仓库:我们做了 int8 模型、int8 词表、编译和测试。
- 无关联声明:本仓库与 Google、FluidInference 没有关联。
