Team Ai
Modelpublic

OpenInfinity/embeddinggemma-2-coreml-int8

sourceHugging Faceapache-2.0updated 13h agoView on Hugging Face
0likes2downloads
Model Card

EmbeddingGemma 2 — Core ML int8 (text + image, compiled)

English · 中文

<a id="english"></a>

English

The text and image encoders of google/embeddinggemma-2, based on FluidInference/embeddinggemma-2-coreml. We changed three things:

  • —int8 weights. Both models have int8 weights, which halves their size. Activations stay fp16.
  • —int8 token table. The token embedding table is int8 with one fp16 scale per row, which also halves it.
  • —Precompiled. Both models are compiled to .mlmodelc, so an app can load them without calling MLModel.compileModel.

On our tests the retrieval quality is the same as the fp32 reference (see Validation). The text model runs entirely on the Apple Neural Engine. The audio encoder is not included.

The output is the same 768-d, L2-normalized embedding as SentenceTransformer("google/embeddinggemma-2").encode(...). Matryoshka truncation to 512, 256 or 128 dimensions works as in the original: keep the first N values and re-normalize.

Files

FileWhatSize
text/EmbeddingGemma2Text.mlmodelcText model, int8 weights. 7 functions share one set of weights.147 MB
text/embeddings.int8Token table, 262,144 × 512 int8, raw, row-major134 MB
text/embeddings.int8.scale.f16One fp16 scale per table row (262,144 values)0.5 MB
text/tokenizer.jsonThe original Gemma tokenizer, unchanged32 MB
vision/EmbeddingGemma2Vision.mlmodelcImage encoder, int8 weights. Functions: vision_70, vision_140, vision_280.155 MB
vision/position_embeddings.f16The patch embedder's x and y position tables, fp163 MB
config.jsonShapes, functions, token ids, task prefixes, image patch layout
SHA256SUMSChecksums of every file

Text-only use needs config.json and text/ (about 313 MB). Image search also needs vision/ (about 158 MB). Sizes on this page are in MB = 10⁶ bytes (FluidInference's README uses MiB).

What changed compared to FluidInference's package

FluidInferenceThis repo
Model weightsfp16int8: linear symmetric, per output channel, for every weight with more than 2048 elements
Token tableembeddings.bf16, 268 MB`embeddings.int8` + per-row fp16 scale, 134 MB: scale = max(abs(row)) / 127
Format.mlpackage`.mlmodelc`, compiled with coremltools 9
Text / image size284 / 306 MB147 / 155 MB
Audio encoderyes, on the GPU, fp16not included

Functions, inputs and outputs are unchanged, so FluidInference's README and its Swift code (FluidUse, EmbeddingGemma2Manager) describe how to call these models. The only difference in calling them is the table lookup (step 3 below). The file layout, config.json and the table format do differ, so FluidUse cannot load this repo without changes.

How the int8 models were made. The coremltools int8 API (linear_quantize_weights) rejects multifunction packages, so we:

  1. 1.split each package into one single-function model per function (via the MIL program);
  2. 2.quantized each model to int8;
  3. 3.merged them back with ct.utils.save_multifunction, which stores identical weights only once;
  4. 4.compiled the result with ct.utils.compile_model.

The compiled models give the same output as the .mlpackage they came from: cosine 1.000000, max difference 0.

Use (text)

  1. 1.Prefix. Add the task prefix, e.g. title: none | text: for documents and task: search result | query: for queries. All prefixes are in config.json.
  2. 2.Tokenize. Use text/tokenizer.json. The sequence must start with <bos> (2) and end with <eos> (1); the tokenizer adds both. If it is longer than 512 tokens, keep the first 511 and end with <eos>.
  3. 3.Look up rows. For each token id, read its row from embeddings.int8 and multiply by scale[id] × √512 (√512 ≈ 22.627417). Pass the result as fp16.
  4. 4.Run. Call the smallest embed_S that fits, S ∈ {32, 48, 64, 128, 256, 512}. Inputs: inputs_embeds [1, S, 512] and attention_mask [1, S], 1 for tokens, right padding. Output: embedding [1, 768]. To embed up to eight short texts in one call, use pack_256; see FluidInference's README.
swift
import CoreML

// Load one function of the compiled model (macOS 15 / iOS 18 or later).
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
config.functionName = "embed_128"
let model = try MLModel(contentsOf: textURL, configuration: config)   // text/EmbeddingGemma2Text.mlmodelc

// Row lookup for one token: int8 row × per-row scale × √512.
func row(_ id: Int, table: UnsafePointer<Int8>, scales: UnsafePointer<Float16>, into out: UnsafeMutablePointer<Float16>) {
    let s = Float(scales[id]) * 22.627417
    for j in 0..<512 { out[j] = Float16(Float(table[id * 512 + j]) * s) }
}
python
import numpy as np, coremltools as ct
from tokenizers import Tokenizer

tok = Tokenizer.from_file("text/tokenizer.json")
table = np.memmap("text/embeddings.int8", dtype=np.int8, mode="r").reshape(262144, 512)
scale = np.fromfile("text/embeddings.int8.scale.f16", dtype=np.float16).astype(np.float32)

ids = tok.encode("task: search result | query: 周末去哪里爬山").ids   # the tokenizer adds <bos> (2) and <eos> (1)
x = np.zeros((1, 64, 512), np.float16); m = np.zeros((1, 64), np.float16)
x[0, :len(ids)] = table[ids] * scale[ids, None] * 22.627417; m[0, :len(ids)] = 1
model = ct.models.CompiledMLModel("text/EmbeddingGemma2Text.mlmodelc",
                                  compute_units=ct.ComputeUnit.CPU_AND_NE, function_name="embed_64")
v = model.predict({"inputs_embeds": x, "attention_mask": m})["embedding"][0]
v /= np.linalg.norm(v)

Use (images)

The steps are the same as in FluidInference's README:

  1. 1.Resize the image, keeping its aspect ratio.
  2. 2.Cut it into 16 px patches.
  3. 3.Run vision_70, vision_140 or vision_280.
  4. 4.Put the resulting tokens into <bos> <|image> … <image|> <eos> and embed that sequence with the text model.

The table rows for the special tokens come from embeddings.int8, as in step 3 above.

Validation

Tests ran on one MacBook Pro (M3 Max, macOS 26.7). Quality was measured with the Neural Engine (cpuAndNeuralEngine) for both text and images. "int8" means int8 models and the int8 token table, which is what this repo ships.

Similarity to the fp32 reference (sentence-transformers, PyTorch fp32):

Cosine, min / mean
64 texts from the simulated knowledge base (48 chunks, 16 queries)0.9995 / 0.9998
100 screenshots, 140 image tokens0.998 / 0.9994

Retrieval quality:

Test setfp32 referenceFluidInference fp16**This repo (int8)**
Simulated personal knowledge base, nDCG@1093.593.293.2
MLQA, Chinese question → English passage, nDCG@1067.868.267.8
Screenshots, image embedding only, MRR@1092.992.992.9
Screenshots, image + OCR-text score fusion, MRR@1096.796.796.8

About the test sets:

  • —Simulated personal knowledge base. 2,829 chunks of notes, Wikipedia pages, open-source documentation and code, mixed Chinese and English, with 113 queries written by an LLM.
  • —MLQA. 200 questions.
  • —Screenshots. 100 screenshots: 60 synthetic app screens and 40 Wikipedia pages, half Chinese and half English, with 100 queries. OCR is Apple Vision.
  • —Score fusion. z(image score) + 0.5 · z(OCR-text score), where z standardizes scores per query.

Speed (M3 Max, one call, warm):

Neural EngineGPU
embed_32 / 48 / 64 / 128 / 256 / 5123.0 / 3.5 / 3.9 / 6.3 / 13.3 / 32.0 msnot measured
vision_70 / vision_140 / vision_280, one image43 ms / 123 ms / not tried35 / 70 / 154 ms
2,740 knowledge-base chunks (82 tokens on average), one per call / with pack_256165 / 189 texts/snot measured

An image needs one embed_S call after the image model. The GPU times are short runs at full GPU clock; under sustained load the GPU slowed down (see Energy).

Neural Engine placement (MLComputePlan): every op runs on the Neural Engine in all seven text functions (3,862 ops each, 3,863 in pack_256) and in vision_70 and vision_140 (2,389 ops each). We did not check vision_280 on the Neural Engine.

int8 has about the same speed as fp16. The gain is size: text model and table 552 → 281 MB, image model 306 → 155 MB.

Energy. Same machine, on AC power in High Power mode. The 100 screenshots ran in a loop for 30 s, each setup twice, and idle power was subtracted. The text model ran on the Neural Engine in every case.

Image model on the Neural EngineImage model on the GPU
Time per screenshot (image model + text model)136 ms in both runs84 ms, then 92 ms
Power above idle4.1 W46 W, then 37 W
Energy per screenshot0.55 J3.4–3.9 J
Image model alone: first run → second run123 ms → 123 ms; 0.50 → 0.51 J70 ms → 146 ms; 3.9 → 1.8 J
  • —The Neural Engine uses about 7× less energy per screenshot. Indexing 10,000 screenshots costs about 1.5 Wh on the Neural Engine and about 10 Wh on the GPU; the battery of this 16-inch MacBook Pro holds 100 Wh.
  • —The GPU is faster only in short runs. It drew about 54 W at first. After about 1.5 minutes of GPU load, the system lowered the GPU clock from about 1.3 GHz to 0.6–0.8 GHz. The image model then took 146 ms per screenshot, slower than the Neural Engine, which stayed at 123 ms.
  • —Recommendation: run vision_70 and vision_140 on the Neural Engine (cpuAndNeuralEngine), especially for background indexing. Use the GPU for vision_280, which we did not try on the Neural Engine. On the GPU the int8 image tokens match FluidInference's fp16 package (cosine ≥ 0.999 for every image function on test inputs).

Limitations

  • —Tested on one machine only: M3 Max with macOS 26.7. The models need macOS 15 / iOS 18 or later, because they are multifunction models.
  • —First load is slow. The first load on a device compiles each function for the Neural Engine. FluidInference reports several minutes for the seven text functions. Core ML caches the result. Apps should warm the functions up in the background, because a .mlmodelc cannot be compiled for the Neural Engine ahead of time.
  • —Fixed shapes. Inputs longer than 512 tokens must be truncated.
  • —`vision_280` was not tried on the Neural Engine. Run it on the GPU.
  • —Small test sets. The test sets are small (100–200 queries each) and some queries are LLM-written. Treat the numbers as a check that int8 loses nothing, not as a benchmark.
  • —No audio. Audio is not included. In our tests EmbeddingGemma 2's audio embeddings were weak for Chinese speech and for cross-language search, and speech-to-text followed by text embedding worked better. If you need the audio encoder, use FluidInference's fp16 package.

License and attribution

  • —License. Apache 2.0, the same as the source models.
  • —Google. EmbeddingGemma 2 is by Google DeepMind. Deployments must follow the Gemma Prohibited Use Policy.
  • —FluidInference. The Core ML conversion (fp16-safe RMSNorm, functions, packing, image pipeline) is FluidInference's.
  • —This repo. We added the int8 models, the int8 token table, the compilation and the tests.
  • —No affiliation. This repo is not affiliated with Google or FluidInference.

<a id="中文"></a>

中文

本仓库提供 google/embeddinggemma-2 的文本编码器和图像编码器, 基于 FluidInference/embeddinggemma-2-coreml 的 Core ML 版本。 我们做了三处改动:

  • —模型权重 int8:两个模型的权重都压成了 int8,体积减半;计算仍用 fp16。
  • —词表 int8:词表也压成了 int8,每行配一个 fp16 缩放系数,体积同样减半。
  • —预先编译:两个模型都编译成了 .mlmodelc,App 可以直接加载,不用再调用 MLModel.compileModel。

在我们的测试中,检索效果和 fp32 原版一样(见下文“验证”)。文本模型完全在 Apple 神经网络引擎(ANE)上运行。 本仓库不含音频编码器。

输出和 SentenceTransformer("google/embeddinggemma-2").encode(...) 相同:768 维、已归一化的向量。 如果想用 512、256 或 128 维,和原版一样取前 N 维,再重新归一化即可。

文件

文件内容大小
text/EmbeddingGemma2Text.mlmodelc文本模型,int8 权重;7 个函数共用一份权重147 MB
text/embeddings.int8词表,262,144 × 512 的 int8 原始数据,按行存放134 MB
text/embeddings.int8.scale.f16词表每行一个 fp16 缩放系数(共 262,144 个)0.5 MB
text/tokenizer.jsonGemma 原版分词器,未改动32 MB
vision/EmbeddingGemma2Vision.mlmodelc图像编码器,int8 权重;函数 vision_70、vision_140、vision_280155 MB
vision/position_embeddings.f16图块的 x、y 位置表,fp163 MB
config.json形状、函数列表、特殊 token、任务前缀、图块排列方式
SHA256SUMS所有文件的校验和

只用文本:需要 config.json 和 text/,约 313 MB。还要搜图片:再加 vision/,约 158 MB。 本页的 MB 指 10⁶ 字节(FluidInference 的 README 用的是 MiB)。

和 FluidInference 原版的区别

FluidInference本仓库
模型权重fp16int8:线性对称,按输出通道量化;只压缩元素数超过 2048 的权重
词表embeddings.bf16,268 MB`embeddings.int8` + 每行一个 fp16 缩放系数,134 MB;scale = 该行绝对值的最大值 / 127
格式.mlpackage`.mlmodelc`,用 coremltools 9 编译
文本 / 图像模型体积284 / 306 MB147 / 155 MB
音频编码器有,在 GPU 上跑,fp16不含

函数名、输入和输出都没变,调用方法可以直接参考 FluidInference 的 README 和它的 Swift 代码 (FluidUse 里的 EmbeddingGemma2Manager)。调用上唯一的区别是查词表(见下面第 3 步)。但文件布局、config.json 和词表格式都不同,FluidUse 不改代码无法直接加载本仓库。

int8 模型是怎么做的:coremltools 的 int8 压缩接口(linear_quantize_weights)不支持多函数模型,所以我们分四步:

  1. 1.通过 MIL 程序把每个函数拆成单独的模型;
  2. 2.逐个压成 int8;
  3. 3.用 ct.utils.save_multifunction 合并回去,相同的权重只存一份;
  4. 4.用 ct.utils.compile_model 编译。

编译后的模型和编译前的 .mlpackage 输出完全一致:余弦相似度 1.000000,最大差值 0。

用法(文本)

  1. 1.加前缀:加上任务前缀,例如文档用 title: none | text: ,查询用 task: search result | query: 。所有前缀都在 config.json 里。
  2. 2.分词:用 text/tokenizer.json 分词。序列必须以 <bos>(2)开头、以 <eos>(1)结尾,分词器会自动加上这两个。 超过 512 个 token 时,保留前 511 个,最后接 <eos>。
  3. 3.查词表:对每个 token id,从 embeddings.int8 取出对应的行,乘以 scale[id] × √512(√512 ≈ 22.627417),转成 fp16。
  4. 4.运行:调用能放下这段文本的最小 embed_S,S ∈ {32, 48, 64, 128, 256, 512}。
  5. 5.输入是 inputs_embeds [1, S, 512] 和 attention_mask [1, S](有 token 的位置为 1,补齐放在右侧);
  6. 6.输出是 embedding [1, 768];
  7. 7.想一次算最多 8 段短文本,用 pack_256,见 FluidInference 的 README。
swift
import CoreML

// 加载编译好的模型中的一个函数(需要 macOS 15 / iOS 18 或更高版本)
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
config.functionName = "embed_128"
let model = try MLModel(contentsOf: textURL, configuration: config)   // text/EmbeddingGemma2Text.mlmodelc

// 查一个 token 的行:int8 行 × 该行的缩放系数 × √512
func row(_ id: Int, table: UnsafePointer<Int8>, scales: UnsafePointer<Float16>, into out: UnsafeMutablePointer<Float16>) {
    let s = Float(scales[id]) * 22.627417
    for j in 0..<512 { out[j] = Float16(Float(table[id * 512 + j]) * s) }
}
python
import numpy as np, coremltools as ct
from tokenizers import Tokenizer

tok = Tokenizer.from_file("text/tokenizer.json")
table = np.memmap("text/embeddings.int8", dtype=np.int8, mode="r").reshape(262144, 512)
scale = np.fromfile("text/embeddings.int8.scale.f16", dtype=np.float16).astype(np.float32)

ids = tok.encode("task: search result | query: 周末去哪里爬山").ids   # 分词器会自动加上 <bos>(2)和 <eos>(1)
x = np.zeros((1, 64, 512), np.float16); m = np.zeros((1, 64), np.float16)
x[0, :len(ids)] = table[ids] * scale[ids, None] * 22.627417; m[0, :len(ids)] = 1
model = ct.models.CompiledMLModel("text/EmbeddingGemma2Text.mlmodelc",
                                  compute_units=ct.ComputeUnit.CPU_AND_NE, function_name="embed_64")
v = model.predict({"inputs_embeds": x, "attention_mask": m})["embedding"][0]
v /= np.linalg.norm(v)

用法(图片)

和 FluidInference 的 README 相同:

  1. 1.按原比例缩放图片;
  2. 2.切成 16 px 的图块;
  3. 3.运行 vision_70、vision_140 或 vision_280;
  4. 4.把得到的 token 放进 <bos> <|image> … <image|> <eos>,再用文本模型算出向量。

特殊 token 的行同样从 embeddings.int8 里查,方法同上面第 3 步。

验证

在一台 MacBook Pro(M3 Max,macOS 26.7)上测试。效果测试时,文本和图像都跑在神经网络引擎上(cpuAndNeuralEngine)。 下面的“int8”指 int8 模型加 int8 词表,也就是本仓库发布的版本。

和 fp32 原版的相似度(参照:sentence-transformers,PyTorch fp32):

余弦相似度,最低 / 平均
模拟个人知识库中的 64 段文本(48 个文本块、16 条查询)0.9995 / 0.9998
100 张截图,140 个图像 token0.998 / 0.9994

检索效果:

测试集fp32 原版FluidInference fp16**本仓库(int8)**
模拟个人知识库,nDCG@1093.593.293.2
MLQA,中文问题找英文段落,nDCG@1067.868.267.8
截图,只用图像向量,MRR@1092.992.992.9
截图,图像和 OCR 文字按分数融合,MRR@1096.796.796.8

测试集说明:

  • —模拟个人知识库:2,829 个文本块,内容包括笔记、维基百科页面、开源文档和代码,中英混合;113 条查询由大模型编写。
  • —MLQA:200 个问题。
  • —截图:100 张,其中 60 张合成的软件界面、40 张维基百科页面,中英各半;查询 100 条;文字识别用 Apple Vision。
  • —分数融合:z(图像分数) + 0.5 · z(OCR 文字分数),z 表示在每条查询内把分数标准化。

速度(M3 Max,单次调用,已预热):

神经网络引擎GPU
embed_32 / 48 / 64 / 128 / 256 / 5123.0 / 3.5 / 3.9 / 6.3 / 13.3 / 32.0 ms未测
vision_70 / vision_140 / vision_280,每张图43 ms / 123 ms / 未试35 / 70 / 154 ms
2,740 个知识库文本块(平均 82 个 token):逐段 / 用 pack_256 拼接每秒 165 / 189 段未测

每张图在图像模型之后,还要调用一次 embed_S。GPU 的耗时是 GPU 满频时短时间跑出来的;持续运行时 GPU 会变慢(见下面的“能耗”)。

算子分配(MLComputePlan):7 个文本函数(每个 3,862 个算子,pack_256 是 3,863 个)以及 vision_70、vision_140 (每个 2,389 个算子)的所有算子都在神经网络引擎上。vision_280 没有在神经网络引擎上检查。

int8 的速度和 fp16 基本一样,好处在体积:文本模型加词表从 552 MB 降到 281 MB,图像模型从 306 MB 降到 155 MB。

能耗:同一台机器,接电源,高性能模式。100 张截图循环跑 30 秒,每种配置跑两轮,扣除空闲功耗。文本模型都在神经网络引擎上。

图像模型在神经网络引擎上图像模型在 GPU 上
每张截图耗时(图像模型 + 文本模型)两轮都是 136 ms84 ms,第二轮 92 ms
比空闲多出的功率4.1 W46 W,第二轮 37 W
每张截图的能耗0.55 J3.4–3.9 J
只跑图像模型:第一轮 → 第二轮123 ms → 123 ms;0.50 → 0.51 J70 ms → 146 ms;3.9 → 1.8 J
  • —神经网络引擎每张截图的能耗约为 GPU 的 1/7:给 1 万张截图建索引,在神经网络引擎上约耗 1.5 Wh,在 GPU 上约 10 Wh; 这台 16 英寸 MacBook Pro 的电池是 100 Wh。
  • —GPU 只在短时间内更快:一开始 GPU 功率约 54 W。GPU 持续工作约 1.5 分钟后,系统把 GPU 频率从约 1.3 GHz 降到 0.6–0.8 GHz, 图像模型每张要 146 ms,比神经网络引擎还慢;神经网络引擎一直稳定在 123 ms。
  • —建议:vision_70 和 vision_140 在神经网络引擎上运行(cpuAndNeuralEngine),后台批量建索引时尤其如此。 vision_280 没有在神经网络引擎上试过,请用 GPU。在 GPU 上,int8 图像模型输出的 token 和 FluidInference 的 fp16 版本一致 (用测试输入检查,每个图像函数的余弦相似度都 ≥ 0.999)。

局限

  • —只在一台机器上测过:M3 Max,macOS 26.7。模型是多函数模型,需要 macOS 15 / iOS 18 或更高版本。
  • —第一次加载很慢:在每台设备上第一次加载时,系统要把每个函数编译给神经网络引擎;FluidInference 报告 7 个文本函数共需几分钟。 Core ML 会缓存编译结果。.mlmodelc 没法提前编译到神经网络引擎这一步,所以建议 App 在后台预热各个函数。
  • —形状固定:超过 512 个 token 的输入必须截断。
  • —`vision_280` 没有在神经网络引擎上试过:请在 GPU 上运行。
  • —测试集很小:每组 100–200 条查询,部分查询由大模型编写。这些数字只用来确认 int8 没有损失效果,不能当作基准成绩。
  • —不含音频:在我们的测试中,EmbeddingGemma 2 的音频向量对中文语音和跨语言搜索效果较弱,先转文字再算文本向量效果更好。 如果需要音频编码器,请使用 FluidInference 的 fp16 版本。

许可证和来源

  • —许可证:Apache 2.0,和原模型相同。
  • —Google:EmbeddingGemma 2 由 Google DeepMind 发布。使用时须遵守 Gemma 禁止用途政策。
  • —FluidInference:Core ML 转换由 FluidInference 完成, 包括防 fp16 溢出的 RMSNorm、各个函数、拼接功能和图像流程。
  • —本仓库:我们做了 int8 模型、int8 词表、编译和测试。
  • —无关联声明:本仓库与 Google、FluidInference 没有关联。