Team Ai
Modelpublic

shirochenkov90/embeddinggemma-2-coreml

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes36downloads
Model Card

EmbeddingGemma 2 (text) — Core ML

Core ML conversion of the text part of google/embeddinggemma-2 for native macOS / iOS apps (Swift + Core ML). Weights are unchanged; only the format is. The vision and audio encoders of the original multimodal model are not included — this package embeds text only.

The package contains the complete sentence-transformers text pipeline, so the output is the final embedding:

EmbeddingGemma2 text model (incl. its internal 512→768 projection) → mean pooling over attention_mask → L2 normalization

Vectors from this model are not compatible with EmbeddingGemma 1 (google/embeddinggemma-300m): switching models requires re-embedding all stored documents. Similarity values are also distributed differently (generally higher), so any similarity thresholds need to be re-tuned.

Files

FileWhat it is
embeddinggemma-2.mlpackage/Core ML model (ML Program, float16)
tokenizer.json, tokenizer_config.jsonTokenizer files from the source model, unchanged
config.jsonModel config from the source model, unchanged (it also describes the vision/audio parts, which are not in this package)
LICENSEApache License 2.0

Core ML interface

NameDirectiondtypeShape
input_idsinputint32(1, 64) or (1, 256)
attention_maskinputint32(1, 64) or (1, 256)
embeddingoutputfloat32(1, 768)
  • —Enumerated shapes: pass exactly (1, 64) or (1, 256); both inputs must have the same shape in one call. Use 64 for short queries, 256 for passages. Inputs longer than 256 tokens must be truncated (the original model supports up to 8192 tokens; this package is limited to 256).
  • —Padding goes on the RIGHT: real tokens first, then pad id 0 with attention_mask = 0. Positions are computed as 0…S-1 inside the model, so left padding would change the result.
  • —Batch size is 1.
  • —The tokenizer prepends `<bos>` (id 2) and appends `<eos>` (id 1). Reproduce exactly that when tokenizing in Swift. Example: проверка → [2, 7877, 144813, 1].
  • —Output is already L2-normalized: cosine similarity = dot product.
  • —Embedding dimension: 768. Matryoshka truncation to 512/256/128 works as in the original model: take the first N values and re-normalize.
  • —Minimum deployment target: macOS 15 / iOS 18 (required for enumerated shapes on two inputs).

Performance note: use .cpuOnly

On Apple Silicon (macOS 15+) Core ML places all operations of this package on the CPU, even with computeUnits = .all (checked with MLComputePlan: 2942 of 2942 ops on CPU). With .all the extra dispatch overhead makes it slower than plain CPU:

computeUnitsTime per 256-token input (Mac, Apple Silicon)
.cpuOnly~216 ms
.all~524 ms

So configure the model with .cpuOnly. On the CPU Core ML computes in float32, which is also why the outputs match PyTorch exactly (cos = 1.000000) despite the float16 weights.

Task prefixes (prompts)

Prepend the prefix to the raw text, exactly as in the original sentence-transformers config (note the trailing space):

Prompt namePrefix
query, SearchQuery, Retrieval-query, Retrieval, Reranking, BitextMining`task: search result \query: `
document, Document, Retrieval-document`title: none \text: `
QuestionAnswering`task: question answering \query: `
FactChecking`task: fact checking \query: `
CodeRetrieval, InstructionRetrieval`task: code retrieval \query: `
Classification, MultilabelClassification`task: classification \query: `
Clustering`task: clustering \query: `
SentenceSimilarity, STS, PairClassification, Summarization`task: sentence similarity \query: `

For search:

text
query    = "task: search result | query: " + text
document = "title: none | text: " + text        // or "title: <title> | text: " + text

Swift usage sketch

swift
import CoreML

let config = MLModelConfiguration()
config.computeUnits = .cpuOnly   // faster than .all for this package, see "Performance note"
let model = try embeddinggemma_2(configuration: config)   // class generated by Xcode from the .mlpackage

/// tokenIds must already contain <bos> (2) at the start and <eos> (1) at the end.
func embed(tokenIds: [Int32]) throws -> [Float] {
    let length = tokenIds.count <= 64 ? 64 : 256
    precondition(tokenIds.count <= length)
    let ids = try MLMultiArray(shape: [1, NSNumber(value: length)], dataType: .int32)
    let mask = try MLMultiArray(shape: [1, NSNumber(value: length)], dataType: .int32)
    for i in 0..<length {
        ids[i] = NSNumber(value: i < tokenIds.count ? tokenIds[i] : 0)   // pad id 0
        mask[i] = NSNumber(value: i < tokenIds.count ? 1 : 0)
    }
    let out = try model.prediction(input_ids: ids, attention_mask: mask)
    let e = out.embedding
    return (0..<e.count).map { Float(truncating: e[$0]) }
}

Verification (this exact package, Apple Silicon, compute units ALL — all ops placed on CPU)

Cosine between the PyTorch pipeline (SentenceTransformer.encode, float32) and Core ML:

TextTokensShapecos(PyTorch, Core ML)
Russian passage (document prefix)153(1, 256)1.000000
Russian short query (query prefix)13(1, 64)1.000000

Semantic check (Core ML): query task: search result | query: про деньги vs document A title: none | text: обсудили бюджет на следующий квартал → cos 0.7320; vs document B title: none | text: починили баг в плеере → cos 0.6348 (PyTorch: 0.7321 / 0.6348).

Cross-script check: скан vs Scan → cos 0.9366 in Core ML (0.9366 in PyTorch).

How it was converted

torch.jit.trace of a single nn.Module wrapping the text pipeline (text model with an explicit bidirectional 4-D attention mask and position ids, mean pooling, L2 norm), then coremltools.convert(..., convert_to="mlprogram", compute_precision=FLOAT16, minimum_deployment_target=macOS15, inputs with EnumeratedShapes [(1, 64), (1, 256)]). Before conversion the wrapper was checked against SentenceTransformer.encode (cos = 1.0000000).

Two conversion details:

  • —coremltools 9.0 has no converter for the boolean a | b used inside the model; it was mapped to MIL logical_or.
  • —Mean pooling divides by the token count before summing, and the pooled vector is rescaled to [-1, 1] before L2 normalization. Both are mathematically identical to the original pipeline and only prevent float16 overflow (activations reach ~2200).

Versions: torch 2.7.0, transformers 5.19.0, sentence-transformers 6.1.0, coremltools 9.0, numpy 2.3.5.

License

Apache License 2.0 — see LICENSE, same as the source model. This repository is a format conversion of Google's model with unchanged weights; the changes are limited to the export described above.