AtomicChat/embeddinggemma-2-GGUF
How to Run EmbeddingGemma 2 Locally
<p style="margin-top: 0; margin-bottom: 0;"> <em>Built from Google's original weights with our own importance matrix. The <a href="https://huggingface.co/datasets/AtomicChat/calib-corpora">calibration corpora</a> behind our builds are public.</em> </p> <div style="display: flex; gap: 8px; align-items: center; margin-top: 10px; margin-bottom: 10px;"> <a href="https://atomic.chat/?utmsource=huggingface&utmmedium=referral&utmcampaign=hfembeddinggemma2&utmcontent=btnatomic"><img src="https://huggingface.co/AtomicChat/Qwen3.8-27B-GGUF/resolve/main/btnatomic.png" width="162" alt="Atomic Chat"></a> <a href="https://discord.gg/8wGSsvmg4V"><img src="https://huggingface.co/AtomicChat/Qwen3.8-27B-GGUF/resolve/main/btndiscord.png" width="119" alt="Discord"></a> <a href="https://github.com/AtomicBot-ai/Atomic-Chat"><img src="https://huggingface.co/AtomicChat/Qwen3.8-27B-GGUF/resolve/main/btngithub.png" width="115" alt="GitHub"></a> </div> <ul style="margin: 0 0 12px 0;"> <li>EmbeddingGemma 2 puts text, code, images and audio into one vector space, and its text model has 270M parameters.</li> <li>See <a href="#how-we-compressed-it-better">how we compressed it better</a> below for the measurements and the reasoning.</li> </ul> <hr style="margin: 0 0 16px 0;">
Lower is better, and Unsloth's file of the same size is 100.
- At about 245 MB our file drifts 32% less than Unsloth's and 17% less than AutoRound's (40% and 21% less on code).
- At about 175 MB it drifts 11% less than both on 30 languages. On code, AutoRound's file is 4% closer.
- Which document a search finds first stays within noise for all three builds; see head to head.
- The
Q8_0file from AtomicChat, Unsloth and ggml-org is the same file.
Pick a file
Every number below is measured, not estimated. How we measured it is at the bottom, and the raw logs are in the metrics repo so you can check any of it yourself.
An embedding model has no next word to predict, so the usual KL divergence does not apply. What matters is whether search still finds what the original finds. same top result is how often a search with this file returns the same best passage as the original weights, over 1,000 queries in 30 languages. cosine to BF16 is how close each vector is to the original's vector for the same text (1 means identical).
AD- means Atomic Dynamic: the type is chosen per tensor from measurements rather than taken from a llama.cpp preset.
Which one to take:
- `Q8_0` if you can spare 310 MB. It is practically the original.
- `AD-Q6_K` for a smaller file whose vectors stay very close to the original's.
- `AD-Q4_K_M` when size matters most. Its top result differs from the original's on about one query in eight, but it finds the right passage as often as the original does (see Search quality below).
[!IMPORTANT] Index your documents and embed your queries with the same file. Vectors from different quantizations are close, not identical, so do not mix them in one index.
[!NOTE] We do not ship F16. Google reports that the model's activations overflow float16 and produce NaN or silently degraded vectors. Use BF16 or one of the quantized files.
Images and audio
Images and audio go through the vision and audio encoders in a separate projector file (--mmproj). The model also embeds video; we have not tested that path.
This is a sanity check, not a benchmark: one image and one 17-second speech clip. In both projector files the image sits closer to its right caption than to a wrong one (0.79 against 0.56), and so does the audio clip (0.59 against 0.52).
Running it
[!NOTE] These files need llama.cpp with thegemma-embedding2architecture, which arrived in llama.cpp PR #30054 on 6 October 2026. An older build stops withunknown model architecture: 'gemma-embedding2'. Apps built on llama.cpp, such as LM Studio and Jan, run the files once their bundled engine includes that change.
llama-server -m embeddinggemma-2-Q8_0.gguf --embeddings --pooling mean \
-c 2048 -b 2048 -ub 2048 -ngl 99The model reads the whole input at once (bidirectional attention), so -ub has to be at least as long as your longest input. Raise -c, -b and -ub together for inputs up to the model's 8K context.
curl http://127.0.0.1:8080/v1/embeddings -H "Content-Type: application/json" -d '{
"input": ["task: search result | query: how do I reset my password?",
"title: none | text: Open Settings, choose Forgot password and follow the link we email you."],
"encoding_format": "float"}'The vectors come back L2-normalised, with 768 dimensions.
Task prefixes. The model was trained with them, and leaving them out costs accuracy:
Shorter vectors. Keep the first 512, 256 or 128 values and normalise again. At 256 dimensions every file keeps the same standing as in the table above (all numbers are in results.json).
Images and audio. Start the server with the projector, then send an OpenAI-style content array:
llama-server -m embeddinggemma-2-Q8_0.gguf --mmproj mmproj-embeddinggemma-2-Q8_0.gguf \
--embeddings --pooling mean -c 4096 -b 4096 -ub 4096 -ngl 99{"input": [{"content": [{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}]},
{"content": [{"type": "input_audio", "input_audio": {"data": "<base64 wav>", "format": "wav"}}]},
"task: search result | query: a newspaper front page about the moon landing"],
"encoding_format": "float"}How these compare to other builds
We downloaded the other publishers' files and measured them with the same harness, the same eval set and the same machine. 1 - cosine is the mean distance of a file's vectors from the original's: lower is better.
- At about 245 MB our vectors are the closest on both sets: 17% closer than AutoRound's and 32% closer than Unsloth's on the multilingual set, and 21% and 40% closer on code. The
same top resultrates differ by less than the noise. Ours is 96.3% (95% interval 95.0-97.4%) and AutoRound's is 97.5% (96.4-98.4%). - At about 175 MB ours is the closest on the multilingual set, 11% ahead of the next file, and the smallest. On code, AutoRound's file is 4% closer than ours and 2.7 MB larger.
- 5-bit: we did not ship one. Ours only tied Unsloth's.
- `Q8_0` from AtomicChat, Unsloth and ggml-org is the same file, tensor for tensor.
- AutoRound is webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF, a GGUF export made with Intel's AutoRound.
Head to head on the same texts
Averages can hide a few outliers, so we also compared each pair text by text. The columns are:
- Our drift vs theirs: our mean 1 - cosine relative to theirs, with a 95% interval from resampling whole documents.
- Texts where ours is closer: the share of single texts where our vector is closer to the original's.
- Same top result: the queries where only one of the two files picks the original's best passage, with an exact sign test.
How to read it:
- The drift differences are real. None of the intervals crosses zero, and at 245 MB our vector is the closer one on 82-99% of texts.
- Search results are a draw. Only one of the two files gets the top passage right on 9 to 161 queries, depending on the pair, out of 1,000 (30 languages) or 400 (code). No pair differs beyond chance.
- The 245 MB file is our clearest win.
- At 175 MB the files are close. AutoRound is ahead on code, probably thanks to its calibrated rounding. Its file also quantizes
per_layer_model_proj, which ours keeps in BF16.
How we compressed it better
Unsloth's and AutoRound's files use one type for almost the whole model: Q6K everywhere, or Q4K everywhere. Ours spend the bits by measurement instead.
The model
The text model is a small Gemma 4-style encoder. A text is cut into tokens, and each token takes its 512-number row from a 262,144-row token table. The vectors then pass through 24 transformer blocks in which every token sees every other one. Five blocks with a local window alternate with one global block. At the end the model takes the mean over all tokens, and an output head turns that mean into the final 768-number vector.
The token table is half of the model: 134M of 270M parameters. The 24 blocks make up another 130M.
We measured before we decided
We quantized one part of the model at a time to Q4K, kept everything else at Q80, and measured how far the vectors moved. The table shows the cost of each megabyte saved, with the token table as 1:
Two parts stand at opposite ends.
The token table is the cheapest place to save bytes. Each token takes its own row, so rounding noise in the table is different for every token. The model then averages all tokens into one vector, and independent noise mostly cancels out. A rounding error inside a transformer weight is different: every token passes through the same matrix, so the error pushes all of them the same way, and pooling cannot remove it.
The output head is the most fragile part. It works on the already-averaged vector, once, so its rounding error shifts every final vector in the same direction and nothing averages it out. Putting the head alone in Q4K saves 0.2 MB and costs more than putting all of attention's q, k and v in Q4K.
Depth matters too
The same test on four bands of six blocks each:
The first and last blocks are 2-3 times more sensitive than the middle.
Costs add up, so a layout can be planned
The four bands measured one at a time add up to what the whole stack costs at once (0.0068 against 0.0067). So within the target size we could add up measured costs, give bits to the parts where they buy the most, and then measure the few best candidates whole.
The result is the picture above:
- the token table takes fewer bits than in the other builds;
- the edge blocks and the output head take more;
- the middle blocks take the fewest bits in the transformer.
About 30 test files went into it, each about 30 seconds to build and measure on a laptop.
What we could not do
- `per_layer_model_proj` (12.6 MB). Unsloth's 4- and 5-bit files and all of AutoRound's quantize it. Stock
llama-quantizealways keeps it in BF16, and so do our files. At equal size we therefore spend about 6-8 MB less on everything else. - Same top result. The better vectors do not show up as a better
same top resultrate: the differences there stay inside the noise.
Notes for anyone building their own
- The imatrix needs three adjustments on this model.
- Without
--override-kv tokenizer.ggml.add_eos_token=bool:false,llama-imatrixstops on an assertion. - It needs
--no-ppl, because there are no logits. - It needs
-ubequal to-c, because attention is bidirectional. - The head, the token table and
per_layer_model_projget no importance data at all. - Check the base. Our BF16 and projector files are tensor-for-tensor identical to ggml-org's conversions of the same checkpoint.
Search quality
Quantization reorders near-ties among the candidates. It does not make search worse. How often a query finds its own passage among all passages of the set (recall@1):
The task is hard on purpose: the query is the first quarter of a chunk and the passage is the rest, so the absolute numbers are low. The point is that they do not move.
The calibration data
About 1.24M tokens from the calibration corpora pool: 60% Wikipedia in 30 languages and 40% source code. Everything is written the way the model sees it at work:
- passages as
title: none | text: ..., ortitle: <file name> | text: ...for code; - queries taken from those passages under all seven task prefixes.
None of it overlaps the held-out eval set. The builder script and the corpus itself are in the metrics repo.
How we measured
- Reference: our BF16 file. We ran it twice and got identical vectors.
- Held-out set: calib-corpora
eval/neutral(81 Wikipedia articles in 30 languages, 1,000 query and passage pairs) andeval/code(100 source files, 400 pairs). Each chunk of text is split in two: the first quarter becomes the query, the rest becomes the passage. - Metrics:
- the cosine between a file's vector and the reference vector for the same text;
- the share of queries whose best passage, out of every passage in the set, is the same as with the reference;
- 95% intervals by bootstrap over documents.
- Setup:
llama-server --embeddings --pooling mean -fa offon an Apple M4 Max (Metal), llama.cpp commit2690873. All files embed the 2,800 texts in about 25 seconds, and the smaller files are not faster on this machine.
Reproducing a file
python convert_hf_to_gguf.py embeddinggemma-2 --no-lazy --outtype bf16 --model-name embeddinggemma-2 \
--outfile embeddinggemma-2-BF16.gguf
python build_calib.py --pool calib-corpora/pool --tokenizer embeddinggemma-2/tokenizer.json --tokens 1000000 -o calib.txt
llama-imatrix -m embeddinggemma-2-BF16.gguf -f calib.txt -o imatrix.gguf -c 512 -b 512 -ub 512 \
--parse-special --no-ppl --override-kv tokenizer.ggml.add_eos_token=bool:false
M='(attn_(q|k|v|output)|ffn_(gate|up|down))\.weight$'
llama-quantize --imatrix imatrix.gguf \
--tensor-type '^output\.weight$=q8_0' --tensor-type '^token_embd\.weight$=iq4_xs' \
--tensor-type '^blk\.\d+\.(inp_gate|proj)\.weight$=q4_k' \
--tensor-type "^blk\.([0-5])\.$M=q5_k" --tensor-type "^blk\.([6-9]|1[01])\.$M=iq4_xs" \
--tensor-type "^blk\.(1[2-7])\.$M=iq4_xs" --tensor-type "^blk\.(1[89]|2[0-3])\.$M=q4_k" \
embeddinggemma-2-BF16.gguf embeddinggemma-2-AD-Q4_K_M.gguf Q4_K_MAD-Q6_K uses these types:
token_embdin q5_k;inp_gateandprojin q8_0;- blocks 0-5 and 18-23 in q8_0;
- blocks 6-17 in q6_k;
outputin q8_0.
The exact rules are in the quantize logs in the metrics repo.
Model details
- Base: google/embeddinggemma-2, revision
914f7f8, Apache 2.0. - Text model: 270M parameters, 24 layers, width 512. It outputs 768-dimensional vectors that can be shortened to 512, 256 or 128, and has an 8K context.
- Projector: the vision encoder (170M) and the audio encoder (300M).
- llama.cpp: needs a build with the
gemma-embedding2architecture (llama.cpp PR #30054, merged on 6 October 2026). We built and tested with commit26908739bc8a.
