Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

Google - Multi-Modal, Document Understanding, Edge AI

EmbeddingGemma 2

EmbeddingGemma 2 puts text, code, images, video and audio in one shared 768-dimension vector space, under Apache 2.0, in about 191 MB on a phone. What changed since v1, what the MTEB numbers actually say, and the float16 warning that will quietly ruin your retrieval.

License Apache 2.0
License Apache 2.0
TL;DR
  • Text, code, images, video and audio in one shared 768-dimension space
  • Modular: 270M text only, 740M with every encoder attached
  • Apache 2.0, where EmbeddingGemma 1 shipped under the Gemma licence
System Requirements
RAMAbout 191 MB text-only on a Pixel 11 Pro quantized, 567 MB full multimodal
GPUNot needed. GGUF for llama.cpp, MLX for Apple Silicon, LiteRT for mobile
VRAMNone required; runs on CPU, phone, browser via WebGPU or any GPU
✓ Apple Silicon

Your RAG pipeline probably has a model for text, another for images, and nothing at all for audio. Three indexes, three sets of vectors that cannot be compared, and a query that only ever searches one of them.

EmbeddingGemma 2 puts text, code, images, video and audio into one shared vector space. Same index, same query, 768 dimensions regardless of what went in. It is Apache 2.0, and the text-only configuration runs in about 191 MB on a phone.

Here is what changed, what the benchmarks actually say, and the one line in the model card that will cost you an afternoon if you skip it.

What actually changed since v1

EmbeddingGemma 1 was a 300M text-only model, and it became one of the most-used embedding models on Hugging Face: 3.6 million downloads in the last 30 days. Plenty of production RAG stacks are running it right now.

EmbeddingGemma (300M)EmbeddingGemma 2
ModalitiesTextText, code, image, video, audio
Text parameters300M270M
Total parameters300M740M with every encoder loaded
Context2,048 tokens8,192 tokens
Output dimensions768 (512 / 256 / 128)768 (512 / 256 / 128)
LicenceGemmaApache 2.0

Two of those rows deserve more attention than they are getting.

The text tower got smaller. 300M down to 270M, while the context window went up 4x. The model you load for plain text retrieval is now lighter than the one it replaces.

The licence changed. v1 shipped under Google's own Gemma Terms of Use. v2 is Apache 2.0. If a custom licence was the reason EmbeddingGemma never made it past your legal review, that reason is gone.

One embedding space for five modalities

This is the part that changes how you would build the pipeline.

Normally, adding images to a text RAG system means a second model, a second index, and a fusion step you have to invent yourself to decide which result wins. Cross-modal search, where a text query finds an image, needs a model trained for that specific pair.

EmbeddingGemma 2 is modular. You load the text tower always, then attach the encoders you need:

ConfigurationParameters
Text only270M
Text + vision440M
Text + audio570M
Everything740M

Whatever you attach, the output is a 768-dimension vector in the same space. A text query and an audio clip land in coordinates you can compare directly with cosine similarity. One index. One nearest-neighbour search. No fusion logic.

The 8,192 token context covers what Google describes as minutes of audio or video, so a clip is one vector rather than a chunking problem.

Worth reading against PageIndex, which argues you should skip vectors altogether and let a model walk a document tree instead. These are genuinely opposed answers to the same question, and the choice turns on whether your corpus is reasoning-shaped or similarity-shaped. Also worth comparing to PixelRAG, which retrieves screenshots rather than parsing text, since the vision tower here covers some of the same ground with a general-purpose model.

The numbers

All of these are Google's own published figures. Nobody here has run MTEB.

BenchmarkEmbeddingGemma (300M)EmbeddingGemma 2
MTEB Multilingual v2, Mean (Task)61.1561.36
MTEB English v2, Mean (Task)69.6768.46
MTEB Code v1, NDCG@1068.7678.68

Code retrieval is the headline: up 9.92 points. If you are building search over a codebase, that is a large jump from a model that also got smaller on the text side.

Multilingual is flat, 61.15 to 61.36, which matches Google's claim that v2 holds v1's multilingual performance.

English went the other way, 69.67 down to 68.46. Google's announcement does not mention that row. It is a 1.21 point drop and it is not catastrophic, but if English text retrieval is the only thing your pipeline does, v1 still scores higher on the benchmark Google itself publishes for both.

For vision, audio and video, Google says the model "sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size." That is a claim without a public number attached, so treat it as a claim.

Running it

The model is on Hugging Face as google/embeddinggemma-2, and the ecosystem landed with it: GGUF builds from ggml-org for llama.cpp and from Unsloth, an MLX build for Apple Silicon, and LiteRT packages for mobile. Google lists transformers, sentence-transformers, vLLM, SGLang, Ollama, LM Studio and transformers.js with WebGPU.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

docs = model.encode(["Postgres stores vectors with pgvector."])
query = model.encode(["how do I store embeddings in postgres"])

Three things will bite you if you do not read the model card.

Do not load it in float16. The card is explicit: the activation range exceeds float16's dynamic range. Use bfloat16. This is the kind of failure that produces quietly wrong vectors rather than a crash, which is far worse, because your retrieval just gets mysteriously bad.

Task prefixes are not optional. The model expects to be told what the embedding is for. Retrieval is asymmetric, so a query and a document get different instructions (SearchQuery and Document), while symmetric tasks like classification and clustering use their own. Embed your documents with the query prefix by mistake and you will lose accuracy for reasons that are invisible in the output.

You can truncate the vectors. Matryoshka training means the 768-dimension output can be cut to 512, 256 or 128 and still work, because the most important information is packed into the earliest dimensions. At 128 dimensions you store a sixth of the bytes per vector. On a million-document index that is the difference between a comfortable database and an expensive one. Measure the recall cost on your own corpus before committing; the point is that the option is free and built in.

If you want somewhere to put the vectors, our pgvector and Supabase RAG build is the pipeline this slots into directly.

Who should upgrade, and who should not

Upgrade if you retrieve code. Nearly 10 points on MTEB Code is not a tuning difference.

Upgrade if you want more than text. This is the whole reason the model exists, and there is no comparable sub-1B Apache 2.0 option that puts five modalities in one space.

Upgrade if the Gemma licence blocked you. Apache 2.0 removes that conversation entirely.

Upgrade if you need the context. 2,048 tokens forces aggressive chunking. 8,192 does not.

Stay on v1 if you do English-only text retrieval, you are happy with your chunk size, and you have already tuned against it. It scores slightly higher on English, it is a known quantity with 3.6 million downloads a month behind it, and "slightly better on the benchmark" plus "already working in production" is a reasonable place to stand.

Limitations

  • Nothing here is independently tested. Every benchmark is Google's own, on Google's harness.
  • It is days old. The repo shows roughly 364 downloads against v1's 3.6 million. Nobody has found the sharp edges yet.
  • The quantized builds are newer still. The GGUF and MLX repos went up with essentially no download history, so treat them as untested paths rather than settled ones.
  • The on-device numbers are one phone. 191 MB and 567 MB are measured on a Pixel 11 Pro with quantization. Your hardware and your quantization will differ.
  • 100+ languages does not mean evenly. The card says so directly: performance is not equal across them.
  • Multimodal quality has no published number. The quality-per-parameter claim is the part of this release with the least evidence attached, and it is also the part doing the most marketing work.

Sources and further reading

Ten minutes: pull google/embeddinggemma-2, embed fifty documents from a corpus you already know, and run your three worst queries against it. Then truncate the vectors to 256 dimensions and run them again. If recall holds, you just cut your index to a third of its size for free.

Not independently tested. Every benchmark figure above is published by Google; we have not reproduced MTEB. Parameter counts, dimensions, context length, licences and the float16 warning were verified directly against the Hugging Face model cards for both versions on the date below.
Date checked: 2026-10-06

Prev Article
Qwen-Image-2.1
Next Article
OpenThinker-32B

Related to this topic: