EmbeddingGemma 2: semantic search on a laptop in 20 lines
Google's new 740M multimodal embedding model runs on a CPU. Here's how to build a working semantic search over your own documents, and the prefixes that make it accurate.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
EmbeddingGemma 2 is Google DeepMind’s new open embedding model. It maps text, code, images, video and audio into one shared 768-dimensional vector space. The text path is only 270M parameters, so it runs on a laptop CPU. That makes it a strong default for local search and RAG.
What’s new compared to v1
- Multimodal: the vision (170M) and audio (300M) encoders are optional modules you can skip loading.
- Much better at code: MTEB code goes from 68.76 to 78.68.
- 8K context, up from 2K.
- Matryoshka embeddings: you can truncate vectors to 512, 256 or 128 dims and lose very little quality.
| Dims | MTEB multilingual | MTEB code | Storage |
|---|---|---|---|
| 768 | 61.36 | 78.68 | 1× |
| 256 | 60.41 | 76.18 | 3× smaller |
| 128 | 57.89 | 71.41 | 6× smaller |
Install
pip install -U sentence-transformers transformersLoad text-only (saves ~470M params)
If you only embed text, skip the vision and audio encoders:
import torchfrom sentence_transformers import SentenceTransformer
use_bf16 = torch.cuda.is_available() and torch.cuda.is_bf16_supported()dtype = torch.bfloat16 if use_bf16 else torch.float32
model = SentenceTransformer( "google/embeddinggemma-2", model_kwargs={"torch_dtype": dtype}, config_kwargs={"vision_config": None, "audio_config": None},)A complete semantic search
EmbeddingGemma 2 was trained with task prefixes. Queries and documents get different prompts, and sentence-transformers applies them for you through prompt_name:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
docs = [ "The northern lights are caused by charged particles from the sun hitting the atmosphere.", "Photosynthesis converts light energy into chemical energy in plants.", "A mixture-of-experts model activates only a subset of its parameters per token.", "Sourdough bread rises thanks to wild yeast and lactic acid bacteria.",]
doc_emb = model.encode(docs, prompt_name="Document", normalize_embeddings=True)query_emb = model.encode("why does bread rise?", prompt_name="SearchQuery", normalize_embeddings=True)
scores = model.similarity(query_emb, doc_emb)[0]for score, doc in sorted(zip(scores.tolist(), docs), reverse=True)[:2]: print(f"{score:.3f} {doc}")If your documents have real titles, format them yourself as title: {title} | text: {content}. prompt_name="Document" assumes title: none.
Shrink your vector DB 3× with Matryoshka
Pass truncate_dim and keep normalize_embeddings=True. Queries and documents must use the same dimension:
doc_emb = model.encode(docs, prompt_name="Document", truncate_dim=256, normalize_embeddings=True)query_emb = model.encode(query, prompt_name="SearchQuery", truncate_dim=256, normalize_embeddings=True)Quality is close to lossless down to 256 dims for text. Before you use 128 dims for images or audio, check it against your own data.
Gotchas
- Never use
float16. The activations overflow fp16 and you silently get NaNs or worse embeddings. Usebfloat16on GPUs that support it andfloat32everywhere else, including CPUs. (More on this.) - Use the right prefix for the task. Classification, clustering and similarity use symmetric prompts (
Classification,Clustering,SentenceSimilarity). Retrieval usesSearchQuery+Document, and code search usesCodeRetrieval.
When to pick it
Pick it for local or on-device search, multilingual RAG, code search, or “search my photos with text”. If you need the best possible English-only retrieval and have a GPU budget, compare it against larger embedding models on your own data first.
Related
Why your EmbeddingGemma 2 vectors are NaN (and the one-line fix)
Loading EmbeddingGemma 2 in float16 silently breaks it. Here's how to detect it and pick the right dtype for your hardware.
google/embeddinggemma-2
Matryoshka embeddings: how many dimensions do you actually need?
A practical rule of thumb for truncating embeddings to cut vector-DB cost, with the numbers from EmbeddingGemma 2.
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash