Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 2 min read

EmbeddingGemma 2: semantic search on a laptop in 20 lines

Google's new 740M multimodal embedding model runs on a CPU. Here's how to build a working semantic search over your own documents, and the prefixes that make it accurate.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

EmbeddingGemma 2 is Google DeepMind’s new open embedding model. It maps text, code, images, video and audio into one shared 768-dimensional vector space. The text path is only 270M parameters, so it runs on a laptop CPU. That makes it a strong default for local search and RAG.

What’s new compared to v1

  • Multimodal: the vision (170M) and audio (300M) encoders are optional modules you can skip loading.
  • Much better at code: MTEB code goes from 68.76 to 78.68.
  • 8K context, up from 2K.
  • Matryoshka embeddings: you can truncate vectors to 512, 256 or 128 dims and lose very little quality.
Dims MTEB multilingual MTEB code Storage
768 61.36 78.68 1×
256 60.41 76.18 3× smaller
128 57.89 71.41 6× smaller

Install

Terminal window
pip install -U sentence-transformers transformers

Load text-only (saves ~470M params)

If you only embed text, skip the vision and audio encoders:

load.py
import torch
from sentence_transformers import SentenceTransformer
use_bf16 = torch.cuda.is_available() and torch.cuda.is_bf16_supported()
dtype = torch.bfloat16 if use_bf16 else torch.float32
model = SentenceTransformer(
"google/embeddinggemma-2",
model_kwargs={"torch_dtype": dtype},
config_kwargs={"vision_config": None, "audio_config": None},
)

EmbeddingGemma 2 was trained with task prefixes. Queries and documents get different prompts, and sentence-transformers applies them for you through prompt_name:

search.py
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
docs = [
"The northern lights are caused by charged particles from the sun hitting the atmosphere.",
"Photosynthesis converts light energy into chemical energy in plants.",
"A mixture-of-experts model activates only a subset of its parameters per token.",
"Sourdough bread rises thanks to wild yeast and lactic acid bacteria.",
]
doc_emb = model.encode(docs, prompt_name="Document", normalize_embeddings=True)
query_emb = model.encode("why does bread rise?", prompt_name="SearchQuery", normalize_embeddings=True)
scores = model.similarity(query_emb, doc_emb)[0]
for score, doc in sorted(zip(scores.tolist(), docs), reverse=True)[:2]:
print(f"{score:.3f} {doc}")

If your documents have real titles, format them yourself as title: {title} | text: {content}. prompt_name="Document" assumes title: none.

Shrink your vector DB 3× with Matryoshka

Pass truncate_dim and keep normalize_embeddings=True. Queries and documents must use the same dimension:

doc_emb = model.encode(docs, prompt_name="Document", truncate_dim=256, normalize_embeddings=True)
query_emb = model.encode(query, prompt_name="SearchQuery", truncate_dim=256, normalize_embeddings=True)

Quality is close to lossless down to 256 dims for text. Before you use 128 dims for images or audio, check it against your own data.

Gotchas

  • Never use float16. The activations overflow fp16 and you silently get NaNs or worse embeddings. Use bfloat16 on GPUs that support it and float32 everywhere else, including CPUs. (More on this.)
  • Use the right prefix for the task. Classification, clustering and similarity use symmetric prompts (Classification, Clustering, SentenceSimilarity). Retrieval uses SearchQuery + Document, and code search uses CodeRetrieval.

When to pick it

Pick it for local or on-device search, multilingual RAG, code search, or “search my photos with text”. If you need the best possible English-only retrieval and have a GPU budget, compare it against larger embedding models on your own data first.

Related