Matryoshka embeddings: how many dimensions do you actually need?
A practical rule of thumb for truncating embeddings to cut vector-DB cost, with the numbers from EmbeddingGemma 2.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Many recent embedding models are trained with Matryoshka Representation Learning (MRL). The first k dimensions of each vector are a useful embedding on their own, so you can cut vectors short and save storage and search time.
The numbers (EmbeddingGemma 2)
| Dims | MTEB (eng) | MTEB (code) | Multimodal (MMEB) | Storage |
|---|---|---|---|---|
| 768 | 68.46 | 78.68 | 59.01 | 1× |
| 512 | 68.41 | 77.24 | 58.38 | 1.5× smaller |
| 256 | 67.78 | 76.18 | 56.24 | 3× smaller |
| 128 | 65.68 | 71.41 | 45.65 | 6× smaller |
Rule of thumb
- Text-only RAG: start at 256. You lose under 1 point on English retrieval and use a third of the storage.
- Code search: use 512 or the full 768, because code drops off faster.
- Images, audio or video: stay at 768 unless your own evals say otherwise. Quality falls sharply at 128.
Two rules that bite people
- Re-normalize after truncating. A cut-down unit vector no longer has unit length, so cosine scores come out quietly wrong. In sentence-transformers, use
truncate_dim=256, normalize_embeddings=True. - Queries and documents must have the same dimension. If you re-index at 256, queries must be encoded at 256 too.
emb = model.encode(texts, truncate_dim=256, normalize_embeddings=True)See the full walkthrough in EmbeddingGemma 2: semantic search on a laptop.
Related
EmbeddingGemma 2: semantic search on a laptop in 20 lines
Google's new 740M multimodal embedding model runs on a CPU. Here's how to build a working semantic search over your own documents, and the prefixes that make it accurate.
google/embeddinggemma-2
Why your EmbeddingGemma 2 vectors are NaN (and the one-line fix)
Loading EmbeddingGemma 2 in float16 silently breaks it. Here's how to detect it and pick the right dtype for your hardware.
google/embeddinggemma-2