Use AIUnderstand AIBuild with AI
Build with AI·Tip·· 1 min read

Matryoshka embeddings: how many dimensions do you actually need?

A practical rule of thumb for truncating embeddings to cut vector-DB cost, with the numbers from EmbeddingGemma 2.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Many recent embedding models are trained with Matryoshka Representation Learning (MRL). The first k dimensions of each vector are a useful embedding on their own, so you can cut vectors short and save storage and search time.

The numbers (EmbeddingGemma 2)

Dims MTEB (eng) MTEB (code) Multimodal (MMEB) Storage
768 68.46 78.68 59.01 1×
512 68.41 77.24 58.38 1.5× smaller
256 67.78 76.18 56.24 3× smaller
128 65.68 71.41 45.65 6× smaller

Rule of thumb

  • Text-only RAG: start at 256. You lose under 1 point on English retrieval and use a third of the storage.
  • Code search: use 512 or the full 768, because code drops off faster.
  • Images, audio or video: stay at 768 unless your own evals say otherwise. Quality falls sharply at 128.

Two rules that bite people

  1. Re-normalize after truncating. A cut-down unit vector no longer has unit length, so cosine scores come out quietly wrong. In sentence-transformers, use truncate_dim=256, normalize_embeddings=True.
  2. Queries and documents must have the same dimension. If you re-index at 256, queries must be encoded at 256 too.
emb = model.encode(texts, truncate_dim=256, normalize_embeddings=True)

See the full walkthrough in EmbeddingGemma 2: semantic search on a laptop.

Related