Embeddings · evidence reviewed October 7, 2026

MTEB evidence for choosing embeddings.

Use embedding benchmarks to choose candidates for a corpus-specific test. This page separates a historical publisher comparison from a newly verified model release. It does not claim to be an exhaustive October leaderboard.

Current candidate coverageHistorical comparisonOfficial MTEB leaderboard

Current candidate · primary source checked 2026-10-07

EmbeddingGemma 2

Google lists the model card as updated October 6, 2026. The card describes 740M total parameters, with a selectively loadable 270M text component, native 768-dimensional embeddings and an 8,192-token context. It names Apache 2.0 as the license. These are publisher specifications. Google model card; dated model-card index.

Publisher benchmark identityMetric / configurationReported score
MTEB, multilingual, v2Mean(Task), mixed task metrics, 768d full precision61.36
MTEB, English, v2Mean(Task), 768d full precision68.46
MTEB, code, v1Mean(Task), NDCG@10, 768d full precision78.68

No comparable score for this model exists in the historical table below. These named v2/v1 results remain separate: they do not establish a rank against scores whose suite revision is unspecified. This page has no CodeSOTA corpus-specific run for EmbeddingGemma 2.

The official usage instructions specify task prefixes and normalization after dimension truncation. Validate the exact checkpoint, precision and document/query formatting used by your serving stack.

Historical source comparison · November 2025 (publisher description)

One publisher table, ten models.

Source identity: KaLM-Embedding-Gemma3-12B-2511 model-card MMTEB table. The row order is the publisher’s Borda rank; it is not descending Mean(Task). Mean(Task) and Mean(TaskType) are different aggregates. Pinned source revision, checked 2026-10-07.

Source Borda rankModelMean(Task)Mean(TaskType)RetrievalSTS
1KaLM-Embedding-Gemma3-12B-251172.3262.5175.6679.02
2llama-embed-nemotron-8b69.4661.0968.6979.41
3Qwen3-Embedding-8B70.5861.6970.8881.08
4gemini-embedding-00168.3759.5967.7179.40
5Qwen3-Embedding-4B69.4560.8669.6080.86
6Qwen3-Embedding-0.6B64.3456.0164.6576.17
7gte-Qwen2-7B-instruct62.5155.9360.0873.98
8Linq-Embed-Mistral61.4754.1458.6974.86
9multilingual-e5-large-instruct63.2255.0857.1276.81
10embeddinggemma-300m61.1554.3162.4974.73

The earlier page merged additional model scores from unidentified snapshots and mislabeled some task columns. Those rows have been withdrawn from this comparison. This table preserves the identities and values found together in the cited publisher report. Dataset-level uncertainty and independent reproduction remain unavailable here.

Choose for your corpus

A benchmark score starts an evaluation.

  1. Choose an exact MTEB suite from the official evaluator. Record its version, task revisions, split and aggregation method.
  2. Hold out queries and relevance labels from your own documents. Keep language mix, chunking, candidate corpus, prompts and retrieval settings fixed across models.
  3. Compare retrieval quality on that fixed corpus, then measure latency, memory and serving cost on the intended hardware. Do not divide an aggregate benchmark score by parameters to invent an efficiency ranking.
  4. Check the selected checkpoint’s code and weight licenses, context limit, query/document instructions and normalization. Model families can contain different releases with different terms.
  5. Evaluate a reranker separately on the same retrieved candidates. Embedding similarity scores and cross-encoder reranker results are different tasks.

The original MTEB paper and the later MMTEB paper describe different benchmark coverage. A language-filtered list of tasks from a current package is not a frozen recreation of either paper’s suite.

Reranking guide · Evidence methodology · Official leaderboard and newer submissions