One small model for text, code, images and audio

EmbeddingGemma 2 is the first version of Google's small embedding model that a team can ship under a standard open licence. Google released it on 6 October 2026 under Apache 2.0. The first EmbeddingGemma carries the Gemma licence tag on its Hugging Face page.

The model maps text, code, images, video and audio into one shared vector space. It is built on the Gemma 4 architecture and sized for phones and laptops.

The model card splits it into modules a team can load separately.

  • A 270M text core, made of a 130M transformer and a 140M embedder
  • An optional 170M vision encoder
  • An optional 300M audio encoder

Hugging Face's API counts 744,371,512 parameters in the published weights. Output vectors have 768 dimensions and can be cut to 512, 256 or 128. The context window is 8,192 tokens, which Google says is four times the first model's.

Text scores hold while code jumps

On the card, the multilingual MTEB score moves from 61.15 to 61.36. Code retrieval moves from 68.76 to 78.68, the jump Google leads with.

Google says the model achieves "leading scores among sub-1B multimodal embedders for its size" on MTEB Code and the audio benchmark MAEB. On Google's own code chart, EmbeddingGemma 2 plots above every other sub-1B model shown. That chart places it at its 270M text size. Qwen3-Embedding-0.6B, which Hugging Face lists under Apache 2.0, plots lower.

The audio chart does not back the claim as written. It plots jina-embeddings-v5-omni-nano above EmbeddingGemma 2. Hugging Face lists that rival at 985,984,512 parameters under CC BY-NC 4.0, a non-commercial licence.

Google chart of Massive Audio Embedding Benchmark scores against model size, with jina-embeddings-v5-omni-nano plotted above EmbeddingGemma 2 Google's audio chart plots a sub-1B rival above EmbeddingGemma 2.
Source: Google, EmbeddingGemma 2 launch post, 6 October 2026

Among sub-1B models on Google's audio chart, EmbeddingGemma 2 leads only if the comparison is limited to commercially usable licences.

Apache weights, a use policy and early downloads

  • Licence: Hugging Face tags the weights apache-2.0. The card still says deployments "must adhere to the Gemma Prohibited Use Policy".
  • Activity: at 03:13 SGT on 7 October, Hugging Face showed 364 downloads in its 30-day count and 218 likes. The first model showed 3,608,007 downloads over the same window.
  • Maturity: usable. Weights are on Hugging Face and Kaggle, and Google lists transformers, sentence-transformers, vLLM, llama.cpp, Ollama and MLX support.
  • Independent scores: the public MTEB results repository had no EmbeddingGemma 2 folder when The Frontier checked at 03:13 SGT. Every figure here is Google's.

Aleph Alpha also put Apache 2.0 on its Kolibri weights this month. See Aleph Alpha Kolibri-1 opens 78B German and English MoE weights under Apache 2.0, with 3.46B active parameters.

Pick it for local code and audio search

  • Pick it for on-device code search or mixed media search in a commercial product, where the first model's Gemma licence was a hurdle.
  • Keep the first model for text-only English and multilingual search if its results already meet your needs.
  • Load only the 270M text core for text work. Google puts that at about 191MB of active RAM on a Pixel 11 Pro with quantisation, against about 567MB for the full model.
  • Use a hosted service if you do not need offline search. Cloudflare's AI Search added image embeddings from Qwen3-VL-Embedding at general availability. See Cloudflare AI Search starts billing on 1 November, and semantic queries cost 7.5 times full-text.

Truncation and float16 break quietly

  • Cutting vectors to 128 dimensions costs about 3.5 points on multilingual text and more than 13 on images and video. The card's MMEB overall score falls from 59.01 to 45.65.
  • Truncated vectors must be normalised again. The card warns that skipping this "degrades ranking quality silently".
  • Queries and documents must use the same dimension.
  • Do not run it in float16. The card says the model then returns NaN or degraded embeddings without raising an error. Use bfloat16 or float32.
  • Prompts matter. The card sets a task prefix for each use, such as code retrieval or question answering.

Google post, model cards and MTEB results

Google's launch post, the four Hugging Face model cards and the MTEB results repository are listed under Sources below.