blog.google web signal

Google ships EmbeddingGemma 2 for on-device multimodal retrieval

TL;DR

  • Google released EmbeddingGemma 2, a 740M-parameter multimodal embedding model built on Gemma 4 and licensed under Apache 2.0.
  • The model has an 8K token context window with 768-dimensional outputs truncatable to 512, 256, or 128 via Matryoshka learning.
  • Quantized on a Pixel 11 Pro, memory is roughly 191MB text-only and 567MB multimodal; MTEB Code reaches 78.68 versus 68.76 for v1.

Google released a 740-million-parameter embedding model that runs on a phone and natively handles text, code, images, audio and video in a single vector space. On the Google blog, Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera describe EmbeddingGemma 2 as "the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space." It is built on Gemma 4 and licensed under Apache 2.0.

The context window is 8K tokens, "4x larger than EmbeddingGemma 1". Outputs are 768-dimensional and truncatable to 512, 256 or 128 via Matryoshka Representation Learning, which the post says allows "up to 6x storage reduction for local vector databases." A single input can hold up to 5.5 minutes of audio, 29 images, or 58 video frames.

Quantized on a Pixel 11 Pro, the model uses "~191MB active RAM for text-only weights and ~567MB for the full multimodal model." On MTEB Code it scores 78.68, up from 68.76 in the first version, a 9.92-point jump. A 270M-parameter text-only variant ships alongside.

Two researchers we track posted the link on launch day.

Weights are on Hugging Face and Kaggle, with integrations into MediaPipe, LiteRT, transformers.js, WebGPU, vLLM, llama.cpp, Ollama, LMStudio, Qdrant, and Unsloth. Google says the original EmbeddingGemma crossed 20 million downloads.

Shared on Bluesky by 2 AI experts