blog.google web signal

Google ships EmbeddingGemma 2, 740M multimodal embedder, Apache 2.0

TL;DR

  • EmbeddingGemma 2 is a 740M-parameter embedder that maps text, code, images, video and audio into one space under an Apache 2.0 license.
  • On MTEB Code it moves from 68.76 to 78.68, a 9.92-point gain over EmbeddingGemma 1, with an 8K context window that is 4x the prior version.
  • On a Pixel 11 Pro with quantization it needs about 191MB of active RAM for the text-only weights, and Matryoshka training lets 768-dim vectors truncate to 128.

Google has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, code, images, video and audio into a single vector space under an Apache 2.0 license, with weights small enough that the text-only variant runs in about 191MB of RAM on a Pixel 11 Pro.

In a Google blog post, Research Engineers Sahil Dua and Henrique Schechter Vera write that the model "has 740 million parameters, making it optimal for on-device inference" and is "released under a commercially permissive Apache 2.0 license." The headline benchmark move is on code retrieval, where the authors report the MTEB Code score climbing "from 68.76 to 78.68" versus EmbeddingGemma 1, a 9.92-point gain.

The model is built as three stackable encoders — a 270M text core with optional 170M vision and 300M audio encoders — sharing one embedding space, and it "features an 8K token context window (4x larger than EmbeddingGemma 1)." Google says that window accommodates up to 5.5 minutes of audio, 29 images, or 58 video frames in a single call.

Storage is handled through Matryoshka Representation Learning: "Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions," which the authors frame as up to a 6x storage reduction. The post reports the first EmbeddingGemma crossed 20 million downloads before this release.

The launch lands in a crowded week for permissive weights on our radar — Mistral previewed a 1-trillion-parameter open-weight Large 4 the same day — but the strategic bet here is different: not a frontier generator, a tiny Apache-2.0 retrieval backbone intended to live inside phones and apps rather than a datacenter.

Shared on Bluesky by 3 AI experts