Google publie EmbeddingGemma 2, embeddings multimodaux 740 M
TL;DR
- Three stackable encoders (270M text, 170M vision, 300M audio) share one embedding space, so a text query can retrieve images or audio without a separate pipeline.
- The text-only config requires 191MB RAM on a Pixel 11 Pro, putting multimodal search within reach of mid-range devices without cloud dependency.
- Matryoshka training lets developers truncate 768-dim vectors to 128 dims, giving a tunable storage-versus-accuracy trade-off per deployment.
EmbeddingGemma 2 tient dans environ 191 Mo de RAM active sur un Pixel 11 Pro en mode texte seul, et autour de 567 Mo lorsqu'on active ses encodeurs vision et audio. Google a publié le modèle le 6 octobre 2026 sous licence Apache 2.0, sur Hugging Face et Kaggle.
Le total affiche 740 millions de paramètres, bâtis sur l'architecture Gemma 4, mais découpés en briques optionnelles : 270 M pour le texte, 170 M pour la vision, 300 M pour l'audio. La fenêtre passe à 8 K tokens, de quoi encoder jusqu'à 5,5 minutes d'audio, 29 images ou 58 trames vidéo dans un même espace d'embedding.
"EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space", écrivent Sahil Dua et Henrique Schechter Vera, Research Engineers chez Google DeepMind. Sur MTEB Code, Google revendique un gain de 9,92 points par rapport à EmbeddingGemma 1, de 68,76 à 78,68. Le billet ne publie pas de chiffres comparatifs face aux concurrents sub-1B.
Via Matryoshka Representation Learning, les développeurs peuvent tronquer les vecteurs de 768 à 512, 256 ou 128 dimensions, pour "up to 6x storage reduction" sur les bases vectorielles locales. Les intégrations annoncées au lancement couvrent MediaPipe, LiteRT, vLLM, Ollama, llama.cpp, transformers.js et Qdrant. La sortie rejoint les 154 autres annonces open source que nous suivons sur les 90 derniers jours.
Ce qu'en disent les autres médias
-
The Next Web Lire →
Frames the release through European data-sovereignty concerns: the model runs with no outbound network call, positioning it as a GDPR-friendly RAG option Google did not explicitly market that way.
The practical result is a retrieval pipeline with no network in it.
-
RuntimeWire Lire →
Details the modular size trade-offs (191MB to 567MB) and benchmarks performance degradation when compressing embeddings; also distinguishes EmbeddingGemma 2 from Google's cloud-only Gemini Embedding 2 API.
Different kinds of input are mapped into the same 768-dimensional space, so a text query can be compared with an image, video or audio clip.
-
arXiv (Google DeepMind) Lire →
The underlying research paper with ablation studies, benchmark tables against SigLIP and Voyage AI, and architectural rationale behind the shared embedding space design.
Shared on Bluesky by 3 AI experts
-
Lynn Cherny @arnicas.bsky.social: EmbeddingGemma 2, a tiny multimodal mobile ready embedding model →
Article original publié par blog.google
Lire l'article original →Titre original : Google publie EmbeddingGemma 2, modèle multimodal de 740 M sous Apache 2.0, 191 Mo RAM en texte seul