OFICIAL Google Blog

EmbeddingGemma 2: an open, lightweight multimodal embedding model

What happened
Based on Google Blog · Oct 06, 2026

Google released EmbeddingGemma 2, a lightweight multimodal embedding model that unifies text, images, audio, video, and code into a single embedding space for on-device applications.

EmbeddingGemma 2: an open, lightweight multimodal embedding model
Google Blog — Google
Key points
·
EmbeddingGemma 2 unifies text, images, audio, video, and code in a single embedding space for on-device processing
·
The model improves code performance by 9.92 points on MTEB Code, raising the score from 68.76 to 78.68
·
EmbeddingGemma 2 is built on Gemma 4 architecture and released under an Apache 2.0 license
Key numbers
·
92-point increase on the MTEB Code benchmark, raising the score from 68.
·
It also sets a new benchmark in quality-per-parameter for sub-1-billion-parameter models across image, video, document, and audio tasks.

Google has launched EmbeddingGemma 2, a new model designed to process text, images, audio, video, and code within a unified embedding space directly on consumer hardware. The model builds on the Gemma 4 architecture and is released under an Apache 2.0 license, making it accessible for commercial use. Unlike its predecessor, which focused solely on text, this version expands multimodal capabilities while maintaining efficiency for on-device inference.

The model introduces significant improvements in code performance, achieving a 9.92-point increase on the MTEB Code benchmark, raising the score from 68.76 to 78.68. It also sets a new benchmark in quality-per-parameter for sub-1-billion-parameter models across image, video, document, and audio tasks. Developers can use it to build cross-modal search tools that operate entirely offline, reducing latency and enhancing data privacy.

EmbeddingGemma 2 is optimized for edge hardware, enabling local file retrieval and semantic search without cloud dependency. It supports applications such as finding specific video clips from voice memos or searching audio recordings using text queries. The model’s shared architecture with Gemma 4 allows for efficient combined use in pipelines, minimizing memory usage while improving performance.

Google provides multiple tools and resources for developers to integrate EmbeddingGemma 2, including the Google AI Edge Gallery’s Instant Media Search and Video Moments Finder. The MediaPipe Decision Task API enables real-time decision engines using multimodal context, while LiteRT supports on-device search and retrieval systems.

Original source → Deals on Clipraptor.com →