EmbeddingGemma 2: A versatile multimodal embedding model for on-device use
EmbeddingGemma 2 is being introduced as a powerful model designed for on-device multimodal embeddings, capable of integrating text, images, audio, and video into a single embedding space.
Originally launched last year, EmbeddingGemma aimed to offer a lightweight solution for delivering high-quality text embeddings, enabling applications to organize, search, and connect information directly on consumer devices. The developer community responded enthusiastically, with the model surpassing 20 million downloads. Developers utilized it to enhance on-device search tools and privacy-focused retrieval augmented generation (RAG) pipelines.
Now, EmbeddingGemma 2 is released, expanding its capabilities beyond text to include code, images, video, and audio within a unified embedding space. Built on the Gemma 4 architecture and licensed under the commercially flexible Apache 2.0 license, EmbeddingGemma 2 features 740 million parameters, making it ideal for on-device inference. The model can efficiently locate a specific video clip from a voice memo or search hours of audio recordings using a text query, all processed by a single, inherently multimodal model.
The technology underpinning EmbeddingGemma 2 is the same as that used in the Gemini Embedding models. It matches the multilingual text performance of its predecessor while improving significantly on code tasks, achieving a 9.92-point increase in MTEB Code performance, from 68.76 to 78.68. This makes it particularly useful for indexing local codebases, performing semantic code searches, and retrieving coding agents. Across images, videos, documents, and audio, it establishes new benchmarks in quality-per-parameter for models under 1 billion parameters, even outperforming some specialized models that are more than twice its size.
Further details on evaluation metrics and model specifications are available in the EmbeddingGemma 2 model card.
EmbeddingGemma 2 empowers edge hardware with advanced capabilities, ensuring data privacy, reducing latency, and enabling developers to build cross-modal search and retrieval systems that function entirely offline. When combined with generative models like Gemma 4, it facilitates RAG pipelines on devices that comprehend complex multimodal data. Sharing the text tokenizer and audio encoder of Gemma 4, both models can run in a unified pipeline with a reduced overall memory footprint.
Users can employ text or images to locate top matches in media libraries based on semantic similarity through Google AI Edge Gallery’s Instant Media Search. Specific moments in videos can be found using text or audio queries via the Video Moments Finder in the same gallery.
EmbeddingGemma 2 can be paired with Gemma 4 for local file retrieval combined with contextual reasoning, as demonstrated in the Google AI Edge Foresight app. Additionally, it can create real-time decision engines that leverage multimodal context for tasks such as classification, routing, and prediction using the MediaPipe Decision Task API.
For more information on building on-device search and RAG systems with LiteRT, refer to the Google AI Edge blog post. Developers can explore comprehensive guides, documentation, and resources for inference and fine-tuning.