Google releases EmbeddingGemma 2 for private multimodal search
The open 740M-parameter model maps text, code, images, audio and video into one embedding space for local search and RAG apps.
Google DeepMind released EmbeddingGemma 2, an open multimodal embedding model for local search and retrieval apps. The model is built on the Gemma 4 architecture, released under Apache 2.0, and has 740 million parameters.
The important change is that EmbeddingGemma is no longer just about text. Google says the new model maps text, code, images, audio and video into a shared embedding space, so developers can search across mixed media with one model. Example use cases include finding a video clip from a voice memo or searching audio recordings with a text query.
Google frames the launch around privacy-first RAG and on-device workloads. The model can run with modular encoders, from about 270M parameters for text-only use to the full multimodal setup, and supports shorter vectors through Matryoshka Representation Learning to reduce local vector database storage. Google says quantized text-only weights can use about 191MB of active RAM, while the full multimodal model can use about 567MB on a Pixel 11 Pro.
For builders, this matters because embeddings are the boring layer that makes search, memory and retrieval feel useful. A capable open model that runs locally lowers the barrier for private document, media and assistant workflows.
Sources
- Google DeepMinddeepmind.google
- Hugging Facehuggingface.co