By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google DeepMind Releases EmbeddingGemma 2 Multimodal Model
Google DeepMind has released EmbeddingGemma 2, an open-source multimodal embedding model designed to convert text, code, images, video, and audio into a unified 768-dimensional vector space. This model, built upon the Gemma 4 architecture, features 740 million parameters and boasts an 8K token context window, representing a significant upgrade from its predecessor. The release is accompanied by an Apache 2.0 license, making it freely available for use and modification. EmbeddingGemma 2 is engineered for applications such as on-device search, classification tasks, and privacy-first Retrieval Augmented Generation (RAG) pipelines. The model's ability to generate embeddings locally ensures that sensitive data remains on the user's device, thereby reducing latency and enabling offline functionality.
The core innovation of EmbeddingGemma 2 lies in its capacity to create a single vector space that encompasses multiple modalities. This means a text query can retrieve relevant images, or a voice memo can be used to find a corresponding video clip. The model is modular, comprising a text and code backbone with 270 million parameters (further divided into a 130M transformer and a 140M embedder), a 170M parameter vision encoder, and an optional 300M parameter audio encoder. Developers can choose to load only the components they require, leading to a flexible parameter count: 270M for text-only, 440M for text and vision, 570M for text and audio, or the full 740M for all modalities. Crucially, all configurations share a single vector space, allowing embeddings generated by a text-only setup to match documents embedded by the complete model.
The context window has been expanded to 8,192 tokens, which is four times larger than that of EmbeddingGemma 1. This increased capacity can accommodate approximately 29 images, 58 video frames, or 5.5 minutes of audio, enhancing its ability to process richer and longer inputs. Google's research team has reported that EmbeddingGemma 2 achieves leading scores among multimodal embedders with fewer than 1 billion parameters on the MTEB Code and MAEB benchmarks. Specifically, it scored 78.68 on MTEB Code v1, an improvement from EmbeddingGemma 1's score of 68.76, and 61.36 on MTEB multilingual v2, slightly surpassing EmbeddingGemma 1's 61.15. The model's weights are now accessible on platforms like Hugging Face and Kaggle, with builds available for Ollama, llama.cpp GGUF, and LiteRT, facilitating immediate deployment by developers.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.