EmbeddingGemma 2 adds images, audio and video to a 740M open model
Google DeepMind's EmbeddingGemma 2 is an Apache 2.0 embedding model with 740M parameters that handles text, images, audio and video on-device.
Google DeepMind released EmbeddingGemma 2 on October 6, 2026. It’s an open embedding model under Apache 2.0, with 740 million parameters, and it maps text, images, audio and video into one shared vector space. A text query can find a photo, a voice memo or a video clip, and it’s built to run on a phone.
An embedding model turns content into vectors, so search works by meaning instead of exact keywords. The first EmbeddingGemma was text-only, and Google says it passed 20 million downloads. This version is built on the Gemma 4 architecture, the open-weights family.
The memory numbers are the part hobbyists should check first. With quantization, Google says it needs about 191MB of active RAM for text-only weights on a Google Pixel 11 Pro, and about 567MB for the full multimodal model. Text-only work needs 270M parameters. Vision (170M) and audio (300M) encoders are optional add-ons.
Storage is another win. With Matryoshka Representation Learning (MRL), a training method that lets output vectors be cut short, you can drop the 768-dimension output to 512, 256 or 128. Google claims up to 6x storage reduction for local vector databases. The context window is 8K tokens, enough for about 5.5 minutes of audio, 29 images or 58 video frames in one input.
Code is the other headline. On the MTEB Code benchmark, the score went from 68.76 to 78.68, a 9.92-point gain. That’s the number to watch if you want local codebase search or retrieval for a coding agent.
How to try it:
- Download the weights from Hugging Face or Kaggle. Google says a Gemini Enterprise Agent Platform listing is coming soon.
- Read the model card for the full evaluation metrics.
- For Android and other apps, use MediaPipe for embedding and retrieval tasks.
The limits: every headline number is Google’s own, and the phone RAM figures come from one device. Nobody outside Google has verified the cross-modal quality yet, so that part is still unproven.
Still, an open, Apache 2.0 multimodal embedder at this size is the thing I’d watch. The text version got used a lot, so if people start building offline search on this one, that will tell us more than any benchmark table.

