Google expands EmbeddingGemma beyond text to images, audio and video
Google's EmbeddingGemma 2 now handles text, images, audio and video in a single embedding space — and runs on a smartphone.

Why it matters
Multimodal embedding models small enough for on-device deployment shift the practical boundary of what AI work can happen locally. For practitioners, this enables new embedding-grounded applications (search, retrieval, clustering) across media types without cloud calls; for enthusiasts, it marks progress in the model capability stack toward unified representations.
The key facts
5 to knowEmbeddingGemma 2 is open multimodal, expanding from text-only (v1, Sept 2025)
Handles text, images, audio, and video in shared embedding space
Small enough to run on a smartphone
Released October 2026
Open weight model (implied by 'open multimodal')
The story so far
Earlier coverage of this storyline
- Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4MarkTechPost
- EmbeddingGemma 2: an open, lightweight multimodal embedding modelGoogle DeepMind Blog
- EmbeddingGemma 2Simon Willison
- This story
Go to the source
SiliconAnglesiliconangle.com
Publisher excerpt: Google LLC today released EmbeddingGemma 2, an open multimodal embedding model small enough to run on a smartphone. The release takes the EmbeddingGemma line beyond text, which was all the first version handled when Google introduced it in September 2025. Images, audio and video now share one…