Google claims EmbeddingGemma 2 outperforms rival embedding models twice its size
Google's 740M embedding model beats rivals 2x its size. On-device RAG just got cheaper.

Why it matters
EmbeddingGemma 2 delivers multimodal (text, image, video, audio, code) embedding at a fraction of memory footprint — 191 MB RAM, open weights — with claimed performance gains over larger competitors. Enables offline RAG and on-device vector search without data egress, shifting the cost/capability frontier for edge and privacy-first deployments.
The key facts
7 to knowEmbeddingGemma 2: 740M parameters
191 MB RAM footprint
Multimodal: text, images, video, audio, code vectorization
Claims to outperform models 2x its size
Open weights release
Runs on-device; enables offline RAG without external servers
Pairs with Gemma 4 for full offline stack
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Google released EmbeddingGemma 2, an open model with 740 million parameters that converts text, images, video, audio, and code into vectors. It runs on-device, needs only about 191 MB of RAM, and outperforms some competing models twice its size, according to Google. Paired with a small open model…