Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Google just shipped a 12B multimodal model with no separate encoder. Here's why that matters for inference costs.

Why it matters
Gemma 4 12B represents a shift toward unified architecture designs that reduce computational overhead while maintaining multimodal capability—a key efficiency play as the industry pivots from scale to optimization.
The key facts
5 to knowGemma 4 12B is encoder-free multimodal architecture
Unified model design (single forward pass for vision + text)
Published by Google DeepMind on June 9, 2026
Positions efficiency and inference cost reduction as competitive advantage
Competes in sub-15B parameter multimodal space
Go to the source
Google DeepMind Blogdeepmind.google