Google DeepMind Releases Gemma 4 12B: An Encoder-Free Multimodal Model with Native audio that runs on a 16 GB laptop
Encoder-free architecture. Native audio. 12B parameters. Gemma 4 runs on 16GB—and Google just open-sourced it.

Why it matters
Google DeepMind's Gemma 4 12B challenges the multimodal model scaling narrative by delivering vision + audio directly to the LLM backbone at consumer-grade hardware specs. Apache 2.0 licensing signals aggressive competition in the open-weights space.
The key facts
5 to knowGemma 4 12B model released
Encoder-free multimodal architecture (vision + audio native)
Runs on 16GB RAM laptop
Apache 2.0 open-source license
Published June 3, 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Gemma 4 12B feeds vision and audio straight into the LLM backbone, running locally under an Apache 2.0 license.