Introducing Gemma 3
Google just shipped Gemma 3—the most capable model that fits on a single GPU. Here's why that changes the game for every startup.

Why it matters
Gemma 3 represents a significant capability leap in efficient, on-device model deployment. For founders and enterprises, this means competitive AI capabilities without massive infrastructure costs—a direct challenge to cloud-dependent model strategies.
The key facts
5 to knowGemma 3 announced by Google DeepMind
Optimized for single GPU/TPU deployment
Positions as 'most capable' in efficient inference class
Published March 12, 2025
Inference efficiency is core differentiation vs. larger closed models
Go to the source
Google DeepMind Blogdeepmind.google
Publisher excerpt: The most capable model you can run on a single GPU or TPU.