Announcing Gemma 3n preview: Powerful, efficient, mobile-first AI
Google just released Gemma 3n—a multimodal open model optimized for on-device AI. Here's why that matters for your inference costs.

Why it matters
Google's Gemma 3n signals a strategic shift toward efficient, mobile-first AI models that can run locally. For founders building consumer AI apps, this opens a new cost/latency tradeoff—and directly competes with proprietary model pricing.
The key facts
5 to knowGemma 3n is open-source and designed for on-device deployment
Multimodal capability includes audio processing
2-in-1 model architecture provides flexibility
Optimized for fast inference on mobile hardware
Positioned as alternative to cloud-dependent models for latency-sensitive applications
Go to the source
Google DeepMind Blogdeepmind.google
Publisher excerpt: Gemma 3n is a cutting-edge open model designed for fast, multimodal AI on devices, featuring optimized performance, unique flexibility with a 2-in-1 model, and expanded multimodal understanding with audio, empowering developers to build live, interactive applications and sophisticated audio-centric…