Gemini 3.1 Flash-Lite: Built for intelligence at scale
Google just shipped Gemini 3.1 Flash-Lite. Faster. Cheaper. The race for inference efficiency just got real.

Why it matters
Google is doubling down on cost-efficient inference with a new lightweight variant positioned to compete in the speed-vs-intelligence tradeoff that's reshaping model economics for deployed applications.
The key facts
4 to knowGemini 3.1 Flash-Lite launched as fastest and most cost-efficient Gemini 3 series model
Positioned for 'intelligence at scale' deployment
Part of Google's multi-tier model strategy (Flash-Lite joining standard Flash and Ultra variants)
Published March 3, 2026 via DeepMind official blog
Go to the source
Google DeepMind Blogdeepmind.google
Publisher excerpt: Gemini 3.1 Flash-Lite is our fastest and most cost-efficient Gemini 3 series model yet.