Gemini 2.5 Flash-Lite is now ready for scaled production use
Gemini 2.5 Flash-Lite hits GA. 1M context window, multimodal, production-ready at scale.

Why it matters
Google's cost-efficient model moves from preview to production, expanding access to enterprise-grade capabilities at lower inference costs—a direct competitive move against OpenAI's smaller model tiers and Claude's pricing strategy.
The key facts
6 to knowGemini 2.5 Flash-Lite now generally available (GA)
Previously in preview status, now stable for production
1 million-token context window
Multimodal capabilities included
Positioned as cost-efficient vs. larger models
Inherits Gemini 2.5 family features
Go to the source
Google DeepMind Blogdeepmind.google
Publisher excerpt: Gemini 2.5 Flash-Lite, previously in preview, is now stable and generally available. This cost-efficient model provides high quality in a small size, and includes 2.5 family features like a 1 million-token context window and multimodality.