ToolsThe story, in brief

New ways to balance cost and reliability in the Gemini API

Google just added two new inference tiers to Gemini API. Cost vs speed is now a dial, not a compromise.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Google's new Flex and Priority tiers give developers granular control over the cost-latency tradeoff, potentially shifting competitive dynamics in the API market where pricing optimization is becoming critical for enterprise adoption.

The key facts

3 to know
  1. Two new inference tiers: Flex and Priority

  2. Focus on balancing cost and latency

  3. Gemini API enhancement

Go to the source

Google AI Blogblog.google

Publisher excerpt: Google is introducing two new inference tiers to the Gemini API, Flex and Priority, to balance cost and latency.
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools