ToolsThe story, in brief

OpenAI Ultrafast mode now available on AI Gateway

OpenAI's Ultrafast tier now live on AI Gateway: 6× per-token cost for sub-100ms latency on GPT-6 Astra.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI is pricing interactive speed as a premium service tier. Practitioners building real-time apps (coding, chat, tool calls) can opt into Ultrafast; cost and regional limits (EU unsupported) will shape adoption.

The key facts

11 to know
  1. Ultrafast service tier now available for GPT-6 Astra via AI Gateway

  2. Billed at 6× standard per-token rate

  3. Supported in US and global regions; EU requests fall back to standard tier

  4. Recommended for workflows with frequent tool calls using Responses API over WebSocket

  5. Published October 5, 2026 on Vercel changelog

  6. Ultrafast service tier billed at 6× standard per-token rate

  7. Supported in US and global regions only; EU requests revert to standard tier

  8. Accessible via AI SDK, Chat Completions API, or Responses API

  9. OpenAI recommends Responses API over WebSocket for frequent tool calls to reduce overhead

  10. Standard tier remains default when no service tier specified

  11. GPT-6 Astra model; published October 5, 2026

The story so far

Earlier coverage of this storyline

  1. Microsoft AI models are now available on AI GatewayVercel Blog
  2. This story

Go to the source

Vercel Blogvercel.com

Publisher excerpt: now supports OpenAI's Ultrafast service tier for GPT-6 Astra, providing faster output for interactive applications and rapid coding iterations.AI Gateway To use Ultrafast, request it for through AI SDK, the Chat Completions API, or Responses API:openai/gpt-6-astra For workflows with frequent tool…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools