OpenAI Ultrafast mode now available on AI Gateway
OpenAI's Ultrafast tier now live on AI Gateway: 6× per-token cost for sub-100ms latency on GPT-6 Astra.

Why it matters
OpenAI is pricing interactive speed as a premium service tier. Practitioners building real-time apps (coding, chat, tool calls) can opt into Ultrafast; cost and regional limits (EU unsupported) will shape adoption.
The key facts
11 to knowUltrafast service tier now available for GPT-6 Astra via AI Gateway
Billed at 6× standard per-token rate
Supported in US and global regions; EU requests fall back to standard tier
Recommended for workflows with frequent tool calls using Responses API over WebSocket
Published October 5, 2026 on Vercel changelog
Ultrafast service tier billed at 6× standard per-token rate
Supported in US and global regions only; EU requests revert to standard tier
Accessible via AI SDK, Chat Completions API, or Responses API
OpenAI recommends Responses API over WebSocket for frequent tool calls to reduce overhead
Standard tier remains default when no service tier specified
GPT-6 Astra model; published October 5, 2026
The story so far
Earlier coverage of this storyline
- Microsoft AI models are now available on AI GatewayVercel Blog
- This story
Go to the source
Vercel Blogvercel.com
Publisher excerpt: now supports OpenAI's Ultrafast service tier for GPT-6 Astra, providing faster output for interactive applications and rapid coding iterations.AI Gateway To use Ultrafast, request it for through AI SDK, the Chat Completions API, or Responses API:openai/gpt-6-astra For workflows with frequent tool…