How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
8x faster. OpenAI's GPT-6 Astra Ultrafast, running on Blackwell, is live in the API and ChatGPT Work.

Why it matters
OpenAI ships a performance-optimized inference tier leveraging NVIDIA Blackwell architecture. Practitioners using the OpenAI API now have a latency option for cost-sensitive or real-time workloads, though the trade-off (quality vs. speed) is not detailed.
The key facts
10 to knowGPT-6 Astra Ultrafast available now in OpenAI API
Available to eligible ChatGPT Work and Codex users
Runs on NVIDIA Blackwell GPUs
Up to 8x faster token generation vs. Astra Standard mode
Uses inference optimizations tied to Blackwell architecture
Quality/accuracy impact of speed tier not disclosed in article
Pricing not disclosed
GPT-6 Astra Ultrafast available now in OpenAI API and ChatGPT Work/Codex
Inference optimization, not model change
Vendor blog; no independent performance verification provided
Go to the source
NVIDIA Blogblogs.nvidia.com
Publisher excerpt: GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x…