ToolsThe story, in brief

How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast

8x faster. OpenAI's GPT-6 Astra Ultrafast, running on Blackwell, is live in the API and ChatGPT Work.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI ships a performance-optimized inference tier leveraging NVIDIA Blackwell architecture. Practitioners using the OpenAI API now have a latency option for cost-sensitive or real-time workloads, though the trade-off (quality vs. speed) is not detailed.

The key facts

10 to know
  1. GPT-6 Astra Ultrafast available now in OpenAI API

  2. Available to eligible ChatGPT Work and Codex users

  3. Runs on NVIDIA Blackwell GPUs

  4. Up to 8x faster token generation vs. Astra Standard mode

  5. Uses inference optimizations tied to Blackwell architecture

  6. Quality/accuracy impact of speed tier not disclosed in article

  7. Pricing not disclosed

  8. GPT-6 Astra Ultrafast available now in OpenAI API and ChatGPT Work/Codex

  9. Inference optimization, not model change

  10. Vendor blog; no independent performance verification provided

Go to the source

NVIDIA Blogblogs.nvidia.com

Publisher excerpt: GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools