Platform WatchJune 30, 2026via The Decoder
OpenAI reportedly cut response costs for guest ChatGPT users by more than half
Why it matters
OpenAI's ability to slash inference costs by 50%+ through optimization, not just scale, demonstrates that the real competitive moat in AI is now operational efficiency and infrastructure optimization. This directly impacts unit economics for every AI company and signals a shift away from raw compute arms races.
Key signals
- Inference costs cut by more than 50%
- GPU requirement for ChatGPT dropped to a few hundred units at peak times
- Optimizations applied to guest ChatGPT users first
- Source: The Information (June 2026)
The hook
More than half. That's how much OpenAI just cut inference costs for ChatGPT—and it signals a major shift in the GPU economics of LLM serving.
According to a report by The Information, OpenAI has cut inference costs for its AI models by more than half. The company applied the optimizations to ChatGPT, where the number of Nvidia GPUs needed dropped to just a few hundred at times.