OpenAI reportedly cut response costs for guest ChatGPT users by more than half
More than half. That's how much OpenAI just cut inference costs for ChatGPT—and it signals a major shift in the GPU economics of LLM serving.

Why it matters
OpenAI's ability to slash inference costs by 50%+ through optimization, not just scale, demonstrates that the real competitive moat in AI is now operational efficiency and infrastructure optimization. This directly impacts unit economics for every AI company and signals a shift away from raw compute arms races.
The key facts
4 to knowInference costs cut by more than 50%
GPU requirement for ChatGPT dropped to a few hundred units at peak times
Optimizations applied to guest ChatGPT users first
Source: The Information (June 2026)
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: According to a report by The Information, OpenAI has cut inference costs for its AI models by more than half. The company applied the optimizations to ChatGPT, where the number of Nvidia GPUs needed dropped to just a few hundred at times.