GPT-5 System Card
GPT-5 isn't one model anymore. OpenAI just split it into three: main, thinking, and nano. Here's what that means for your inference costs.

Why it matters
OpenAI has architected GPT-5 as a multi-tier routing system, enabling developers to optimize latency and cost by selecting appropriate model variants for different task complexities—a strategic shift toward efficient model deployment at scale.
The key facts
5 to knowGPT-5 split into three variants: gpt-5-main, gpt-5-thinking, gpt-5-thinking-nano
Unified model routing system for task-specific optimization
Lightweight versions available for cost-sensitive workloads
Published August 7, 2025 on OpenAI official channel
Focus on latency and performance differentiation across model tiers
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: This GPT-5 system card explains how a unified model routing system powers fast and smart responses using gpt-5-main, gpt-5-thinking, and lightweight versions like gpt-5-thinking-nano, optimized for different tasks and developer use.