Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock
OpenAI's GPT-5.6 models now available on Bedrock with explicit prompt caching — a direct cost lever for teams running repeated inference on large contexts.

Why it matters
Practitioners running GPT workloads on AWS can now reduce inference costs via fine-grained prompt caching control, shifting from model choice to operational optimization. This expands OpenAI's reach beyond first-party platforms and makes token economics more competitive.
The key facts
9 to knowOpenAI GPT-5.6 Sol, Terra, Luna now GA on Amazon Bedrock
Explicit prompt caching feature ships with the models
Enables selective caching of prompt sections for cost reduction
Migration path from existing GPT deployments
Published July 2026
OpenAI GPT-5.6 Sol, Terra, and Luna now generally available on Amazon Bedrock
Explicit prompt caching feature enables precise control over which prompt sections are cached and reused
Cost reduction and inference optimization for existing GPT workloads
Migration path for enterprises already using GPT on Bedrock
Go to the source
AWS Machine Learning Blogaws.amazon.com
Publisher excerpt: OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock, along with explicit prompt caching that gives you precise control over which parts of your prompt are cached and reused. Learn how to get started, set up explicit caching, and migrate existing GPT workloads to reduce…
