ToolsThe story, in brief

Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock

OpenAI's GPT-5.6 models now available on Bedrock with explicit prompt caching — a direct cost lever for teams running repeated inference on large contexts.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Practitioners running GPT workloads on AWS can now reduce inference costs via fine-grained prompt caching control, shifting from model choice to operational optimization. This expands OpenAI's reach beyond first-party platforms and makes token economics more competitive.

The key facts

9 to know
  1. OpenAI GPT-5.6 Sol, Terra, Luna now GA on Amazon Bedrock

  2. Explicit prompt caching feature ships with the models

  3. Enables selective caching of prompt sections for cost reduction

  4. Migration path from existing GPT deployments

  5. Published July 2026

  6. OpenAI GPT-5.6 Sol, Terra, and Luna now generally available on Amazon Bedrock

  7. Explicit prompt caching feature enables precise control over which prompt sections are cached and reused

  8. Cost reduction and inference optimization for existing GPT workloads

  9. Migration path for enterprises already using GPT on Bedrock

Go to the source

AWS Machine Learning Blogaws.amazon.com

Publisher excerpt: OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock, along with explicit prompt caching that gives you precise control over which parts of your prompt are cached and reused. Learn how to get started, set up explicit caching, and migrate existing GPT workloads to reduce…
Read original report
Back to today's editionMore tools news

The wider picture

View all
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools01

OpenAI nabs key Patreon execs ahead of upcoming announcement

OpenAI is building a creator-focused product suite with deep domain expertise (Patreon's co-founder + product + engineering leads). This signals a major new revenue and engagement vector for ChatGPT — and a direct threat to Patreon's existing creator economy.

The Verge AI
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Tools02

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

A practical, deployment-ready layer for a common production pattern — classification over generation — that practitioners can drop into existing LLM stacks immediately. No fine-tuning required.

MarkTechPost
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools03

SpeakON Ships a MagSafe AI Voice Button With Its Own Microphone

A hardware-first approach to voice AI workflow — MagSafe button with onboard mic addresses the real friction in voice-to-text-to-action. Relevant to practitioners building voice UX and to the broader consumer AI tooling wave.

MarkTechPost