FrontierThe story, in brief

From hard refusals to safe-completions: toward output-centric safety training

GPT-5 ditches hard refusals. OpenAI's new 'safe-completions' approach trains models to answer risky questions—safely.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI is redefining AI safety training away from blocking dangerous requests toward nuanced output control. This shifts how enterprises balance security with usability, and sets a new standard for handling dual-use prompts that competitors will need to match.

The key facts

5 to know
  1. GPT-5 introduces 'safe-completions' safety training methodology

  2. Moves from hard refusals to output-centric safety approach

  3. Improves both safety and helpfulness metrics

  4. Designed to handle dual-use prompts with nuance

  5. Published August 7, 2025 on OpenAI official channel

Go to the source

OpenAI Blogopenai.com

Publisher excerpt: Discover how OpenAI's new safe-completions approach in GPT-5 improves both safety and helpfulness in AI responses—moving beyond hard refusals to nuanced, output-centric safety training for handling dual-use prompts.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier