FrontierThe story, in brief

OpenAI is now using AI to attack its own AI, and it's working better than humans ever did

84%. That's how often OpenAI's GPT-Red finds vulnerabilities in its own models—6x better than human red teamers. The future of AI safety isn't human oversight. It's AI attacking AI.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI has shifted from human-led red teaming to AI-powered adversarial testing, dramatically improving vulnerability detection rates. This signals a fundamental change in how frontier labs approach model safety and hardening—moving from manual processes to automated, scalable attack discovery that feeds directly into production model improvements.

The key facts

5 to know
  1. GPT-Red achieves 84% success rate finding attacks in test scenarios

  2. Human red teamers achieve 13% success rate by comparison

  3. Self-play training methodology used for vulnerability discovery

  4. Results integrated into GPT-5.6 Sol hardening pipeline

  5. Represents 6.5x improvement over human red teaming baseline

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training. Human red teamers manage just 13 percent. The results feed directly into hardening models like GPT-5.6 Sol.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier