FrontierThe story, in brief

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

OpenAI's new GPT-Red isn't a product—it's a red-team LLM that made GPT-5.6 its most robust release yet.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI is using adversarial LLM training (GPT-Red as sparring partner) as a core safety methodology for its flagship models. This signals a shift from external red-teaming to internal, automated robustness testing—a capability differentiation that affects how companies should evaluate model safety claims.

The key facts

4 to know
  1. GPT-Red: adversarial LLM built by OpenAI for internal red-teaming

  2. GPT-5.6 released as 'most robust release yet' after training against GPT-Red

  3. Automated cyberattack simulation embedded in training loop

  4. Published: July 15, 2026 by MIT Technology Review

Go to the source

MIT Technology Reviewtechnologyreview.com

Publisher excerpt: OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red made the model its…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier