OpenAI details GPT-Red, an AI that attacks its own models to find flaws
OpenAI just automated red teaming. GPT-Red now finds its own vulnerabilities before humans can exploit them.

Why it matters
Red teaming—the expensive, manual process of breaking AI systems to find flaws—is shifting from human security teams to autonomous AI agents. This changes how companies think about AI safety governance, disclosure timelines, and competitive advantage in finding vulnerabilities first.
The key facts
5 to knowGPT-Red is an internal AI system built to autonomously attack OpenAI's own models
Focuses on prompt injection vulnerabilities and security flaws
Replaces manual red teaming work traditionally done by human security teams
Operates at scale and speed beyond human capability
Published July 15, 2026 on SiliconANGLE
Go to the source
SiliconAnglesiliconangle.com
Publisher excerpt: OpenAI Group PBC today detailed GPT-Red, an internal artificial intelligence system it built to attack its own models and surface prompt injection vulnerabilities before they reach users. Red teaming is the job of hammering software to find its weak points, work that normally falls to human…