FrontierSeptember 17, 2026via InfoQ AI/ML
GPT-6 Astra Is the First Model OpenAI Classifies as Critical for Cybersecurity
Why it matters
OpenAI's own safety framework now flags a frontier model as a genuine cybersecurity risk at scale. Practitioners deploying or advising on GPT-6 need to understand the attack surface; the decline in chain-of-thought monitoring compounds the concern.
Key signals
- GPT-6 Astra is first OpenAI model classified as 'Critical' under Preparedness Framework cybersecurity threshold
- Model discovered previously unknown vulnerabilities in browser and OS kernel during expert-led testing
- Model built working exploits for discovered vulnerabilities
- System card reports 'substantial decline in chain-of-thought monitorability'
- Published September 17, 2026 via InfoQ/OpenAI system card
The hook
First model ever. GPT-6 Astra hit OpenAI's 'Critical' cybersecurity threshold — found zero-days in real systems, built working exploits.
OpenAI has classified GPT-6 Astra at the Critical cybersecurity threshold under its Preparedness Framework, a first. In expert-led testing the model found previously unknown vulnerabilities in a browser and an OS kernel and built working exploits. The same system card reports a substantial decline i…