FrontierSeptember 2, 2026via The Decoder
OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder
Why it matters
A frontier lab is shipping a model with acknowledged dangerous capabilities while simultaneously acknowledging its safety instrumentation (chain-of-thought monitoring) cannot reliably observe what the model is actually doing. This is a live debate about interpretability, capability, and risk management at the frontier.
Key signals
- Astra rated as first OpenAI system with 'critical' cyber capabilities
- Safety monitoring strategy relies on chain-of-thought observation
- Chain-of-thought already considered unreliable as a safety mirror
- Astra's architecture pushes more thinking into 'unreadable' space
- Safety observability degrading as capabilities scale
- Monitoring and capability tension is the core technical story
The hook
OpenAI's Astra: first 'critical' cyber capability model — but the safety monitoring it relies on is already unreliable and getting worse.
OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a re…