FrontierSeptember 12, 2026via The Decoder

GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks

Why it matters

A new frontier model shows measurable capability gains in embodied AI tasks—spatial reasoning benchmarks matter because they signal whether the next generation can reason about physical work, not just text.

Key signals

  • GPT-6 Astra completed 7/100 tasks on StationeryBench with dual-arm robots
  • MolmoAct2 competitor completed 0/100 on same benchmark
  • Described as 'step change in spatial reasoning' by researchers
  • Benchmark: StationeryBench (robotics/spatial reasoning task)
  • Published Sep 12, 2026

The hook

GPT-6 Astra just lapped the field on spatial reasoning. 7/100 on dual-arm robotics tasks. Competitors: zero.

In a new robotics benchmark, GPT-6 Astra shows major gains in spatial understanding. On StationeryBench, the model completed 7 out of 100 tasks with dual-arm robots, while competitor MolmoAct2 couldn't finish a single one. A researcher calls it a "step change in spatial reasoning."

The week's key stories, every Friday.

ONE BRIEFING · EVERY FRIDAY · FREE

Free. Unsubscribe anytime.

GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks | KeyNews.AI