New benchmark confirms AI models still perform poorly at visual perception
No frontier model cracks 60%. Moonshot's PerceptionBench reveals visual perception—not reasoning—is the real bottleneck.

Why it matters
A rigorous capability benchmark isolates where multimodal models actually fail: not in logic, but in basic image understanding. This reshapes how practitioners should think about vision-language tradeoffs and where to invest eval effort.
The key facts
5 to knowMoonshot AI released PerceptionBench benchmark
No frontier model reaches 60% accuracy on visual perception
GPT-5.6 Sol leads by narrow margin
Isolates visual perception from reasoning—finding errors occur at image-reading stage, not downstream logic
Multimodal AI weakness is foundational, not reasoning-layer
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage.