FrontierThe story, in brief

New benchmark confirms AI models still perform poorly at visual perception

No frontier model cracks 60%. Moonshot's PerceptionBench reveals visual perception—not reasoning—is the real bottleneck.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A rigorous capability benchmark isolates where multimodal models actually fail: not in logic, but in basic image understanding. This reshapes how practitioners should think about vision-language tradeoffs and where to invest eval effort.

The key facts

5 to know
  1. Moonshot AI released PerceptionBench benchmark

  2. No frontier model reaches 60% accuracy on visual perception

  3. GPT-5.6 Sol leads by narrow margin

  4. Isolates visual perception from reasoning—finding errors occur at image-reading stage, not downstream logic

  5. Multimodal AI weakness is foundational, not reasoning-layer

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier