FrontierThe story, in brief

New math benchmark reveals AI models confidently solve problems that have no solution

No model breaks 50% on spotting unsolvable problems. Gemini 3 Pro leads at 30% on research math—but that's not the story.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A new 439-task benchmark exposes a critical gap in AI reasoning: models scale compute to solve harder problems, but gain zero ability to recognize when problems have no solution. This reveals a fundamental limitation in current scaling approaches that matters for deployment in research and professional settings.

The key facts

6 to know
  1. SOOHAK benchmark: 439 handwritten tasks built by 64 mathematicians

  2. 99 tasks deliberately unsolvable

  3. Google Gemini 3 Pro: 30% on research-level problems

  4. No model exceeds 50% accuracy on identifying unsolvable tasks

  5. Increased compute improves problem-solving but not error detection

  6. Benchmark designed to measure gap between flashy benchmarks and broad research capability

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: A consortium of 64 mathematicians built SOOHAK, a new AI benchmark with 439 handwritten tasks, including 99 that are deliberately unsolvable. Google's Gemini 3 Pro leads on research-level problems at 30 percent. But no model cracks 50 percent on spotting broken tasks. More compute makes models…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier