New math benchmark reveals AI models confidently solve problems that have no solution
No model breaks 50% on spotting unsolvable problems. Gemini 3 Pro leads at 30% on research math—but that's not the story.

Why it matters
A new 439-task benchmark exposes a critical gap in AI reasoning: models scale compute to solve harder problems, but gain zero ability to recognize when problems have no solution. This reveals a fundamental limitation in current scaling approaches that matters for deployment in research and professional settings.
The key facts
6 to knowSOOHAK benchmark: 439 handwritten tasks built by 64 mathematicians
99 tasks deliberately unsolvable
Google Gemini 3 Pro: 30% on research-level problems
No model exceeds 50% accuracy on identifying unsolvable tasks
Increased compute improves problem-solving but not error detection
Benchmark designed to measure gap between flashy benchmarks and broad research capability
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: A consortium of 64 mathematicians built SOOHAK, a new AI benchmark with 439 handwritten tasks, including 99 that are deliberately unsolvable. Google's Gemini 3 Pro leads on research-level problems at 30 percent. But no model cracks 50 percent on spotting broken tasks. More compute makes models…