WorkThe story, in brief

AI models would rather guess than ask for help, researchers find

22 AI models tested. Almost none ask for help when they don't know. Here's why that matters for deployment.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Academic research reveals a critical behavioral gap in multimodal AI models—they hallucinate rather than seek clarification—with implications for real-world reliability and user trust in production systems.

The key facts

10 to know
  1. ProactiveBench benchmark tests whether multimodal models request help when visual information is missing

  2. 22 models tested; almost none exhibit help-seeking behavior

  3. Models default to guessing/hallucination over asking clarifying questions

  4. Reinforcement learning approach shown as potential remediation pathway

  5. Addresses fundamental safety/reliability concern for deployment-ready AI systems

  6. ProactiveBench benchmark tests 22 multimodal language models

  7. Almost no models ask for help when visual information is missing

  8. Models default to guessing/hallucination over clarification requests

  9. Reinforcement learning approach shows potential fix

  10. Published April 2026

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: ProactiveBench tests whether multimodal language models ask users for help when visual information is missing. Out of 22 models tested, almost none ask for what they need, but a simple reinforcement learning approach hints at a fix.
Read original report
Back to today's editionMore work news

The wider picture

View all
Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI illustration by KeyNews
Work01

The Emerging M&A Map For AI Agent Security

As agents move from pilots to production with real system access, enterprise security models are breaking. The M&A map is forming around who controls agent permissions, monitoring, and governance — a new class of identity management problem that practitioners need to architect for now.

Crunchbase News
Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
AI illustration by KeyNews
Work02

AI privacy budgets: Ask for the calculation, not the claim

Enterprise AI buyers are accepting privacy budget numbers without verification. This deep dive explains what questions to ask vendors about federated learning privacy claims, and why the gap between contractual promises and operational evidence is where real exposure lives.

CIO
Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
AI illustration by KeyNews
Work03

Andrew Kelley Interview: Why He Built Zig, Banned AI Contributions, and Moved Zig off GitHub

Open-source governance is shifting in response to AI-generated contributions. Zig's formal ban and migration off GitHub signals broader industry concern about code quality, maintainer burden, and the cultural impact of automated submissions — a flashpoint for how AI changes the work of software development.

InfoQ AI/ML