WorkThe story, in brief

AI chatbots reading X-rays can be dangerously confident even when they're wrong

AI radiologists are confidently wrong. RadLE 2.0 benchmark reveals models fail at knowing when to defer to humans—a critical gap before clinical deployment.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

As AI models enter high-stakes domains like medical imaging, the ability to recognize uncertainty and defer decisions is as important as raw accuracy. This research surfaces a fundamental safety gap that could block enterprise adoption in healthcare.

The key facts

5 to know
  1. RadLE 2.0 benchmark tests AI confidence calibration in radiology

  2. Multiple AI models deliver wrong diagnoses with high confidence

  3. Human radiologists outperform AI on uncertainty recognition

  4. Key issue: AI lacks ability to defer to humans when uncertain

  5. Implication: Critical blocker for autonomous clinical AI deployment

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work