Leading AI chatbots avoid harm but fall short in high-risk conversations, startup’s new benchmark finds
Leading AI chatbots fail at the conversations that matter most. A new benchmark exposes why.

Why it matters
As AI moves into high-stakes domains like mental health support, new safety benchmarks reveal critical gaps in how leading models handle suicide risk, eating disorders, and health misinformation—forcing a reckoning on where guardrails actually work.
The key facts
5 to knowmpathic (Seattle startup) released mPACT benchmark
Evaluates AI model performance on suicide risk, eating disorders, misinformation conversations
Clinician-led benchmark design
Leading models show harm-avoidance but performance gaps in high-risk scenarios
Benchmark focuses on deployment readiness in mental health and clinical contexts
Go to the source
GeekWiregeekwire.com
Publisher excerpt: Seattle-based mpathic released mPACT, a clinician-led benchmark that evaluates how leading AI models handle conversations involving suicide risk, eating disorders, and misinformation.

