Introducing MentalHealthBench
OpenAI releases MentalHealthBench, the first expert-informed eval for AI safety in mental health — a benchmark that matters as agents enter healthcare.

Why it matters
As AI agents move into high-stakes domains like mental health support, rigorous domain-specific benchmarks become critical infrastructure. MentalHealthBench sets a template for evaluating both capability AND safety in sensitive conversations — a frontier labs concern as models scale into regulated spaces.
The key facts
5 to knowExpert-informed benchmark design for mental health conversations
Evaluates both helpfulness and safety in AI responses
Addresses realistic mental health use cases
Published by OpenAI as part of frontier safety/eval work
Represents growing attention to domain-specific benchmarking beyond general-purpose metrics
The story so far
Earlier coverage of this storyline
- AI safety conversations have gotten unbelievableTechCrunch AI
- Nvidia CEO Jensen Huang emerges as Trump's top ally in AI safety debateCNBC Technology
- This story
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.