OpenAI introduces MentalHealthBench, a benchmark for AI mental-health conversations
OpenAI has published MentalHealthBench, a new benchmark built to measure how AI systems handle realistic mental health conversations rather than generic safety prompts.
What's new
OpenAI describes the project directly: "MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations."
The benchmark was built with more than 80 licensed mental health experts from 22 countries, spanning 19 languages and close to 20 subspecialties. It scores model behavior on things like safety, seeking context before responding, preserving user agency, and giving actionable guidance, rather than just checking for refusals.
Conversations in the benchmark cover three acuity levels — non-acute, high-acuity, and emergencies — and include adults, teens, caregivers, and clinicians as the people on the other side of the conversation, across multiple languages and regions.
Context
Mental health has become one of the more scrutinized use cases for consumer AI chatbots, with regulators, clinicians, and the press all pushing labs to show their work rather than assert it. A benchmark built with licensed experts across dozens of countries and languages is OpenAI's answer to the criticism that model safety testing in this area has been thin, English-centric, or self-graded without outside input.
It also follows a pattern other labs have used before: publish the eval alongside the claim, so outside researchers can run it against competing models instead of taking a vendor's word for it.
Why it matters
The benchmark's headline result, per OpenAI, is steady improvement across AI systems in handling these conversations, with newer models showing particular gains in appropriately seeking context before giving advice. That is a meaningful shift from earlier complaints that chatbots either refused sensitive conversations outright or answered too confidently without checking on the person's situation first. Because the benchmark is open, it becomes a reference point other labs will likely get measured against, whether or not they choose to publish their own scores on it.
Corroborating sources
- Openai
https://openai.com/index/introducing-mentalhealthbench
“MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.”