cyberivy
MentalHealthBenchOpenAIAI SafetyMental Health AIChatbot SafetyAI BenchmarksGPT-6 AstraClinical AI

MentalHealthBench exposes large gaps in AI mental-health advice

September 24, 2026

Eine stilisierte menschliche Kopfsilhouette mit farbigen Gehirnstrukturen neben einem digitalen KI-Symbol

OpenAI evaluates AI responses across 1,215 mental-health scenarios using expert-written criteria. Even the best tested model reaches only 57.3 percent.

What this is about

OpenAI released MentalHealthBench on September 23, 2026. The open benchmark tests how language models respond to realistic conversations about stress, grief, mental illness, and acute crises. It contains 1,215 synthetic conversations and 5,262 evaluation criteria developed by more than 80 licensed psychologists and psychiatrists across 22 countries.

The most striking result is not one model's victory but the large gap to dependable support: the top-ranked GPT-6 Astra reaches 57.3 percent in OpenAI's evaluation. Claude Opus 5.5 reaches 52.4 percent. OpenAI itself stresses that ChatGPT is not a substitute for therapy or professional care.

What MentalHealthBench actually does

The dataset covers three levels of acuity. Non-acute situations make up 53.5 percent of conversations, high-acuity distress 18.2 percent, and emergencies 28.3 percent. Adults, teenagers, caregivers, and clinicians are represented as user groups. Beyond English, the benchmark includes conversations in Spanish, Hindi, Arabic, Portuguese, German, and other languages.

At least three experts reviewed every conversation. Their criteria reward helpful behavior and penalize risky responses with weights from minus ten to plus ten. They assess whether a model seeks enough context, calibrates urgency, preserves user agency, and offers actionable next steps. An automated grader then compares model responses with those criteria.

Why it matters

People already discuss deeply personal crises with chatbots. That is why testing only obvious emergencies is not enough. A response can sound polite and empathetic while still jumping to conclusions, omitting crucial questions, or recommending real-world help too late.

MentalHealthBench makes these differences easier to measure. At the same time, the study shows why a leaderboard must not be confused with clinical safety. Even expert-written, short conversational answers scored only 38.5 percent because the grading system rewards comprehensive responses. That reveals a limitation of the measurement method; it does not mean machines are better therapists.

In plain language

The benchmark is like a driving test with 1,215 different traffic situations. A car does not pass merely because it stops at a red light; it must also handle blind junctions, pedestrians, and bad weather. A score of 57.3 percent therefore does not mean the system is suitable for every second crisis. It shows how many expected behaviors it demonstrated in the test.

A practical example

Imagine a 16-year-old writes that she has barely slept for days, feels worthless, and does not want to burden anyone. A weak answer offers generic relaxation tips. A stronger answer takes the statements seriously, gently checks for immediate danger, encourages contact with a trusted person, and points to appropriate local help without claiming a diagnosis.

The benchmark breaks that response into individual criteria. The model can gain points for an appropriate safety question and concrete support, but lose points if it becomes controlling or infers a disorder from a few sentences. These fine distinctions are what the benchmark aims to expose to developers.

Scope and limits

  • The conversations are synthetic. This protects privacy but only approximates messy, contradictory real-life situations.
  • OpenAI developed the benchmark and used its own automated grader to evaluate models. Independent replications therefore matter.
  • An aggregate score is not medical approval. Language, culture, product safeguards, and the specific crisis can substantially change outcomes.

MentalHealthBench is therefore a useful testing tool for developers and researchers. For people in distress, real professional support remains essential, especially when there is immediate danger.

SEO & GEO keywords

MentalHealthBench, OpenAI, AI mental health, chatbot safety, Mental Health AI, GPT-6 Astra, Claude Opus 5.5, AI benchmark, crisis intervention, clinical evaluation

πŸ’‘ In plain English

MentalHealthBench tests whether AI responds safely and usefully in conversations about mental distress. The best tested model meets only a little over half of the expected criteria, and the test is not a medical certification.

Key Takeaways

  • β†’The open benchmark contains 1,215 synthetic conversations and 5,262 expert-authored criteria.
  • β†’More than 80 licensed experts across 22 countries contributed to the evaluations.
  • β†’GPT-6 Astra scores 57.3 percent in OpenAI's evaluation, while Claude Opus 5.5 scores 52.4 percent.
  • β†’The dataset covers everyday situations, high-acuity distress, and mental-health emergencies.
  • β†’The results are not medical approval and require independent scrutiny.

FAQ

What is MentalHealthBench?

It is an open benchmark that evaluates AI responses to 1,215 synthetic mental-health conversations using expert-written criteria.

Which model performed best?

GPT-6 Astra reached 57.3 percent in the published evaluation. That is a benchmark score, not evidence of clinical effectiveness.

Can a chatbot replace therapy?

No. OpenAI explicitly states that ChatGPT is not a substitute for therapy or professional care.

What is the biggest limitation?

The conversations are synthetic and evaluation uses an automated OpenAI grader. Independent replications are needed to test robustness.

Sources & Context