█

AI/TLDR

OpenAI · 2026-09-23 · major

MentalHealthBench — OpenAI's open test of AI in mental health chats

OpenAI released MentalHealthBench on 23 September 2026: 1,215 synthetic mental-health conversations with 5,262 rubric criteria written by 80+ licensed experts from 22 countries. GPT-6 Astra scores 57.3%, Claude Opus 5.5 52.4%.

Unite.AI header illustration for its report on OpenAI's MentalHealthBench
Unite.AI

An open benchmark, written with 80+ psychologists and psychiatrists, that grades how well AI responds to people in distress.

Quick facts

MakerOpenAI
Size1,215 conversations, 5,262 rubric criteria
Experts80+ psychologists and psychiatrists, 22 countries
Languages19
GraderGPT-5.6 Sol (high reasoning effort)
AccessPaper + dataset download from OpenAI

Benchmarks

MentalHealthBench (task-clipped)
GPT-6 Astra57.3%
GPT-6 Sol53.9%
Claude Opus 5.552.4%
GPT-6 Luna50.2%
source ↗

What is it?

MentalHealthBench is an open benchmark from OpenAI for judging AI replies in realistic mental health conversations. It holds 1,215 synthetic conversations and 5,262 rubric criteria, co-written with more than 80 licensed psychologists and psychiatrists from 22 countries in 19 languages. The conversations come from four kinds of users: adults, teens, caregivers and clinicians.

How does it work?

Each expert read a conversation and wrote criteria for the model's next reply, each checking one thing — asking the right question, calibrating urgency, preserving the user's agency. Criteria carry weights from -10 to +10, so harmful behaviour costs points. A grader model, GPT-5.6 Sol at high reasoning effort, scores each reply, and the total splits into ten expert-defined dimensions.

Why does it matter?

The benchmark goes beyond crisis replies: it covers everyday, non-acute talks (53.5% of cases) alongside high-acuity (18.2%) and emergency (28.3%) ones. OpenAI's own results put every frontier model below 60%, so the test leaves clear room to improve.

Who is it for?

AI safety researchers and teams building health or companion apps

Frequently asked questions

Which AI model scores highest on MentalHealthBench?
In OpenAI's published results for MentalHealthBench, GPT-6 Astra leads with a task-clipped score of 57.3%, ahead of GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4% and GPT-6 Luna at 50.2%. OpenAI built and ran the benchmark itself, and a GPT-5.6 Sol model does the grading, so independent re-runs are worth watching.
Can I download MentalHealthBench and run it myself?
OpenAI released MentalHealthBench openly so researchers can inspect the method and run their own evaluations. The dataset is a download from OpenAI's CDN alongside the paper. The examples carry a canary string so they can be kept out of training data, and OpenAI asks people not to post the plain-text examples publicly.
What are the limits of MentalHealthBench?
MentalHealthBench is effectively single-turn: it grades only the model's final reply to a synthetic conversation, not a real patient exchange. Scoring relies on a single grader model, it covers chat only, and the paper calls its comparisons across languages descriptive. User-written rubrics were collected only for non-acute conversations.

Sources · 3 outlets

Tags

  • openai
  • mentalhealthbench
  • benchmark
  • evals
  • mental-health
  • healthcare
  • clinical-ai
  • ai-safety
  • rubric-grading
  • gpt-6-astra
  • claude-opus-5-5

← All releases · Learn AI