OpenAI · 2026-09-23 · major
MentalHealthBench — OpenAI's open test of AI in mental health chats
OpenAI released MentalHealthBench on 23 September 2026: 1,215 synthetic mental-health conversations with 5,262 rubric criteria written by 80+ licensed experts from 22 countries. GPT-6 Astra scores 57.3%, Claude Opus 5.5 52.4%.

An open benchmark, written with 80+ psychologists and psychiatrists, that grades how well AI responds to people in distress.
Quick facts
| Maker | OpenAI |
|---|---|
| Size | 1,215 conversations, 5,262 rubric criteria |
| Experts | 80+ psychologists and psychiatrists, 22 countries |
| Languages | 19 |
| Grader | GPT-5.6 Sol (high reasoning effort) |
| Access | Paper + dataset download from OpenAI |
Benchmarks
| GPT-6 Astra | 57.3% | |
|---|---|---|
| GPT-6 Sol | 53.9% | |
| Claude Opus 5.5 | 52.4% | |
| GPT-6 Luna | 50.2% |
What is it?
MentalHealthBench is an open benchmark from OpenAI for judging AI replies in realistic mental health conversations. It holds 1,215 synthetic conversations and 5,262 rubric criteria, co-written with more than 80 licensed psychologists and psychiatrists from 22 countries in 19 languages. The conversations come from four kinds of users: adults, teens, caregivers and clinicians.
How does it work?
Each expert read a conversation and wrote criteria for the model's next reply, each checking one thing — asking the right question, calibrating urgency, preserving the user's agency. Criteria carry weights from -10 to +10, so harmful behaviour costs points. A grader model, GPT-5.6 Sol at high reasoning effort, scores each reply, and the total splits into ten expert-defined dimensions.
Why does it matter?
The benchmark goes beyond crisis replies: it covers everyday, non-acute talks (53.5% of cases) alongside high-acuity (18.2%) and emergency (28.3%) ones. OpenAI's own results put every frontier model below 60%, so the test leaves clear room to improve.
Who is it for?
AI safety researchers and teams building health or companion apps
Frequently asked questions
- Which AI model scores highest on MentalHealthBench?
- In OpenAI's published results for MentalHealthBench, GPT-6 Astra leads with a task-clipped score of 57.3%, ahead of GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4% and GPT-6 Luna at 50.2%. OpenAI built and ran the benchmark itself, and a GPT-5.6 Sol model does the grading, so independent re-runs are worth watching.
- Can I download MentalHealthBench and run it myself?
- OpenAI released MentalHealthBench openly so researchers can inspect the method and run their own evaluations. The dataset is a download from OpenAI's CDN alongside the paper. The examples carry a canary string so they can be kept out of training data, and OpenAI asks people not to post the plain-text examples publicly.
- What are the limits of MentalHealthBench?
- MentalHealthBench is effectively single-turn: it grades only the model's final reply to a synthetic conversation, not a real patient exchange. Scoring relies on a single grader model, it covers chat only, and the paper calls its comparisons across languages descriptive. User-written rubrics were collected only for non-acute conversations.