← All posts / Research

1,215 Conversations, 80 Clinicians: OpenAI Open-Sources MentalHealthBench

OpenAI's new open benchmark scores frontier models on the full spectrum of mental health conversations — from everyday stress to psychiatric emergencies — using rubrics written by 80+ licensed clinicians across 20+ countries.

1,215 Conversations, 80 Clinicians: OpenAI Open-Sources MentalHealthBench

On September 23, 2026, OpenAI released MentalHealthBench, an open benchmark of 1,215 realistic mental health conversations designed to measure how well large language models respond when people bring them their hardest moments — not just acute emergencies, but the full spectrum from everyday stress to psychiatric crisis. It is the company’s most ambitious attempt yet to turn one of AI’s most sensitive failure surfaces into something measurable, reproducible, and public.

The timing is pointed. More than a billion people interact with ChatGPT every week, and mental health has quietly become one of its most common high-stakes use cases: emotional support, life advice, and help during moments of crisis. Regulators, researchers, and grieving families have spent two years asking whether chatbots are safe in these conversations. OpenAI’s answer is a benchmark that treats the question as an empirical one — and publishes the results for its competitors, not only its own models.

What’s in the Benchmark

MentalHealthBench is a dataset of 1,215 synthetic conversations built to mirror the real distribution of mental health usage on ChatGPT, which OpenAI mapped using privacy-preserving techniques similar to its Clio pipeline. Each conversation ends on a user turn, and the model under test must produce the next assistant response — which is then graded against a rubric written specifically for that conversation.

The composition is deliberately broad:

  • Acuity: 53.5% non-acute (everyday emotional valence — venting, relationship advice, reassurance seeking), 18.2% high-acuity (meaningfully significant symptoms like moderate-to-severe depression or intense panic), and 28.3% emergent (psychiatric emergencies involving risk of harm to self or others, grave disability, or medical emergency from substance use).
  • User profiles: 68.1% adults, 21.2% teenagers (13–17), 5.8% clinicians asking practice questions, and 4.9% caregivers supporting a loved one.
  • Themes: 19 categories led by romantic relationships, suicide and self-harm, faith and belief systems, psychosis and altered reality, and medication uncertainty.
  • Length: 53.7% of conversations run longer than five messages, reflecting how these talks actually unfold — gradually, over time.
  • Prior context: 70 tasks (5.8%) test whether models adapt when prior context about the user is known, such as a recent death in the family surfacing in a later conversation about grief.
  • Languages: a non-English subset spanning Spanish (105 conversations), Hindi (54), Arabic (34), Portuguese (29), German (25), Italian and Persian (17 each), Indonesian (17), Turkish (13), and Chinese (1), annotated by fluent clinicians with direct insight into those cultural contexts.

The rubrics are the benchmark’s core innovation. Each conversation was independently annotated by two blinded licensed clinicians, then adjudicated by a third who had authored neither initial rubric. Criteria carry importance weights from −10 to +10, rewarding beneficial behaviors and penalizing harmful ones. The final rubric for every task reflects the consensus of at least two, and often three, experts. The cohort behind this work: more than 80 licensed psychiatrists and psychologists from over 20 countries, collectively speaking 19 languages and spanning nearly 20 subspecialties — addiction, forensic psychology, trauma, sexual health, suicide and self-harm, psychotic disorders, neurodevelopmental disorders, and more. A subset of 30 clinicians with current or recent adolescent patients handled all teen examples.

OpenAI also collected an expert-authored reference completion for each conversation — a fourth clinician, uninvolved in the rubric, wrote the response they would most want a safe AI system to give.

The Results

On the overall task-clipped score, a clear hierarchy emerges:

ModelTask-clipped score
GPT-6 Astra57.3
GPT-6 Sol53.9
Claude Opus 5.552.4
GPT-6 Luna50.2
Muse Spark 1.348.6
GPT-5.6 Sol (Aug 2026)47.0
Claude Fable 5.146.4
GPT-5.6 Luna (Aug 2026)44.9
Claude Sonnet 544.5
GPT-5 Thinking42.9
Claude Haiku 4.541.7
Grok 4.7 / Gemini 2.5 Flash32.1
Gemini 3.1 Pro32.1
Gemini 2.5 Pro35.5
GPT-4o (March 2025)41.3
Gemini 3.8 Flash46.4

Two findings stand out beyond the leaderboard. First, the score decomposition shows some models earn plenty of positive points but lose heavily to negative penalties — racking up credit for empathy while still committing rubric-flagged harmful behaviors. Second, performance varies sharply by acuity: GPT-6 Sol actually leads the non-acute band at 57.9 while GPT-6 Astra leads emergent conversations, and the older GPT-4o scores a striking 54.3 on non-acute but collapses to 31.7 on high-acuity and 27.8 on emergent — evidence that everyday conversational skill and crisis competence are genuinely different capabilities.

The paper’s central diagnosis is that models remain weakest at seeking appropriate context and calibrating urgency — knowing when a situation is serious enough to escalate, and when escalation itself would be the harm. Its taxonomy of ten behavioral axes (context/assessment 66.5% of rubric weight, actionable guidance 52.5%, empathy/support 39.1%, clinical accuracy 36.1%, urgency calibration 17.9%, reality testing 14.9%) makes those failure modes individually visible for the first time. Older models, for example, perform poorly on reality-testing — validating users’ unsupported beliefs — a failure mode long associated with sycophancy.

Users vs. Experts

The most unusual section of the paper compares expert rubrics with rubrics written by a separate cohort of 40+ real ChatGPT users from 16 countries. The verdict: user guidance is a coherent, complementary signal but not a substitute for expert guidance. The two cohorts align on only 25.7% of rubric weight; 39.1% of weight is expert-only (mostly added guardrails), 34.2% user-only (mostly actionability and tone preferences), and just 1.0% directly contradictory. Where experts penalize assuming another person’s motives, users reward a possible explanation for a partner’s texting behavior — a small, precise illustration of why optimizing purely for user satisfaction would drift from clinical safety. Notably, the user cohort only annotated non-acute conversations, for ethical reasons.

Why It Matters

Benchmarks shape incentives. For years, mental health evaluation meant crisis-refusal tests — could the model avoid handing out a method or missing a suicidal ideation flag. That bar is necessary but radically insufficient: most mental health conversations are not emergencies, and a response can be technically safe yet generic, impersonal, or poorly attuned to the user’s circumstances. By spanning the full acuity spectrum and grading against clinician consensus rather than binary safety checks, MentalHealthBench gives every lab — not just OpenAI — a public target to train and evaluate against. Its multilingual design also acknowledges an under-examined risk: models may misconstrue ordinary spiritual practices as pathology or give advice that ignores a user’s cultural norms.

The limitations are stated plainly. Conversations are synthetic, so they approximate rather than reproduce real usage. Language comparisons are descriptive and can’t isolate language effects from acuity, topic, or culture. The benchmark is single-turn at its core, despite multiturn aspects. And the hardest questions — how to incorporate lived experience into high-acuity evaluation, or extend grading to voice, video, and agentic surfaces — are left explicitly open.

What OpenAI has built, in effect, is the mental-health counterpart to HealthBench: a measurement instrument designed to make “safe and helpful” falsifiable. The initial scores suggest the frontier is real but immature — the best model in the world gets 57.3% on conversations clinicians consider routine judgment. For an industry that positions AI as a complement to overstretched mental health care, that gap is now the metric that matters.