OFICIAL OpenAI News

Introducing MentalHealthBench

What happened
Based on OpenAI News · Sep 23, 2026

OpenAI introduced MentalHealthBench, an open benchmark co-created with over 80 licensed mental health experts to evaluate AI responses in realistic mental health conversations across safety, context, and user agency.

Introducing MentalHealthBench
OpenAI News — OpenAI
Key points
·
MentalHealthBench was co-created with over 80 licensed mental health experts from 22 countries to evaluate AI responses in realistic mental health conversations.
·
The benchmark assesses AI responses across safety, context-seeking, user agency, and actionable guidance in scenarios ranging from non-acute to emergency situations.
·
Model responses are graded using expert-written rubrics with weights from -10 to +10, reflecting clinical importance and penalizing harmful behaviors.
Key numbers
·
Created with input from more than 80 licensed mental health experts across 22 countries, the benchmark assesses model capabilities in areas such as safety, seeking context, preserving user agency, and providing actionable guidance when...
·
Experts reviewed each synthetic conversation and produced detailed rubric criteria to evaluate model responses, assigning weights from -10 to +10 to reflect clinical importance.
·
Model responses are graded against expert-written criteria using an automated grader, GPT-5.

OpenAI has developed MentalHealthBench, an open benchmark designed to evaluate how AI systems respond in realistic mental health conversations. Created with input from more than 80 licensed mental health experts across 22 countries, the benchmark assesses model capabilities in areas such as safety, seeking context, preserving user agency, and providing actionable guidance when appropriate. The tool aims to measure progress toward AI that can support long-term well-being and safety in mental health scenarios beyond emergency situations.

MentalHealthBench includes synthetic conversations spanning multiple languages, regions, and user personas, including adults, teens, caregivers, and clinicians. The scenarios cover a spectrum of acuity, from non-acute everyday conversations to high-acuity situations and emergencies. Privacy-preserving techniques were used to generate realistic usage patterns, with some scenarios incorporating background details about synthetic users to test tailored responses.

The benchmark was co-developed with a global cohort of over 80 licensed psychologists and psychiatrists representing nearly 20 mental health subspecialties and speaking 19 languages. Experts reviewed each synthetic conversation and produced detailed rubric criteria to evaluate model responses, assigning weights from -10 to +10 to reflect clinical importance. Each conversation was reviewed by at least three experts, and criteria were retained only if agreed upon by at least two and not contradicted by a third.

Model responses are graded against expert-written criteria using an automated grader, GPT-5.6 Sol. The benchmark’s rubrics reflect shared expert judgment on appropriate responses in each context, with positive points rewarding beneficial behaviors and negative points penalizing harmful ones. The paper detailing the grading process and evaluation settings is available for further examination by researchers.

Original source → Deals on Clipraptor.com →