Introducing MentalHealthBench
OpenAI introduced MentalHealthBench, an open benchmark co-created with over 80 licensed mental health experts to evaluate AI responses in realistic mental health conversations across safety, context, and user agency.
OpenAI has developed MentalHealthBench, an open benchmark designed to evaluate how AI systems respond in realistic mental health conversations. Created with input from more than 80 licensed mental health experts across 22 countries, the benchmark assesses model capabilities in areas such as safety, seeking context, preserving user agency, and providing actionable guidance when appropriate. The tool aims to measure progress toward AI that can support long-term well-being and safety in mental health scenarios beyond emergency situations.
MentalHealthBench includes synthetic conversations spanning multiple languages, regions, and user personas, including adults, teens, caregivers, and clinicians. The scenarios cover a spectrum of acuity, from non-acute everyday conversations to high-acuity situations and emergencies. Privacy-preserving techniques were used to generate realistic usage patterns, with some scenarios incorporating background details about synthetic users to test tailored responses.
The benchmark was co-developed with a global cohort of over 80 licensed psychologists and psychiatrists representing nearly 20 mental health subspecialties and speaking 19 languages. Experts reviewed each synthetic conversation and produced detailed rubric criteria to evaluate model responses, assigning weights from -10 to +10 to reflect clinical importance. Each conversation was reviewed by at least three experts, and criteria were retained only if agreed upon by at least two and not contradicted by a third.
Model responses are graded against expert-written criteria using an automated grader, GPT-5.6 Sol. The benchmark’s rubrics reflect shared expert judgment on appropriate responses in each context, with positive points rewarding beneficial behaviors and negative points penalizing harmful ones. The paper detailing the grading process and evaluation settings is available for further examination by researchers.