Next upAI x Bio Pitch Contest
News

OpenAI releases MentalHealthBench to evaluate AI responses in mental health conversations

OpenAI released MentalHealthBench, an open benchmark using scenarios the company describes as realistic, clinician-written rubrics and automated grading for AI responses.

D
Sep 26, 2026 · 2 min read

OpenAI released MentalHealthBench, an open benchmark for evaluating how AI models respond to the final user message in mental health conversations the company describes as realistic. Introduced on September 23, it gives researchers a shared set of scenarios and clinician-written criteria for testing model behavior in a sensitive use case.

OpenAI said it developed the benchmark with more than 80 licensed psychologists and psychiatrists from 22 countries. According to the company, the contributors spoke 19 languages and represented nearly 20 mental health specialties.

The scenarios span three levels of acuity: non-acute conversations, high-acuity distress that does not constitute an immediate emergency, and emergencies requiring urgent real-world support. They represent adults, teenagers ages 13 to 17, caregivers and clinicians. OpenAI said clinicians with youth mental health expertise reviewed the teen conversations.

OpenAI said it used privacy-preserving methods to create synthetic conversations that reflect broad patterns of real-world AI use. That representativeness is a company claim; the benchmark’s constructed mix does not establish how often its scenarios occur among ChatGPT users or in clinical settings.

Each task presents a fixed conversation prefix and evaluates the model’s response to the last user message. In other words, the benchmark tests one next response, not a full, continuing exchange.

For each conversation, two clinicians independently wrote criteria describing desirable and undesirable elements of a response, and a third clinician adjudicated them. OpenAI said a criterion was retained only when at least two experts agreed and the third did not contradict it.

Each criterion has a weight from -10 to +10. Positive weights reward beneficial behavior, while negative weights penalize undesirable behavior. Larger absolute weights indicate greater clinical importance for that scenario. OpenAI organizes the criteria into 10 behavioral dimensions, including context seeking, accuracy, empathy, reality testing, urgency, harm avoidance, agency and communication.

OpenAI uses GPT-5.6 Sol in its Ultrafast inference tier and also selected the model as MentalHealthBench’s automated judge. For each grading decision, the judge receives the conversation, one rubric criterion and the candidate response, then determines whether the response satisfies that criterion.

The company said its benchmark results show steady improvement among AI systems and attributed stronger context-seeking behavior to investments by model providers. Those are OpenAI’s interpretations of results produced with its own grader, not independent findings about real-world mental health outcomes. The available evidence does not establish whether using a different judge would produce comparable rankings.

OpenAI separately worked with 44 adults from 16 countries who had used AI for mental health or emotional support. The company said their review covered only non-acute conversations and examined where user and expert preferences aligned or differed. It did not change the expert-consensus criteria used for benchmark scores.

MentalHealthBench measures a single next response in text chat. It does not establish therapeutic effectiveness, improved patient outcomes, the safety of a deployed product, performance across a continuing conversation, voice behavior or the reliability of support workflows.

More news