September 23, 2026
AI mental-health answers are now assessed in dialogue: 1,215 scenarios
On Sep 23, OpenAI released MentalHealthBench with 1,215 synthetic dialogues for evaluating the next AI response in conversational context. The dataset can be downloaded as a ZIP archive containing 5,262 criteria weighted from −10 to +10.

Previous AI mental-health evaluations mostly focused on emergency cases using broad criteria. MentalHealthBench evaluates the next response in the context of a specific dialogue, from everyday topics to emergency cases.
You will need to assemble the run. The ZIP does not include a ready-made runner, model responses, or evaluation logs. OpenAI’s protocol asks users to generate 4 independent responses for each task and preserve the original system message with the user context.
In the official table, GPT-6 Astra scored 57.3% versus 32.1% for GPT-4o March 2025 on average task score, not clinical outcomes.
Source
