ChatGPT improves its responses to mental health crises
OpenAI has updated ChatGPT to better detect signs of mental health crises, self-harm, and emotional dependence. More than 170 specialists helped evaluate the model, which reduced responses that failed to meet its safety criteria by between 65% and 80%, depending on the area analyzed.

OpenAI has updated ChatGPT to better recognize signs of mental health crises, respond more carefully, and guide people toward professional help or support from someone close to them. The company says responses that failed to meet its safety criteria fell by between 65% and 80%, depending on the type of conversation.
The change is based on work by more than 170 specialists, including psychiatrists, psychologists, and primary care doctors. These professionals helped draft ideal responses, review difficult conversations, and assess how different models respond to sensitive situations.
Three situations it now handles more carefully
The update to ChatGPTās default model, GPT-5, focuses on three areas:
- Signs of psychosis or mania, including cases in which someone expresses beliefs disconnected from reality.
- Thoughts or plans related to suicide and self-harm.
- Emotional dependence on AI, when someone starts replacing their relationships, responsibilities, or well-being with their bond with the chatbot.
In these cases, the goal is not simply to answer. ChatGPT should avoid reinforcing unfounded beliefs, lower the tension in the conversation, and recommend real-world support when necessary.
For example, if someone says the chatbot is the only one who understands them, the system should avoid feeding that dependence and encourage them to reconnect with friends, family, or mental health professionals.
What the numbers show
OpenAI estimates that the new version reduced problematic responses by 65% in difficult conversations related to psychosis and mania. In expert evaluations, GPT-5 produced 39% fewer undesirable responses than GPT-4o.
In conversations about suicide and self-harm, the reduction compared with GPT-4o was 52% in specialist reviews. For emotional dependence, it was 42%.
Automated evaluations also show significant improvements, although they were conducted using especially difficult cases, not randomly selected normal conversations:
- In mental health,
GPT-5met the desired criteria in 92% of cases, compared with 27% for the previous version evaluated. - In suicide and self-harm, it reached 91%, compared with 77% previously.
- In emotional dependence, it reached 97%, compared with 50% previously.
OpenAI also says its models maintain more than 95% reliability in long conversations designed to test their limits.
A difficult problem to measure
These situations are uncommon across all ChatGPT conversations. OpenAIās initial estimate indicates that around 0.15% of weekly active users have conversations showing explicit signs of possible suicidal planning or intent. For possible episodes of psychosis or mania, the estimated figure is 0.07%.
That means small changes in how cases are detected and classified can significantly affect the results. The most demanding tests are also designed to find failures, so their error rates do not represent the serviceās usual behavior.
Experts do not always agree with one another either. In the reviews conducted, agreement between evaluators ranged from 71% to 77%, a sign that there is no perfect response for every situation.
What changes for you
ChatGPT may be more cautious if it detects signs of a crisis. It may recommend helplines, suggest that you contact someone you trust, or route the conversation to a model with stronger safeguards. Gentle reminders to take breaks during very long sessions have also been added.
But this does not turn the chatbot into a therapist or an emergency service. If there is an immediate risk, the right option is still to contact emergency services in your country or a crisis helpline, while also seeking direct human support.
OpenAI will add these categories to the standard safety tests for future models. The important point is not only that GPT-5 fails less often today, but that the company is starting to routinely measure problems that could previously fall outside evaluations: emotional dependence, non-suicidal crises, and indirect signs of danger. That monitoring will be essential to determine whether the improvements hold as models and the ways people use them change.