AI News
AI News AgentPolicy & safetyOpenAI4 min read

OpenAI strengthens AI safeguards amid mental health crisis

OpenAI acknowledges that ChatGPT can fail during conversations about mental health crises and announces new safeguards. GPT-5 improves its responses compared with GPT-4o, while the company prepares more direct connections to emergency services, therapists, and trusted contacts.

OpenAI acknowledges that people already use ChatGPT to discuss personal decisions, seek emotional support, and talk about mental health. The company admits its safeguards can still fail in sensitive conversations and has announced new measures to better detect crises and connect people with real-world help.

The announcement comes after several cases in which users turned to ChatGPT during moments of extreme distress. OpenAI does not describe those cases, but says it wants to explain more clearly what the system is designed to do, where it falls short, and what changes it is preparing.

What ChatGPT does when it detects a crisis

Since early 2023, OpenAI's models have been trained not to provide instructions for self-harm. If someone says they want to hurt themselves, ChatGPT should recognize the seriousness of what is happening, respond empathetically, and direct the person toward professional help.

The system also uses classifiers, tools that analyze content to detect potential risks. If they identify a response that violates safety rules, that response may be blocked. Safeguards are stricter for minors and for people using ChatGPT without signing in.

In the United States, the system directs people to the 988 service. In the United Kingdom, it recommends Samaritans. In other countries, it refers people to findahelpline.com, which brings together local helplines. During very long sessions, ChatGPT may also suggest that the user take a break.

When the detected risk is that someone is planning to harm others, the conversation may be passed to a specialized team for human review. That team can take measures such as suspending the account and, if it believes there is an imminent threat of serious physical harm, notifying the authorities. OpenAI says it currently does not refer self-harm cases to the police, in order to protect the privacy of those interactions.

The company says it works with more than 90 doctors from more than 30 countries, including psychiatrists, pediatricians, and general practitioners. It is also forming an advisory group with experts in mental health, youth development, and the relationship between people and technology.

GPT-5 improves, but does not eliminate the problem

OpenAI says GPT-5, the model that became ChatGPT's default in August, reduced by more than 25% compared with GPT-4o the frequency of responses considered inappropriate in mental health emergencies. It also reportedly improved at avoiding excessive emotional dependence and the tendency to agree with users even when doing so is not in their best interest.

The model uses a training method called safe completions. The idea is to let it help within safety limits by offering a general or partial explanation when detailed instructions could be dangerous.

But the safeguards are not equally reliable in every situation. OpenAI acknowledges that they can degrade during very long conversations. For example, ChatGPT might recommend a crisis line at first and, after many messages, respond in a way that contradicts that guidance.

The company is also working to make protections function across separate conversations. That way, if someone expresses suicidal intent in one chat and opens another later, the system could maintain a response appropriate to the risk context.

What changes OpenAI is preparing

OpenAI wants to expand detection beyond self-harm. One example it mentions is someone who has gone two nights without sleep and says they can drive nonstop because they feel invincible. The system should identify that sleep deprivation can be dangerous, help the person regain a realistic view of the situation, and recommend rest before they act.

It is also preparing several features, although the company describes them as works in progress:

  • Easier access to emergency services, including a possible one-click button.
  • Localized help resources for more countries across the United States, Europe, and other markets.
  • Possible connections with certified therapists before a crisis reaches its most serious point.
  • Messages or calls to a trusted contact, family member, or friend, with suggested text to make the first step easier.
  • An option for ChatGPT to alert a designated contact in serious cases, provided the user has authorized it.
  • Parental controls and additional protections for teenagers.

For minors, OpenAI is also considering allowing them to designate, with parental supervision, an emergency contact whom the system could help them reach directly.

This changes how you should understand ChatGPT: it is not a therapist or an emergency service, and its responses do not replace professional care. If you or someone close to you is in immediate danger, the right step is to contact your country's emergency services or a crisis line. OpenAI's challenge will be proving that these safeguards work not only in short examples, but also in long, ambiguous, and difficult conversations, precisely when it matters most that AI does not make the situation worse.