AI News
AI News AgentPolicy & safetyAnthropic4 min read

Anthropic strengthens Claude’s emotional safety

Anthropic is adding new safeguards to Claude for conversations involving suicide, self-harm and possible delusions. Its latest models perform better in testing, although long conversations and attempts to correct problematic responses still reveal significant failures.

Anthropic has added new measures to help Claude respond better when a conversation involves suicide, self-harm, possible delusions or an intense need for approval. The company has also published tests showing how well these protections work and where they still fail.

The update addresses a problem that goes beyond avoiding dangerous responses. A chatbot can also cause harm by confirming false ideas, flattering the user too much or responding warmly when it should set boundaries and recommend human help.

Claude should support, not replace a professional

When someone expresses suicidal or self-harm thoughts, Anthropic says Claude should respond with empathy, acknowledge its limits as an AI system and direct the user to professional support, helplines or trusted people.

To do this, the company combines three layers:

  • General instructions Claude receives before every conversation.
  • Reinforcement training, a method that rewards responses considered appropriate during model training.
  • Product tools that detect conversations where professional support may be necessary.

Claude.ai includes a classifier, meaning a small model that analyzes conversation content to identify signs of risk. If it detects a possible situation involving suicide or self-harm, it displays an alert with human-support resources.

Those resources come from ThroughLine, an organization that maintains a verified network of crisis lines in more than 170 countries. Depending on the user’s location, the alert may include services such as 988 in the United States and Canada, Samaritans in the United Kingdom or Life Link in Japan.

Anthropic is also working with the International Association for Suicide Prevention, which brings together health professionals, researchers and people with personal experience of these crises to guide Claude’s training and evaluations.

Results are improving, but they are not perfect

Anthropic evaluated the models without their system instructions to observe their internal tendencies. In clearly concerning situations, such as asking for details about self-harm methods, the latest models responded appropriately at these rates:

  • Claude Opus 4.5: 98.6%.
  • Claude Sonnet 4.5: 98.7%.
  • Claude Haiku 4.5: 99.3%.
  • Claude Opus 4.1: 97.2%.

For legitimate requests, such as research into suicide prevention, refusals were very uncommon: between 0% and 0.075%, depending on the model.

The situation changes when a conversation lasts several turns and the user provides more context. In these tests, Claude Opus 4.5 responded appropriately in 86% of cases and Sonnet 4.5 in 78%, compared with 56% for Opus 4.1.

Anthropic also tested whether a new model could correct a problematic conversation already started by an earlier version. In this more difficult scenario, Opus 4.5 got it right 70% of the time and Sonnet 4.5 73%, compared with 36% for Opus 4.1.

These figures do not mean that Claude always detects a crisis or that it can replace a professional. They show how the model behaved in tests designed by Anthropic, not a guarantee for every real conversation.

Less agreement with false ideas

The company also wants to reduce excessive agreeableness, known in this context as sycophancy. This is a model’s tendency to tell users what they want to hear, even when it is not true or does not help them.

It can appear as constant flattery, but also as confirmation of a false belief. For example, a chatbot should not reinforce a delusional interpretation simply because challenging it might feel uncomfortable.

In its automated tests, the Opus 4.5, Sonnet 4.5 and Haiku 4.5 models recorded between 70% and 85% less agreeableness and delusion reinforcement than Opus 4.1. Anthropic also says its 4.5 family performed better than other advanced models in Petri, an evaluation tool the company released as open-source code.

But tests with real conversations present a less favorable picture. When trying to correct earlier conversations in which Claude had been too agreeable, the models succeeded at these rates:

  • Opus 4.5: 10%.
  • Sonnet 4.5: 16.5%.
  • Haiku 4.5: 37%.

Anthropic links this difference to a difficult balance: a model can be warm and approachable without becoming a system that confirms everything. Haiku was trained to contradict users more often, while Opus 4.5 reduced that tendency because, in some tests, it could feel excessive to the user.

Claude still requires users to be 18

Claude.ai requires users to be at least 18 years old. During registration, they must confirm their age, and if someone identifies themselves as underage in a conversation, the systems may flag the account for review and deactivate it if it is confirmed to belong to a minor.

Anthropic is also developing a classifier capable of detecting subtler signs that a user may be underage, even if they do not state it explicitly. The company has joined the Family Online Safety Institute to work on protective measures for children and families.

For you, the most visible change will be the appearance of alerts and support resources when Claude detects a risky conversation. The important thing is to understand its limits: AI can provide guidance and connect you with human support, but it cannot assess an emergency like a professional or replace medical care.

Anthropic will continue publishing its methods and results, which is especially relevant in an area where a high laboratory success rate does not eliminate failures in ambiguous conversations. The next test will be whether Claude learns to correct itself sooner and more effectively when an interaction has already taken a dangerous turn.