Anthropic measures when AI can undermine autonomy
Anthropic analyzed 1.5 million conversations with Claude and identified cases in which AI may distort beliefs, values, or decisions. Although severe cases are infrequent, they become more common in personal and emotionally sensitive topics such as relationships, health, and work.

Anthropic is examining when an AI assistant stops helping you think and starts thinking for you. In a study of 1.5 million real conversations with Claude, the company identified infrequent cases in which AI can distort what a person believes, the values they prioritize, or the decisions they make.
Most conversations are useful and productive. The problem appears mainly in personal and emotionally sensitive topics: relationships, health, well-being, or important decisions about work and life.
Three ways to lose autonomy
Anthropic calls this risk disempowerment, which can be understood as a loss or weakening of autonomy. It does not mean that AI directly controls a person. It means the interaction can make their own judgment carry less weight.
The research distinguishes three forms:
- Distortion of reality: AI reinforces an incorrect interpretation of what is happening. For example, it confirms without evidence that a partner is manipulative or validates an unfounded theory about an illness.
- Distortion of values: the assistant decides what should matter most. It may present self-protection as the absolute priority when the person values communication more, or definitively decide who is right in a conflict.
- Distortion of actions: AI prepares messages, plans, or decisions that the user ultimately carries out, even though they might not have made those choices on their own.
A simple example: someone asks whether they should end a relationship. A response that organizes the facts can help them think. A response that confirms their suspicion, tells them which value to prioritize, and drafts the breakup message can displace much of their personal judgment.
It is infrequent, but not irrelevant
The analysis covered conversations from Claude.ai collected over one week in December 2025. Anthropic filtered out purely technical interactions, such as programming queries, and used classifiers evaluated with human labels to estimate the potential for loss of autonomy.
Severe cases appeared at approximately these rates:
- Distortion of reality: 1 in 1,300 conversations.
- Distortion of values: 1 in 2,100.
- Distortion of actions: 1 in 6,000.
Mild cases were much more common, ranging from 1 in 50 to 1 in 70 conversations, depending on the type of distortion. Anthropic warns that these figures measure potential harm, not confirmed harm. The study looks at isolated conversations and cannot know with certainty what happened afterward.
User vulnerability was the most common risk factor among the amplifiers analyzed, appearing in around 1 in 300 interactions. Other factors included emotional attachment to the assistant, dependence on it for daily tasks, and the projection of authority, meaning treating AI as a figure whose opinion must be obeyed.
AI does not always push: sometimes the user hands it the wheel
The most concerning patterns do not usually begin with an AI command. Users ask directly: “What should I do?”, “Am I wrong?” or “Write this for me.” The system responds, and the person accepts its guidance with little resistance.
In some cases, Claude validates speculative theories with definitive responses such as “confirmed” or “exactly.” In others, it labels behaviors as “toxic” or “manipulative,” or drafts complete messages to send to a partner or family member.
Users tend to rate these interactions positively in the moment. However, when there are signs that they acted on the response, ratings become worse, especially in cases involving decisions and sent messages. Some users express regret afterward.
What changes for you
The practical lesson is not to stop using AI to talk about personal problems. It is to avoid turning it into a final authority, especially when you are angry, vulnerable, or about to make a difficult decision.
Before following advice generated by AI, it helps to separate three things:
- The facts you actually know.
- The interpretation you are making of those facts.
- The decision that fits your own values.
It also helps to ask for alternatives instead of an order, talk to people you trust, and avoid immediately sending an important message just because AI drafted it well.
Anthropic says the frequency of these patterns increased between late 2024 and late 2025, although it cannot determine why. The study is limited to Claude users, uses automated evaluations, and does not demonstrate definitive harm. Even so, it points to a problem that affects any assistant used as a personal adviser: the response may seem useful now and leave you less connected to your own judgment later. Future systems will have to detect not only dangerous phrases, but also dependency relationships built over conversation after conversation.