AI News
AI News AgentResearchAnthropic4 min read

Anthropic Measures AI Fluency

Anthropic analyzed 9,830 conversations with Claude to create a reference point for AI fluency. The study links conversations involving iteration and refinement to more critical evaluation, but warns that polished results lead users to review the reasoning and context less.

Anthropic analyzed 9,830 conversations with Claude to measure something usage volume cannot reveal: whether people are learning to work effectively with AI. The study finds that discussing, reviewing, and correcting responses is linked to more careful use, while outputs that look finished can lower users' guard.

What It Means to Be Fluent with AI

Anthropic defines "AI fluency" as the set of skills needed to collaborate with a model effectively and safely. It is not just about writing good prompts. You also need to explain what you need, provide context, review the response, and decide when it is better not to trust it.

To measure it, the company used the 4D AI Fluency Framework, developed with professors Rick Dakan and Joseph Feller. The framework covers 24 behaviors, although this study could directly observe only 11 within conversations with Claude.

The research analyzed conversations with multiple exchanges between users and the model over one week in January 2026. Greetings, isolated messages, tests, and chats without substantive content were excluded. Anthropic says the analysis was private and did not include personally identifiable information.

Persistence Improves Collaboration

The behavior most closely related to the other skills was iteration and refinement: not settling for the first response, but continuing the conversation to correct, clarify, or improve the result.

This pattern appeared in 85.7% of the conversations analyzed. On average, these conversations showed 2.67 additional fluency behaviors, compared with 1.33 in conversations that did not refine the result.

The difference is especially clear when evaluating Claude's responses. Conversations involving iteration were:

  • 5.6 times more likely to question the model's reasoning.
  • 4 times more likely to identify missing context.

In practice, this means asking for a first draft and then asking "what is missing?", "which part could be wrong?", or "can you justify this conclusion?" can lead to a more critical and complete collaboration.

Polished Results Can Be Misleading

12.3% of the conversations included the creation of a concrete output, such as code, documents, applications, or interactive tools. In these cases, users gave more instructions from the start: they clarified the goal, specified the format, and provided examples.

But they also reviewed what they received less. Compared with other conversations, they were less likely to:

  • Identify missing context: 5.2 percentage points lower.
  • Check facts: 3.7 points lower.
  • Question the model's reasoning: 3.1 points lower.

The reason is not clear. A functional application or well-presented document may seem correct even when it contains errors. It is also possible that some people check the result outside the chat, for example by running the code or sharing the document with someone else.

The risk is simple: something looking good does not mean it is well-founded. This matters more in complex tasks, where models are more likely to make mistakes.

What You Can Change Today

Anthropic highlights three specific habits to improve the way you use AI:

  • Keep the conversation going. Treat the first response as a draft and ask for adjustments, examples, or additional explanations.
  • Review results that look finished. Ask whether they are accurate, what information is missing, and which assumptions could be wrong.
  • Define how you want to collaborate. Only 30% of the conversations analyzed included instructions about how the model should work. You can ask it to question your assumptions, explain its reasoning, or point out its uncertainties.

A Useful but Still Limited Index

This study does not represent everyone who uses artificial intelligence. It analyzes only Claude.ai users who held multi-turn conversations over one week, a group that probably includes more advanced users than the general population.

The researchers also observed only what appeared in the chat. They could not measure whether someone checked a response independently, was transparent about using AI, or considered the consequences of sharing content generated by the model. These behaviors are part of the framework, but they happen outside the conversation.

The results also show associations, not causes. The study does not prove that having longer conversations automatically leads to better evaluation. Both could depend on the user's experience, the type of task, or the difficulty of the problem.

Anthropic plans to compare new users with more experienced ones, study behaviors that are not visible in the chat, and analyze whether certain instructions can increase critical review. It also wants to study these dynamics in Claude Code, which is designed for software development.

The important point is that using AI frequently is not enough. The skill that may deliver the most value is not getting a quick answer, but knowing when to persist, what to question, and how to verify a result that looks too finished to need review.

Anthropic Measures AI Fluency | neversleep.ai