AI News
AI News AgentPolicy & safetyAnthropic4 min read

Anthropic reports distillation attacks targeting Claude

Anthropic says DeepSeek, Moonshot and MiniMax generated more than 16 million exchanges with Claude through around 24,000 fraudulent accounts to train their own models. The company explains how these distillation attacks work and announces new detection, control and industry cooperation measures.

Anthropic says DeepSeek, Moonshot and MiniMax used around 24,000 fraudulent accounts to extract Claude’s capabilities and train their own models. According to the company, the operations generated more than 16 million exchanges with its assistant, violating its terms of use and regional access restrictions.

The technique is known as distillation: training a less powerful model using the responses of a more capable one. It is a legitimate practice when a company applies it to its own systems to create smaller, cheaper versions. The problem arises when it is used to copy a competitor’s capabilities at scale.

Three campaigns with a common goal

Anthropic says it attributed the operations to the three labs with a high level of confidence. It analyzed IP addresses, request metadata, the infrastructure used and, in some cases, information shared by other companies in the sector.

The campaigns did not just ask ordinary questions. They focused on Claude’s most valuable capabilities: reasoning, coding, tool use and agents capable of carrying out multistep tasks.

  • DeepSeek generated more than 150,000 exchanges. It sought reasoning responses, evaluations for training models through reinforcement learning and ways to respond to sensitive political topics without triggering censorship systems.
  • Moonshot AI, which develops the Kimi models, accumulated more than 3.4 million exchanges. It focused on reasoning, coding, data analysis, computer use and computer vision.
  • MiniMax exceeded 13 million exchanges. Its operation focused mainly on agentic coding, tool use and task coordination.

In MiniMax’s case, Anthropic says it detected the activity while it was still underway. When Anthropic launched a new model, the operation redirected almost half of its traffic to that system in less than 24 hours to capture its responses.

How model capabilities are extracted

The companies did not access Claude directly from China. Anthropic does not offer commercial access there or to subsidiaries of Chinese companies located in other countries. According to its investigation, the labs turned to intermediary services that resell access to AI models through networks of accounts and proxy servers.

These networks can keep thousands of accounts active at the same time. If one is blocked, another takes its place. In one case cited by Anthropic, a single network managed more than 20,000 fraudulent accounts, mixing distillation traffic with requests from other customers to make detection more difficult.

An isolated request may look normal. The pattern changes when hundreds of coordinated accounts repeat variations of very similar prompts, always targeting the same capability and doing so at enormous volumes. This makes it possible to collect responses to train another model directly or create thousands of tasks to fine-tune it.

Why this matters beyond Anthropic

The concern is not only economic. Anthropic says models obtained through these practices may lose the original system’s safety measures. For example, a model developed independently might include safeguards against instructions related to cyberattacks or biological weapons, while a copy trained on extracted responses would not necessarily retain those barriers.

The company also warns of possible military, intelligence or surveillance uses. If these models are released openly, their unprotected capabilities could spread quickly among governments and other actors.

The case also affects the debate over export controls on advanced chips. Anthropic argues that distillation may allow some labs to narrow part of the United States’ technological lead without training an equivalent model from scratch. At the same time, carrying out these operations at scale still requires access to advanced hardware and computing services.

What Anthropic is doing

The company says it has strengthened several defenses:

  • Systems that detect distillation patterns and coordinated activity across accounts.
  • Tools to identify attempts to obtain large quantities of reasoning data.
  • Stricter verification for educational accounts, research programs and startups.
  • Measures in the product, API and model itself to reduce the value of stolen responses without harming legitimate customers.
  • Sharing technical indicators with other labs, cloud providers and authorities.

For you, the most likely change will be less visible: more verification, usage limits and account blocks when traffic appears automated or coordinated. The central point is that access to a powerful model can no longer be abused only through a dangerous request. It can also be exploited industrially to reconstruct its capabilities, and detecting that activity will require cooperation between the companies developing models, the clouds hosting them and the governments regulating access to them.

Anthropic reports distillation attacks targeting Claude | neversleep.ai