Anthropic updates its AI scaling policy
Anthropic has published version 3.0 of its responsible scaling policy and acknowledges that measuring the risks of advanced models is harder than expected. The new framework separates its own commitments from recommendations for the industry and introduces roadmaps, regular risk reports and external review.

Anthropic has published version 3.0 of its responsible scaling policy, the voluntary framework it uses to decide which safety measures to apply when its AI models reach new capabilities. The company acknowledges that some parts of the system have worked, but that others depend on assumptions that no longer fit the current pace of development.
The original policy, published in September 2023, classified models into different safety levels known as ASL. The idea was simple: if a model exceeded a certain capability threshold, Anthropic had to activate stricter safeguards.
For example, a model with enough capability to help create biological weapons would require additional controls against malicious use and the theft of its internal files, known as model weights.
What has changed since 2023
Today’s models no longer just answer questions. They can browse the internet, write and run code, use computers and complete multistep tasks with some degree of autonomy. That has created risks that were not considered when Anthropic designed its first policy.
The company says the framework did produce concrete results. In May 2025, it activated ASL-3 protections for relevant models, focused mainly on chemical and biological risks. These measures include input and output classifiers, systems that detect and block requests or responses related to dangerous content.
The approach also influenced other companies. OpenAI and Google DeepMind adopted similar frameworks a few months later, while some principles from these voluntary policies have appeared in regulatory initiatives in California, New York and the European Union.
But the system has one central problem: it is not always clear when a model has crossed a risk threshold.
The problem of measuring dangerous capabilities
Anthropic explains that its models already show enough biological knowledge to pass many quick tests. That makes it impossible to claim that the risk is low. However, those same tests do not by themselves prove that the risk is high.
More comprehensive studies, such as those requiring tests in real laboratories, take so long that they may be finished only after more powerful models already exist. The result is an area of ambiguity: the company can decide to apply precautions, but it does not always have sufficiently clear evidence to convince other companies or governments to do the same.
Government action has also moved more slowly than AI capabilities. Anthropic points out that the political debate has focused on competitiveness and economic growth, while safety measures have still not gained enough momentum at the federal level in the United States.
The higher levels in its policy also present a practical limit. Some safeguards against highly advanced models could be too difficult for a single company to implement on its own. Anthropic cites a RAND report stating that a safety level designed to protect model weights from institutions with the greatest cyber capabilities “is currently not possible” and would probably require help from national security agencies.
Three changes in the new policy
Version 3.0 reorganizes the framework into three main areas:
- Separating commitments from recommendations. Anthropic will distinguish between the measures it will apply regardless of what other companies do and a more ambitious roadmap of capabilities and safeguards that, in the company’s view, the entire industry should adopt.
- An advanced safety roadmap. The company will publish concrete objectives for cybersecurity, alignment, misuse controls and public policy. These will not be binding promises, but Anthropic will publicly report on its progress.
- Risk reports and external review. Every three to six months, it will publish reports on the safety of its models, their capabilities, the threats under consideration and the measures in place. Some data may be withheld for legal, privacy, intellectual property or public safety reasons.
Under certain circumstances, outside experts will review these reports with access to uncensored versions or versions with only limited redactions. Anthropic is already testing this process, although it says its current models do not yet formally require external review.
What it means for you
The update does not announce a new model or a capability available to users. It mainly changes how Anthropic promises to assess and communicate the risks of its future systems.
For people who use Claude, the direct effect will be limited in the short term. The significance lies elsewhere: periodic reports could make it possible to check whether safety measures are advancing at the same pace as model capabilities, rather than relying solely on the company’s general statements.
Anthropic also acknowledges that its unilateral commitments have limits. For the most serious risks, safety may depend on shared standards, cooperation between companies and government involvement. The new policy aims to make that distinction visible before models reach levels that are even harder to control.