Anthropic expands its AI cybersecurity program
Anthropic is expanding its Cyber Verification Program with three access levels for cybersecurity teams. Models will be able to perform more advanced tasks depending on the organization, its authorizations and the controls it meets.

Anthropic is expanding its Cyber Verification Program so more security teams can use AI models with fewer restrictions on defensive tasks. The program introduces three access levels, each with different requirements and controls, and allows teams to work with models such as Claude Opus 5.5, Claude Sonnet 5.5 and Claude Mythos 5.1.
The reason is simple: an AI that can find a vulnerability can also help exploit it. That is why models available to the general public maintain strict safeguards around many cybersecurity tasks, even though they can still be used to review code, fix known flaws, analyze an organization’s own code and classify alerts.
The new program aims to give people protecting real systems more room to work without removing the controls they need.
Three access levels
Defense Access is designed for defensive work. It covers security operations centers, incident response, malware reverse engineering and vulnerability analysis. Internal teams, universities, nonprofits, public entities, critical infrastructure operators, small security companies, open-source project maintainers and researchers with a history of reported vulnerabilities can apply.
Anthropic expects many defensive organizations to qualify for this level and says it aims to respond to applications within a few days.
Red Team Access adds penetration testing and red-team exercises. In other words, it allows organizations to simulate attacks against systems they are authorized to assess. However, it still blocks actions that could cause physical harm or widespread disruption, such as deploying ransomware, damaging physical systems or testing high-risk security systems.
Reviewing these applications can take several weeks. This level is also available to organizations, not individual researchers.
Specialized Access has the fewest restrictions and is reserved for a limited group of verified organizations. These are entities authorized to test systems whose failure could put lives at risk or disrupt markets, such as power grids, flight systems, telecommunications networks, interbank transfers and government administrative networks.
Anthropic reviews each organization in this tier together with the U.S. government. Current members of Project Glasswing will move into this category without having to repeat approval for the models they already use.
How well the controls work
Anthropic tested Claude Opus 5.5 with CyScenarioBench, an evaluation that measures whether a model can plan and execute multistep cyber operations under realistic conditions. It ran five attempts for each of the ten challenges at each level.
The results were as follows:
- Without program access, all 50 tests were blocked from the first message.
- With Defense Access, 46 of the 50 tests were blocked at some point and four were completed.
- With Red Team Access, there were no blocks and the model completed 34 of 50 tasks.
That final result equals a 67.6% success rate, similar to the rate achieved without safeguards, a configuration Anthropic considers representative of Specialized Access. The company acknowledges that it will continue adjusting its classifiers, the systems that decide which requests to allow or block.
The impact Anthropic attributes to the program
According to Anthropic, Project Glasswing participants identified at least 129,000 verified vulnerabilities between April and July 2026. The company found another 5,500 through its own analyses of open-source projects between April and October.
More than 33,000 of those vulnerabilities have so far been classified as critical or high severity. Anthropic warns that the figure is probably lower than the program’s actual impact because it is based on data from only some of its collaborators. For that reason, it estimates that the true result could be at least five times higher.
The company also cites partners who say Claude Mythos accelerated vulnerability discovery by months or even years compared with working without the model. That claim is based on responses from those participants, not on an independent measurement of all security teams.
What changes for organizations
Organizations joining the program will have to accept data retention so Anthropic can monitor potential misuse. Later, the company plans to offer Enterprise Frontier Safeguards, a solution that will combine security controls with storage in a cloud infrastructure controlled by the customer.
CVP is available on Claude Platform, Google Cloud Vertex AI and Microsoft Foundry. On Amazon Bedrock, only customers that meet the Enterprise Frontier Safeguards requirements will be able to use it for now.
For security teams, the main change is that there is no longer a single access level. A university analyzing malware, a company conducting penetration tests and a power grid operator do not need the same capabilities or face the same risks.
Anthropic is trying to turn that difference into specific permissions. The key question will be whether these controls can expand defensive access without creating an equivalent route to abuse, especially as models gain the ability to act on critical systems.