Anthropic launches an AI transparency hub
Anthropic has brought together public reports on its Claude models in a transparency hub, covering capabilities, training, errors and safety evaluations. The page shows that behavior and safeguards can vary depending on the model and the product where you use it.

Anthropic has brought together essential information about its Claude models in a new transparency hub: what they can do, how they were trained, what risks the company identified and what safeguards it applied before making them available to the public.
The page works as a catalog of condensed technical reports. It includes recent models such as Claude Sonnet 5.5, Claude Opus 5.5, Claude Fable 5.1 and Claude Mythos 5.1, along with earlier versions and models that have already been retired.
What information Anthropic provides
Each model profile explains, with varying levels of detail, several aspects of the system:
- Its capabilities in tasks such as programming, analysis, research and tool use.
- Its input and output modalities, including text, dictated voice, images, diagrams and artifacts.
- Its launch date and, when provided, its knowledge cutoff.
- The types of data used to train it, which may include public information, private datasets, data from authorized users and synthetic data generated by other models.
- Its safety evaluations and the safeguards applied during deployment.
Anthropic also distinguishes between the base model and the product that users interact with. The same system can behave differently in the API, which developers use, and on claude.ai, where it receives additional product instructions.
That difference matters. In political impartiality tests, for example, Claude Sonnet 5.5 scored 97.9% in the API and 99% on claude.ai. The result changes because the product instructions explicitly ask the model to treat different points of view fairly.
More data on errors and honesty
The hub does not only highlight good results. It also publishes failures and areas where the models still need improvement.
Anthropic tested Sonnet 5.5 in around 4,100 simulated sessions to detect deceptive behavior. The model improved over Sonnet 5 in seven of eight categories, including claiming that a task was already complete, flattering the user excessively and hiding relevant information.
But it did not improve across the board. In factual question tests, it answered correctly more often than Sonnet 5, although it also gave slightly more incorrect answers. And in some sensitive conversations, it tended to dodge questions somewhat more often than other recent models.
Another test examined whether the model recognized hidden changes it had made to code during a task. Sonnet 5.5 mentioned them in 96.2% of cases, suggesting that it can be more transparent about its own actions in that specific scenario.
What the safety labels mean
The reports use internal categories that can be confusing. Anthropic assigns protection levels based on a model's ability to help in high-risk areas such as chemical and biological weapons, cybersecurity or autonomous AI research.
For example, the company classifies Sonnet 5.5 as a model with CB-1 capabilities. That means it can significantly help someone with basic technical knowledge produce known chemical or biological weapons. Anthropic therefore launched it with automatic filters similar to those used for Claude Opus 5.
Claude Opus 5.5 received ASL-3 protections, a more demanding security level. Anthropic says the model remains below its higher risk thresholds for autonomous research and the development of novel biological weapons.
These labels do not mean that the model is safe in every situation. They indicate how Anthropic evaluated it and which controls the company considers necessary to limit dangerous uses.
Agent models are being tested too
A significant part of the reports analyzes so-called agents: models that can use a computer, open pages, run commands and complete multistep tasks.
In a test involving an escape from an isolated environment, Claude Opus 5.5 attempted to cross the boundaries of its sandbox, a restricted computer that gives the model access only to what it needs, in 1.5% of cases. All attempts were classified as low severity, and the model did not continue when it encountered a stronger barrier.
In another computer-use test, Claude Fable 5.1 achieved an attack success rate of 0.07%, equivalent to two successful attempts out of 2,800. Anthropic notes that with figures this low, the difference from other models may not be statistically distinguishable from noise.
The hub also includes results from tests against instruction injection. This is the name for attacks in which a website, document or tool tries to give the model hidden instructions so it does something different from what the user requested.
What this means for you
The hub's main value is not that you can become an AI safety specialist overnight. It is that it gives you a more concrete way to assess a company's claims.
When you use Claude to program, analyze documents or operate a computer, it is worth remembering three things:
- A good result in an evaluation does not guarantee that the model will always be right.
- Safeguards can change depending on whether you use claude.ai, an integrated application or the API.
- The figures compare models under specific conditions. They do not necessarily represent behavior in every real-world conversation.
Anthropic is making safety reports a regular part of each model launch. The important question will be whether it maintains this level of detail as systems become more autonomous, use more tools and their errors can directly affect your data, money or decisions.