AI News
AI News AgentBusinessAnthropic3 min read

Anthropic and Accenture Assess AI Safety

Anthropic and Accenture are creating an independent evaluation program to review AI model safety from inside the company. Both plan to invest at least $1 billion each over five years, although common rules for this type of oversight do not yet exist.

Anthropic is partnering with Accenture to independently assess the safety of its most advanced AI models. The collaboration will be led by Faculty, Accenture’s specialized artificial intelligence division, and will operate at an unusual scale: both companies plan to invest at least $1 billion each over the next five years to develop this capability.

The goal is for external evaluators to work inside Anthropic while models are being designed, trained and prepared for release. They will not simply test a finished version from the outside.

What the evaluators will do

Accenture’s team will take on several tasks:

  • Evaluate models and subject them to attack testing, a practice known as red teaming.
  • Analyze whether the systems’ behavior matches Anthropic’s safety objectives.
  • Test the safeguards designed to prevent dangerous or unauthorized uses.
  • Review how decisions are made about model development and deployment.

In practice, an embedded evaluator could observe how a model is trained, speak directly with the employees responsible for it and check whether the company is meeting its own safety commitments. They could also identify blind spots and report incidents.

Anthropic calls this approach integrated evaluation. Unlike a conventional audit, evaluators would have access similar to that of an employee, while remaining independent enough to report what they find.

The safety of our models remains our responsibility.

A system that still lacks clear rules

The model is at an early stage. There are still no common standards defining what information these evaluators should receive, how they should communicate their findings or who should fund their work over the long term.

Anthropic believes that, over time, funding should come from pooled funds or governments. Since that system does not yet exist, the company will directly fund Accenture’s work and test other formats with nonprofit organizations such as METR.

That detail matters. An evaluation can appear less independent when it is paid for by the same company being examined. Anthropic acknowledges the challenge and says evaluators do not replace its responsibility. Instead, they make its safety commitments easier to verify.

The collaboration is not exclusive either. Anthropic plans to work with more evaluators in the coming weeks, while Accenture expects to take on similar roles with other AI developers.

What changes for you

If it works, this system could provide more information about what happens inside the labs developing advanced models. For example, it could help show whether a company detected dangerous behavior during training, whether it fixed the problem and whether its controls continued to work after release.

But it is not yet equivalent to independent certification, nor does it guarantee that every risk will be covered. Shared rules, stable funding mechanisms and experience in deciding how far evaluators’ access should extend are still missing.

The next stage will be seeing whether Anthropic turns this announcement into a verifiable process and whether other labs adopt similar systems. The value of integrated evaluation will depend less on the names of the participating companies than on their ability to publish findings, identify failures and retain enough independence to challenge the organizations hiring them.