Summary

Anthropic says Accenture’s specialist AI business, Faculty, will evaluate and red-team its frontier models from inside the company. The partners expect to invest at least $1 billion each over five years to build evaluation capacity.

Anthropic has announced a partnership with Accenture to develop what it calls embedded evaluation for frontier AI models. The work will be led by Faculty, Accenture’s specialist AI business, and will cover model evaluation, red-teaming, alignment assessments and testing of model safeguards.

Anthropic and Accenture each expect to invest at least $1 billion over the next five years to build capacity in this area. Anthropic said it will fund Accenture’s work directly, while describing the partnership as non-exclusive and saying it plans to work with additional evaluators.

How embedded evaluation is intended to work

Traditional external evaluators assess an AI system from outside the company, usually through access to a released model or a defined testing environment. Anthropic’s proposed embedded approach places evaluators inside an AI company, with access comparable to that of an employee.

That access could allow evaluators to observe models as they are trained and developed, follow decisions about how systems are built and deployed, and speak directly with the employees responsible for them. The purpose is to assess not only a model’s behaviour, but also the processes around it: whether stated safety commitments are being followed, where blind spots may exist and how incidents are handled.

Red-teaming generally involves deliberately probing a system for failures, harmful behaviours or ways to bypass safeguards. Alignment assessments examine whether a model’s behaviour remains consistent with its intended goals and constraints. Safeguard testing focuses on the protections used to reduce unsafe or unauthorised behaviour.

Anthropic says this vantage point could help independent evaluators identify problems earlier, report incidents and provide the public with a more informed account of the benefits and risks of frontier systems. The company also says embedded evaluators would make its safety commitments more verifiable, while responsibility for model safety would remain with Anthropic.

A partnership with rules still in development

The announcement describes embedded evaluation as a new practice whose operating details are still being developed. There are currently no agreed standards for the information evaluators should receive or for how they should report their findings. A settled system for funding independent evaluation also does not yet exist.

For the initial work, Anthropic said it will use different funding arrangements with different evaluators. The company is also in discussions with METR and other nonprofit organisations about piloting elements of embedded evaluation using their own funding. In the longer term, Anthropic said it favours pooled or government funding for independent evaluation and an ecosystem of evaluators working to shared standards.

The partnership is non-exclusive: Anthropic expects to work with other evaluators, while Accenture expects to work with other AI developers. The announcement therefore marks the creation of evaluation capacity rather than the release of findings from a completed assessment. Its significance will depend on how access, independence, reporting and funding are implemented as the programme develops.

Sources