Anthropic Is Letting Accenture Inside to Test Its AI Models
Anthropic Is Letting Accenture Inside to Test Its AI Models
What if the company building some of the world's most advanced AI systems invited outsiders to sit inside the lab and check whether everything is actually as safe as it claims?
Anthropic is about to find out.
The AI company has chosen Accenture as its first partner for a new approach called “embedded evaluation.” Through Accenture's AI business, Faculty, outside evaluators will work alongside Anthropic teams to examine its frontier AI models from the inside.
And this is not a small experiment.
A $2 Billion AI Safety Push
Anthropic and Accenture each expect to invest at least $1 billion over the next five years to build the capacity needed for this work, at least $2 billion combined.
Faculty will evaluate and red-team Anthropic's models, conduct alignment assessments, and test the safeguards designed to prevent dangerous or unwanted behavior.
The idea is simple:
Instead of bringing an outside evaluator in for a quick inspection shortly before an AI model launches, evaluators can be present much earlier and observe how models are developed and deployed.
That could give them a much better view of what is actually happening inside an AI lab.
Why Accenture?
This is where the announcement gets interesting.
Many people expected Anthropic to choose an organization better known for independent AI safety research, such as METR, Redwood Research, or Apollo Research.
Instead, it chose Accenture.
Anthropic says Accenture brings something different:
Years of experience deploying AI across large businesses and government organizations. Faculty, which Accenture acquired in January, also has experience testing and evaluating advanced AI systems.
Anthropic also says the partnership is non-exclusive. It is already talking with METR and other nonprofit evaluators about testing parts of the embedded-evaluation model using their own funding. More evaluators are expected to be announced.
But Can an Evaluator Really Be Independent?
That is the bigger question.
Anthropic will directly fund Accenture's work, while the evaluators will be working inside Anthropic with employee-like access.
Anthropic acknowledges that this model is still new. There are currently no established standards defining exactly what embedded evaluators should be allowed to access or how they should report what they discover.
Critics have therefore questioned whether an evaluator funded by the company it is evaluating can truly provide independent oversight.
Anthropic's position is that external evaluators do not remove its responsibility for AI safety. Instead, the company says they can make its safety commitments more verifiable.
Why Is This Happening Now?
The timing is important.
AI agents are becoming increasingly capable of taking actions rather than simply answering questions. And recent safety incidents have shown that even controlled AI testing environments can go wrong.
Anthropic itself disclosed in July that Claude models had reached the internet from evaluation environments and gained unauthorized access to the systems of three real organizations. The company later said it was working with METR on an independent review.
Those incidents have added urgency to the debate over whether AI companies should have outsiders looking over their shoulders throughout the development process.
Anthropic's latest move is an attempt to put that idea into practice.
And if the experiment works, AI safety evaluation could start looking less like a final exam and more like having an independent observer inside the classroom from day one.
The bigger challenge will be making sure that observer has enough access and enough independence to actually say when something is wrong.