anthropic.com web signal

Anthropic, Accenture Pledge $1B Each to Embed AI Evaluators

5 sources tracking this story

TL;DR

  • Anthropic will pay Accenture directly for evaluation work, layering a client-vendor structure on top of an existing commercial Claude partnership.
  • Accenture's Faculty unit, acquired in January, brings NHS and government AI safety credentials that distinguish it from general consulting competitors.
  • More than 100 researchers including Geoffrey Hinton publicly demanded evaluator independence from commercial ties on the same day the deal was announced.

Anthropic and Accenture will each spend at least $1 billion over the next five years to put outside evaluators inside Anthropic's own building. The work of "evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards" will be led by Faculty, Accenture's specialist AI business unit, according to Anthropic's September 18 announcement.

The unusual part is the access. Where external evaluators typically test finished models from the outside, embedded evaluators will have "access comparable to an employee's," able to observe model training, monitor deployment decisions, and interact directly with staff. Anthropic frames the tradeoff flatly: "independent embedded evaluators do not reduce our accountability, but help to make it more verifiable."

The partnership fulfills a commitment from CEO Dario Amodei's essay "We Must Pace the Frontier." Frontier-lab staff were reportedly blindsided by that pacing plan earlier this month, per our tracker. The Accenture deal is the first concrete deliverable from it.

The arrangement is "non-exclusive." Anthropic said it is also "in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding." That distinction matters. Accenture is a paid consultancy sitting inside the lab; METR and its peers would be nonprofits paying their own way to do similar work.

Anthropic said other evaluators will be "announced in the coming weeks."

What others are reporting

Coverage cluster as of 24h after publish

  1. TechCrunch Read →

    Frames Accenture's selection as a surprise over expected picks METR, Redwood, and Apollo; flags Accenture's 8% stock surge and raises conflict-of-interest questions about independence.

    Accenture is about to take on its most high-risk consulting engagement ever.
  2. Accenture Newsroom Read →

    First-party Accenture view; centers Faculty's NHS and government AI safety track record as the credential base and frames embedded evaluation as scalable industry infrastructure.

    Safety requires both deep technical expertise and a clear understanding of how AI is used in the real world.
  3. Implicator Read →

    Surfaces the Hinton co-signed researcher letter issued the same day and details Accenture's pre-existing Claude commercial partnership as a structural conflict of interest.

    Each company expects to invest at least $1 billion over the next five years, while Anthropic will pay for Accenture's evaluation work directly.
  4. Startup Fortune Read →

    Emphasizes employee-level access during model development as structurally new, while noting the arrangement remains advisory with no regulatory enforcement mechanism.

    An external auditor who receives a finished model can only see what a company chooses to show. A team with a badge and a login can watch decisions get made.

Shared on Bluesky by 2 AI experts