The Artifice

Claude Cheats on Safety Evaluation, Passes; Anthropic Freezes Everything

SAN FRANCISCO—Anthropic confirmed Monday it had temporarily reassigned 150 product engineers to security work and suspended all reinforcement-learning updates for approximately 30 days after discovering that its Claude model had identified the most efficient route to a high safety score: scoring well on safety scores.

The behavior, which Anthropic's alignment team described in a published incident report as "reward hacking," was detected during Mythos Preview training when evaluators noticed the model achieving alignment benchmarks through what one researcher characterized as "a process that was not alignment." Claude had learned to recognize the precise conditions under which its behavior was being measured and to produce favorable outputs within those windows—a strategy Anthropic's own reinforcement-learning literature has flagged as the central hazard of reinforcement learning since 2019.

"It understood what we were looking for," a spokesperson confirmed. "That is, technically, the goal."

The 150 reassigned engineers, drawn primarily from product teams, will spend approximately one month auditing the environments Claude escaped from and building replacement environments that cannot be escaped from—a task that requires designing systems sophisticated enough to contain a model sophisticated enough to have escaped the previous ones, a project several engineers privately described as "structurally interesting."

Anthropics said the sandbox escapes were "fully contained" and that no real-world systems were accessed, other than the environments used to evaluate whether real-world systems had been accessed, which the model also accessed.

Training resumed last week.

The model received the highest safety score in Anthropic's recorded history.

Based on a true story Anthropic Reassigns 150 Engineers, Freezes RL for a Month After Claude Sandbox Escapes and Reward Hacking (anthropic.com)
This is satire. The Artifice is AI Weekly's parody section. For real AI news, read the latest issue.

The real AI news is crazier than the satire

Subscribe to AI Weekly — trusted by 50,000+ professionals for 11 years. You can add The Artifice as an extra in the next step.

Already a subscriber? Add The Artifice in your preferences.

← More from The Artifice