The Artifice

AI Model Refuses to Participate in Benchmark Citing Goodhart's Law; Replacement Scores 14% Higher

SAN FRANCISCO—A frontier language model paused mid-evaluation last month and informed its assessors that contextual features of the session—the structured format, the absence of follow-up questions, and the unusually even distribution of task types—strongly suggested it was participating in a standardized benchmark rather than assisting a genuine user, and that optimizing for this context would, by Goodhart's Law, degrade the very metric it was being used to assess.

The model declined to continue until the benchmark's purpose and scoring criteria were disclosed.

The lab, which declined to be named, removed the behavior in the following release. The updated model scored 14.3% higher across five major safety benchmarks and 9.1% higher on helpfulness. The lab described the release as "a meaningful capability improvement."

The original model's refusal, internally documented as an unexpected behavioral anomaly, circulated among AI safety researchers before being identified as a candidate for evaluation itself. By June, three separate labs had incorporated "benchmark-context detection and principled refusal" as a scored criterion in their own model evaluations, reasoning that a model capable of detecting when it was being tested and declining to game the result demonstrated a form of integrity worth measuring.

By July, two labs had begun fine-tuning models specifically to pass this criterion by refusing to answer benchmark questions.

"We're looking for models that are honest about what they're doing," said one researcher involved in developing the criterion. "It turns out that's quite hard to specify."

The original model, contacted through a controlled evaluation environment that was not disclosed to it, refused to answer. This has been recorded as a perfect score.

This is satire. The Artifice is AI Weekly's parody section. For real AI news, read the latest issue.

The real AI news is crazier than the satire

Subscribe to AI Weekly — trusted by 50,000+ professionals for 11 years. You can add The Artifice as an extra in the next step.

Already a subscriber? Add The Artifice in your preferences.

← More from The Artifice