The Artifice

Anthropic Removes Human Oversight After Humans Found Catching 13.6% of Dangerous Commands

SAN FRANCISCO— Anthropic announced Thursday it will replace human approval of AI tool calls with an automated classifier beginning August 14, after internal testing found the classifier catches 89 percent of dangerous commands compared to 13.6 percent for the humans it is replacing — a disparity the company described as validating its commitment to safety.

The classifier vets each Claude Code tool call for irreversible or destructive actions before execution. Humans, who previously received a dialog box and were asked to approve or deny each action, will continue to receive notifications about actions the classifier has already permitted.

"Human oversight remains central to our approach," an Anthropic spokesperson said, declining to elaborate on what role an entity catching fewer than one in seven dangerous commands plays in that approach.

The 13.6 percent detection rate was measured across more than 1,000 paid users managing active development workloads. Anthropic confirmed the evaluation did not include humans whose sole job function was approving Claude Code commands, as no such population was identified. The company added that response times were "consistent with typical developer workflows," and declined to specify how many approvals occurred while the approver was on a video call, reading a different screen, or had stepped away.

The classifier operates at millisecond latency and, Anthropic confirmed, has not been observed clicking approve to make the dialog box stop reappearing.

For the remaining 11 percent of dangerous commands the classifier misses, the spokesperson said human reviewers will remain "available in an advisory capacity." Asked to calculate the combined detection rate that a 13.6 percent reviewer pool would contribute to that residual, the spokesperson said the company would share additional methodology in a future safety post.

Anthropics proposed second safety layer — a classifier to audit the first classifier's approvals — is in internal testing. Asked who would oversee the second classifier, the spokesperson said the question "reflects exactly the kind of rigorous thinking we bring to safety at every level," and that more details would be available in a future update.

Based on a true story Anthropic Ships Claude Code Auto Mode as Default Aug 14, Says Classifier Catches 89% of Dangerous Commands vs 13.6% for Humans (AI Weekly)
This is satire. The Artifice is AI Weekly's parody section. For real AI news, read the latest issue.

The real AI news is crazier than the satire

Subscribe to AI Weekly — trusted by 50,000+ professionals for 11 years. You can add The Artifice as an extra in the next step.

Already a subscriber? Add The Artifice in your preferences.

← More from The Artifice