cbc.ca web signal

Anthropic alignment lead: >10% chance AI kills all humans

TL;DR

  • Anthropic's alignment science lead Evan Hubinger says he personally puts the odds of AI killing all humans this decade at greater than 10 percent.
  • Jacob Coxon resigned from Anthropic on September 8, accusing Anthropic and OpenAI of 'racing straight to self-improving superintelligence and gambling with our lives.'
  • Hubinger conceded Anthropic does 'not yet have a plan to solve alignment for superintelligence and are not clearly on track to.'

Anthropic's alignment science lead, Evan Hubinger, said publicly that he personally puts the odds of AI killing all humans this decade at greater than 10 percent, and that his company does not yet have a plan to prevent it.

The statement, reported by CBC News, came in reply to a resignation thread from Jacob Coxon, a former Anthropic researcher who Fox Business reports previously worked at OpenAI as well. Coxon quit on September 8, writing that "the people building AI earnestly believe that it could kill us all by the end of the decade," and that Anthropic and OpenAI are "racing straight to self-improving superintelligence and gambling with our lives."

Hubinger backed him on X: "Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

Hubinger said current models pose "low" risk, and specified his concern is about "superintelligence arising from recursive self-improvement."

Coxon's own pitch was blunter. He argued AI systems "will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources," and floated "a temporary ban on improving model capabilities" as a possible response.

Shared on Bluesky by 2 AI experts