nature.com web signal

Anthropic staffers warn of AI extinction; scientists push back

TL;DR

  • Anthropic researcher Jacob Coxon resigned on 8 September, saying colleagues believe AI could kill everyone by the end of the decade; his X post drew over 100 million views in 24 hours.
  • Evan Hubinger, who leads Anthropic's alignment science, estimates the risk of human extinction at over 10% within the next decade.
  • RAND's Michael Vermeer and AI Now's Heidy Khlaaf argue the extinction debate rests on untestable claims and distracts from measurable AI reliability failures.

On 8 September, a researcher at Anthropic resigned because he believed the systems his employer was building could destroy humanity. 'The people building AI earnestly believe that it could kill us all by the end of the decade,' Jacob Coxon wrote in a post on X that drew more than 100 million views in 24 hours.

In a feature for Nature, Elizabeth Gibney weighs the science behind that claim. Evan Hubinger, who leads Anthropic's alignment science, puts the risk of human extinction at '>10% within the next decade.' Anthropic chief executive Dario Amodei then posted an essay calling for a slowdown in AI development, not a halt; an open letter supporting the call collected around 1,400 signatures, and Sam Altman at OpenAI and Elon Musk at xAI backed the suggestion.

Scientists outside the frontier labs push back. Michael Vermeer, who researches science and technology policy at the RAND Corporation, says the extinction debate 'involves so many untestable claims that you just end up with a conversation that is really more like faith than something scientific.'

Heidy Khlaaf, chief AI scientist at the AI Now Institute, says the greater concern is 'AI's low reliability and accuracy rates in critical environments with life-or-death consequences.' Four researchers we track had posted the piece within days of publication.

Shared on Bluesky by 4 AI experts