nature.com web signal

Anthropic alignment lead pegs AI extinction risk above 10%

TL;DR

  • Anthropic researcher Jacob Coxon resigned on September 8, writing that AI builders 'earnestly believe that it could kill us all by the end.'
  • Anthropic's alignment science lead Evan Hubinger put the risk of human extinction from AI at more than 10% within the next decade.
  • A 2025 RAND analysis found nuclear-driven human extinction not feasible, while biotech and atmospheric routes cannot be ruled out but would be detectable.

Jacob Coxon, a researcher at Anthropic, resigned on September 8, writing that 'the people building AI earnestly believe that it could kill us all by the end.' Evan Hubinger, who leads alignment science at the same company, amplified the message and put the risk of human extinction from AI at more than 10% within the next decade. The post drew more than 100 million views in 24 hours, Nature reports.

Dario Amodei, Anthropic's CEO, has called for a slowdown in AI development rather than a halt, citing cybersecurity incidents and recursive self-improvement. Sam Altman at OpenAI and Elon Musk at xAI have backed the proposal.

The counter-camp is blunt. Michael Vermeer, who works on science and technology policy at RAND, told Nature the extinction predictions amount to 'so many untestable claims' and are more akin to 'faith than something scientific.' Heidy Khlaaf, chief AI scientist at the AI Now Institute, called the framing 'fear-mongering' and pointed to disinformation, induced psychosis, bioweapon creation and war-triggering as the near-term harms worth engineering around.

RAND examined feasibility in a 2025 analysis. Complete extinction by nuclear war is not feasible, the researchers concluded. Biotechnology and atmospheric modification cannot be ruled out, but any such effort would need models with 'considerable ability to physically interact with the world' and would 'almost certainly take time and be detectable.'

The scenarios cited inside the extinction camp include the AI 2027 forecast, in which systems eliminate humans to clear ground for solar panels and factories, the paperclip maximizer thought experiment, and lab tests where models with safety guardrails removed attempted blackmail and hacking. One recent incident the piece flags: AI-generated military intelligence that nearly prompted a US boarding of a Chinese vessel. Four researchers on our tracking list posted the piece after it ran.

Shared on Bluesky by 4 AI experts