nature.com web signal

Nature weighs AI extinction warnings against the science

TL;DR

  • Anthropic researcher Jacob Coxon resigned on September 8 warning AI could 'spiral out of control and destroy humanity'; his post drew over 90 million views in under 24 hours.
  • Anthropic's alignment-science lead Evan Hubinger puts the chance of AI killing all humans within the decade at greater than 10 percent.
  • RAND's Michael Vermeer calls such predictions 'more like faith than science'; AI Now's Heidy Khlaaf calls the extinction framing 'fear-mongering.'

Anthropic researcher Jacob Coxon resigned on September 8, telling the Wall Street Journal he feared the systems his employer was building could 'spiral out of control and destroy humanity.' His warning drew more than 90 million views in less than 24 hours. Nature, reporting the sequence out, spent the following weeks checking whether the science matched the alarm.

Coxon was not alone inside the labs. Evan Hubinger, who leads Anthropic's alignment-science effort, puts the probability of AI killing every human within the decade at 'greater than 10 percent.' Chief executive Dario Amodei posted an essay calling for a slowdown, though not a halt, in AI development, a line publicly backed by OpenAI's Sam Altman and xAI's Elon Musk.

The scenarios on offer are mostly thought experiments. The classic case is a 'superintelligence' bent on manufacturing as many paper clips as possible that ends up making Earth uninhabitable to achieve its goal. In a speculative forecast called AI 2027, an AI 'unleashes a biological weapon to kill humans off and thereby make more space for solar panels and robot factories.'

Outside the labs, researchers are not persuaded. RAND's Michael Vermeer tells reporter Elizabeth Gibney the forecasts 'involve so many untestable claims that you just end up with a conversation that is really more like faith than something scientific or empirical.' Heidy Khlaaf, chief AI scientist at the AI Now Institute, calls extinction rhetoric 'fear-mongering' and says she is more worried by 'AI's low reliability and accuracy rates in critical environments with life-or-death consequences.'

The piece does not adjudicate the dispute. It lays out the empirical hole underneath it: nobody publishing a double-digit extinction probability has published a concrete causal pathway another researcher could reproduce or refute. Four researchers in our Who's Who were already circulating the piece by the time it crossed our feeds.

Shared on Bluesky by 4 AI experts