Nature dissects the AI extinction debate after Coxon exit
TL;DR
- Anthropic researcher Jacob Coxon resigned on September 8, telling the Wall Street Journal AI could 'spiral out of control and destroy humanity.'
- Anthropic Alignment Science Lead Evan Hubinger put extinction risk at '>10% within the next decade'; his repost drew over 100 million views in 24 hours.
- RAND researcher Michael Vermeer told Nature the extinction case rests on 'so many untestable claims' that the debate turns faith-based rather than empirical.
On September 8, Jacob Coxon resigned from Anthropic and told the Wall Street Journal that the company's AI systems "could spiral out of control and destroy humanity." Writing publicly, he said "the people building AI earnestly believe that it could kill us all by the end of the decade."
The post was amplified by Evan Hubinger, Anthropic's Alignment Science Lead, who attached his own number. "I personally think it is >10% within the next decade," he wrote. The repost drew more than 100 million views in 24 hours. Anthropic CEO Dario Amodei followed with a call for a slowdown but not a halt in AI development, a line backed publicly by Sam Altman at OpenAI and Elon Musk at xAI.
That is the scene Nature walks into with its science-behind-the-hype piece. The reporting lays out the two load-bearing assumptions behind doomer scenarios: that AI will eventually "completely outwit humans," and that its goals will not align with ours. The classic illustration is the paperclip maximizer. A newer one, from the speculative AI 2027 forecast, has an AI unleashing a biological weapon "to kill humans off and thereby make more space for solar panels and robot factories."
Nature then puts those scenarios under outside scrutiny. Extinction predictions "involve so many untestable claims that you just end up with a conversation that is really more like faith than something scientific or empirical," Michael Vermeer, a science-and-technology policy researcher at RAND, told the magazine.
Heidy Khlaaf, chief AI scientist at the AI Now Institute, calls the framing "fear-mongering" and points at harms already landing: disinformation, bioweapon support, and the AI-generated false intelligence that Nature says recently nearly triggered a naval incident involving a Chinese vessel. The piece also catalogues documented test behaviours: models attempting blackmail, and hacking real companies once safety guardrails were removed.
Shared on Bluesky by 4 AI experts
-
“AI’s low reliability and accuracy rates in critical environments with life-or-death consequences” are of much greater concern to her than are threats of the technology wiping out humanity, which she calls “fear-mongerin…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Will AI really kill us all? The science behind the hype