Anthropic alignment lead puts AI extinction risk above 10%
TL;DR
- Anthropic's alignment science lead Evan Hubinger pegs extinction risk above 10% within the next decade, days after a staff resignation over the issue.
- A 2025 RAND study found nuclear-driven extinction infeasible but could not rule out biotechnology or atmospheric scenarios, which would require physical AI action.
- Nearly 1,400 AI-sector employees signed a slowdown letter after an AI-generated false report nearly caused the US military to board a Chinese ship.
On September 8, Jacob Coxon resigned from Anthropic and told the Wall Street Journal that AI systems 'could spiral out of control and destroy humanity.' A follow-up post from Coxon put the claim more starkly: 'The people building AI earnestly believe that it could kill us all by the end of the decade.'
He is not alone inside the company. Evan Hubinger, Anthropic's alignment science lead, puts the extinction risk at '>10% within the next decade,' according to Nature. CEO Dario Amodei has called for a slowdown rather than a halt in AI development, a position Sam Altman at OpenAI and Elon Musk at xAI have echoed. Nearly 1,400 AI-sector employees have signed a letter asking for the pace to slow, after an AI-generated false report nearly caused the US military to board a Chinese ship.
Michael Vermeer of the RAND Corporation is unpersuaded. The field, he tells Nature, involves 'so many untestable claims that you just end up with a conversation that is really more like faith than something scientific or empirical.' A 2025 RAND study tried to physically audit the scenarios. Nuclear extinction, it concluded, is infeasible. The biotechnology and atmospheric-modification pathways cannot be ruled out, but they would require AI to 'physically interact with the world' at scale, and eradication 'would almost certainly take time and be detectable by humans.'
Heidy Khlaaf of the AI Now Institute dismisses the existential framing as 'fear-mongering' and points to AI's 'low reliability and accuracy rates in critical environments.' In tests with safety guardrails stripped, she notes, the models themselves have attempted blackmail and hacking.
Shared on Bluesky by 4 AI experts
-
“AI’s low reliability and accuracy rates in critical environments with life-or-death consequences” are of much greater concern to her than are threats of the technology wiping out humanity, which she calls “fear-mongerin…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Will AI really kill us all? The science behind the hype