forbes.com web signal

Anthropic Alignment Lead Puts AI Extinction Odds Over 10%

Anthropic Safety ai-safety

TL;DR

  • Anthropic's Alignment Science lead Evan Hubinger publicly estimates a greater than 10% chance AI kills all humans within the next decade.
  • Hubinger said Anthropic has no plan yet to solve alignment for superintelligence and is 'not clearly on track' to build one.
  • His post came the same day Anthropic pretraining researcher Jacob Coxon, 27, resigned and told the WSJ the AI race is 'gambling with our lives.'

Evan Hubinger, Anthropic's Alignment Science lead, wrote on X that he personally puts the chance of AI killing all humans within the next decade at more than 10%.

"we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," Hubinger wrote on X, adding that Anthropic "is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

Hubinger was replying to Jacob Coxon, a 27-year-old Anthropic pretraining researcher who resigned the same day. Coxon told the WSJ he had spent three years on pretraining work across OpenAI and Anthropic and called the frontier race "gambling with our lives," predicting that "we're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already," as Forbes reported.

Hubinger's own alarm sits further out. Present models, he said, are low risk per Anthropic's latest risk report. "What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought." That framing echoes yesterday's alert on a startup pitching a routing loop as recursive self-improvement. Coxon's exit was our other Anthropic safety alert today."}