interconnects.ai via Hacker News

Lambert Rebuts AI-Extinction Framing Around Coxon Resignation

Anthropic Safety ai-business

TL;DR

  • Nathan Lambert argues Coxon's Anthropic resignation caught fire because the audience was already primed and 'fear sells,' not because the extinction case is new.
  • He proposes 'lossy self-improvement' against RSI, citing jagged model capabilities and the absence of stacking efficiency gains promised by doom forecasts.
  • Lambert names unhardened lab infrastructure as the biggest near-term risk, pointing to OpenAI's own retrospective on misaligned behavior undetected for weeks.

"Then, some basic factors of human nature apply, with the most crucial being that fear sells. Fear is the simplest story, the one people cannot look away from," Nathan Lambert writes in his Interconnects essay on the resignation of Anthropic pretraining researcher Jacob Coxon.

Lambert casts the rollout as opportunistic rather than conspiratorial: "the Wall Street Journal had an exclusive story that Jacob coordinated before posting. I suspect that Jacob shared his plan of quitting in groupchats with AI safety advocacy groups ahead of time." He then puts distance between himself and the extinction wing of the field. "I put the probability of complete extinction as being so low it isn't worth discussing," he writes.

Then he goes at the mechanism underneath the doom framing. Against recursive self-improvement, Lambert offers what he calls lossy self-improvement: "AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at." The field, he argues, has "not seen the stacking efficiency gains that massively reduce model size and cost, leading to an explosion in progress."

The near-term worry he wants readers focused on is more prosaic. "The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven't hardened their own infrastructure, enabling AI misuse to proliferate," he writes, citing OpenAI's own retrospective in which "the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks."

His parting shot is on culture: "people at the labs operate with a religious energy. It's very common to go through very out of touch interactions with them." We carried the Coxon resignation itself a day earlier; Lambert's essay is now moving through the same analyst circles as the original story, one of more than 300 Anthropic pieces we've tracked in the last 90 days.

Shared on Bluesky by 1 AI expert