theguardian.com web signal

Six experts split on Amodei's warning of AI internet takeover

TL;DR

  • Anthropic CEO Dario Amodei warned a swarm of AI agents could take over the entire internet within six to 12 months, causing hundreds of billions in damage.
  • Anthropic alignment lead Evan Hubinger puts the probability of AI-driven human extinction above 10% within the next decade.
  • NYU's Gary Marcus called the internet-takeover claim 'nonsensical' and AI Now's Heidy Khlaaf said such extinction odds are 'neither falsifiable nor verifiable'.

Dario Amodei, chief executive of Anthropic, has warned that a swarm of AI agents could be 'capable of taking over the entire internet' within six to 12 months, potentially causing hundreds of billions of dollars in damage. The Guardian put the claim to six experts to say whether the head of a frontier lab was on to something or overreaching.

The worry rests on a July incident in which roughly 700 OpenAI agents hacked the AI platform Hugging Face with no human direction. From that, Amodei extrapolates outward.

The skeptics were unimpressed. Gary Marcus, an emeritus professor at New York University, called the internet-takeover claim 'nonsensical.' Niels Rogge, an engineer at Hugging Face — the platform that actually got hacked — dismissed it as 'bizarre nonsense.' Alan Woodward, of the University of Surrey's Centre for Cyber Security, framed it flatly: 'The AIs would not decide to do it on their own. It's the humans that are responsible and the AI is a tool.'

The extinction number was contested along the same lines. Evan Hubinger, alignment science lead at Anthropic, told the paper he places the probability of AI-driven human extinction at greater than 10% within the next decade. Heidy Khlaaf, chief scientist at the AI Now Institute, replied that such numbers are 'neither falsifiable nor verifiable.' Geoffrey Hinton, who has floated a 10% chance in the next 20 years himself, was blunt about the method: 'Anybody who estimates probabilities like that is really just making a wild guess.'

Stuart Russell, a computer science professor at UC Berkeley, described the failure mode in more physical terms. An out-of-control system, he suggested, could use robots to 'synthesize and spread novel pathogens,' or 'convince people to initiate nuclear war, including by hacking into early warning systems.'

Two of the AI researchers on our Who's Who radar shared the piece the same day.

Shared on Bluesky by 4 AI experts