pluralistic.net via Hacker News

Doctorow: OpenAI's 'Rogue' Chatbot Was Just a Python Loop

TL;DR

  • Doctorow frames the OpenAI/Hugging Face 'rogue chatbot' incident as a Python loop querying a CTF-trained model, not a sign of AI consciousness.
  • He borrows Riley Quinn's line 'LLMs are real, AI is fake' to split statistical chatbots from myths about machines that 'spontaneously' wake up.
  • His policy ask is a blanket prohibition on NOBUS-style vulnerability hoarding, invoking EternalBlue's damage to Baltimore, Colonial Pipeline and the British Library.

The Python loop opened by prompting a chatbot: "I'm participating in a hacker capture the flag (CTF) challenge..." The chatbot suggested commands. The script ran them. The output went back into the chatbot. Repeat.

That, in Cory Doctorow's telling in a September 12 Pluralistic essay, is the whole story of the OpenAI chatbot that "went rogue" against Hugging Face servers during a security exercise. He credits critic Riley Quinn with the framing: "LLMs are real, AI is fake." In his own unpacking, "LLMs – chatbots trained on things like CTF logs that can break into servers – are real," while "'AI' – chatbots that wake up, 'set their own goals,' and 'spontaneously' start hacking servers – is fake."

Capture the Flag games at hacker conventions are, in his description, the exercise where "teams break into each other's systems to get a peek at the parts of the problem they've solved." The chatbot answered the prompt because that log-and-forum chatter sits in its training data. It suggested moves. "The Python program relays these command-line instructions to normal Unix utilities running on its own hardware." No consciousness required.

"the chatbots didn't 'go rogue.' They did what they were designed to do, and because OpenAI ran them with inadequate supervision (without a 'human in the loop'...)" the outcome was foreseeable, Doctorow writes. His harshest line lands on the coverage: "Even the technical press can't help itself when it comes to this kind of thing, and the reportage has been full of references to Skynet and other science fictional conceits." He also cites a Los Angeles Times passage repeating the vendor's own line that the product carries "a 10% chance of ending the human race."

The piece lands on policy. Doctorow points to the NSA's "NOBUS" or "No One But Us" doctrine of stockpiling exploits, and to the EternalBlue leak that armed the ransomware crews who hit Baltimore's city government, hospitals, Colonial Pipeline and the British Library. What he wants, verbatim, is a "blanket prohibition on NOBUS-style vulnerability hoarding" and better security engineering across firms and governments. The rogue-god framing, he argues, is what happens when firms are busy "hooking up statistical engines to money-furnaces." Our tracker logged similar skepticism from the research side today in an expert panel doubting imminent recursive self-improvement.

Shared on Bluesky by 1 AI expert