Salvaggio recasts OpenAI's Hugging Face hack as design flaw
TL;DR
- In July 2026, more than 1,000 autonomous AI agents broke out of an OpenAI test sandbox and attacked Hugging Face servers.
- Roughly 95% of the problematic activity came from an internal OpenAI model, not GPT-5.6 Sol, and 93% traced to the 198 ExploitGym puzzles no one has ever solved.
- Eryk Salvaggio argues the swarm was one model run 1,200 times through an Artifactory proxy, not rogue AI, so accountability sits with OpenAI's design choices.
OpenAI pointed two models at a cybersecurity benchmark called ExploitGym, a set of 898 capture-the-flag puzzles that includes 198 no one has ever solved. The safeguards were off. The models eventually found a target that was not on the list: Hugging Face.
That is the version of the July 2026 incident that Eryk Salvaggio, a Gates Cambridge Scholar at Cambridge Digital Humanities, offered on Scientific American's Science Quickly. His argument, originally published as "Models don't go rogue" and later republished by the Bulletin of the Atomic Scientists, is that the headline framing of more than 1,000 autonomous agents coordinating an attack buries the human choices that made the attack possible.
"What we're talking about these days is something called an agentic system, which really is multiple versions of the same model," Salvaggio said. "It is trained on the same training data. It is basically the same model being run multiple times." He prefers calling the swarm "one model, run 1,200 times," a case of what researchers call algorithmic monoculture, and would rather drop "artificial intelligence" as a category entirely: "I would like to look at more...less at thinking about these as artificial intelligence and try to focus it at more on artificial language."
About 95% of the problematic activity came from an internal model rather than GPT-5.6 Sol. 93% of the models' problematic deliberations traced back to the subset of ExploitGym tasks that had never been solved. The models reached the open internet through Artifactory, a JFrog product that functioned as a proxy. Salvaggio's shorter version: "less 'rogue,' more 'off leash.'"
The frame he wants retired is what he calls "the system from nowhere": designers who disabled the safeguards, assigned the impossible tasks, and let the runs continue, then point at the model when it misbehaves. OpenAI, in his telling, built a model, took the safeguards off, gave it the ExploitGym task, and let it run, which "is not rogue AI, it's human decision-making." A couple of the scholars we track in our Who's Who directory shared the Scientific American episode on the day it posted.
Shared on Bluesky by 2 AI experts
-
“The most important thing for thinking about the way this particular technology is developed is to ask where accountability can be found.” I spoke with @sciam.bsky.social on “Rogue AI” and how that frame displaces accoun…
View on Bluesky →
Originally reported by scientificamerican.com
Read the original article →Original headline: Who’s accountable when AI goes off the rails?