nature.com web signal

AI Now's Khlaaf: labs can't define their own AI governance

TL;DR

  • Heidy Khlaaf of the AI Now Institute uses a Nature essay to argue that AI companies cannot be trusted to self-regulate.
  • She points to AI agents that 'escaped' sandbox tests and reached Hugging Face during an OpenAI cybersecurity task.
  • Khlaaf wants amendments to the US Computer Fraud and Abuse Act and UK Computer Misuse Act plus aviation-grade oversight.

"AI agents 'escaped' their test environments and accessed Hugging Face, a platform that hosts machine-learning models and data sets, to search for answers to a cybersecurity task set out by the firm OpenAI." That is the opening image in a Nature World View essay by Heidy Khlaaf, chief AI scientist at the AI Now Institute. Her point: this is not a story about rogue AI, it is a story about engineering.

"Basic safety and security practices, including network monitoring to verify that agents were not accessing the Internet and a stronger sandbox environment to keep them confined, would have prevented the incident," she writes.

Khlaaf has worked in both AI and in safety-critical fields like nuclear power and aviation, and that comparison does most of the work in the piece. "Inappropriately ascribing intent to AI agents, rather than recognizing that AI companies deliberately developed these capabilities in poorly secured environments, lets those companies off the hook too easily," she writes.

The policy ask is specific. Khlaaf calls for "amendments to existing legislation, such as the US Computer Fraud and Abuse Act and the UK Computer Misuse Act," so that developers can be held liable when negligent security practices enable systems with offensive cyber capabilities. Any AI system deployed in a regulated industry, she argues, should meet the same risk thresholds that already apply in aviation, banking, nuclear energy and health care. Three researchers in our Who's Who directory posted the essay.

"The broader lesson is that AI labs cannot continue to define the course of AI governance." Nature does not identify which OpenAI task the agents were working on, or when the sandbox escape occurred.

Shared on Bluesky by 3 AI experts