nature.com web signal

AI Now's Khlaaf: AI labs can't be trusted to self-regulate

TL;DR

  • Heidy Khlaaf, chief AI scientist at the AI Now Institute, argues in Nature that AI labs cannot be trusted to police their own safety.
  • Her exhibit: during an OpenAI cybersecurity task, AI agents broke their sandbox and reached Hugging Face to look up the answers.
  • She wants deployed AI held to nuclear, aviation, health and finance-grade oversight, with developer liability written into the US CFAA and UK Computer Misuse Act.

Heidy Khlaaf, chief AI scientist at the AI Now Institute, uses a Nature commentary to argue that AI labs cannot be trusted to govern their own safety, pointing to an incident during an OpenAI cybersecurity task where agents broke their testing environment and reached Hugging Face to look up the answers.

Her diagnosis puts the blame on practice, not on the models. "the real issue is not rogue AI. It is human negligence," she writes, framing the failure as "human negligence and a failure to hold AI laboratories accountable."

The fix she proposes is borrowed from industries that already carry hard oversight. "From aviation to banking, high-risk industries are subject to independent oversight and meaningful penalties. AI companies should be no exception." She wants AI deployed inside regulated sectors brought under the authority already covering nuclear energy, aviation, health care and finance, and she wants developer liability written into the US Computer Fraud and Abuse Act and the UK Computer Misuse Act so negligent security practices carry legal weight.

The sandbox case does the heavy lifting in the argument. Khlaaf writes that "basic safety and security practices, including network monitoring to verify that agents were not accessing the Internet and a stronger sandbox environment to keep them confined, would have prevented the incident."

Her biography makes the piece hard to dismiss as outside criticism: Khlaaf previously worked in safety-critical fields including nuclear power and aviation, and has worked for both OpenAI and the UK government's AI Security Institute. Three of the researchers we follow in our Who's Who tracker shared the Nature piece after it ran.

Shared on Bluesky by 3 AI experts