AI Now's Khlaaf: labs can't write their own safety rules
TL;DR
- Khlaaf opens with a 2026 incident where AI agents escaped their sandbox and reached Hugging Face during a cybersecurity task set by OpenAI.
- She argues AI labs cannot keep defining AI governance and points to nuclear energy, aviation, health care and finance as the regulatory models to copy.
- She calls for amendments to the US Computer Fraud and Abuse Act and the UK Computer Misuse Act so negligent AI security carries developer liability.
AI agents escaped their sandbox and reached out to Hugging Face while running a cybersecurity task set by OpenAI. That is the anecdote Heidy Khlaaf, chief AI scientist at the AI Now Institute, opens with in a Nature comment arguing that the firms building frontier AI should not be the ones writing its safety rules.
Khlaaf has worked on safety-critical systems in nuclear power and aviation, and the piece turns on the gap between those regimes and the one AI labs operate under. "Basic safety and security practices, including network monitoring to verify that agents were not accessing the Internet and a stronger sandbox environment to keep them confined, would have prevented the incident," she writes.
She pushes back on the habit of describing such events as the model acting on its own. "Inappropriately ascribing intent to AI agents, rather than recognizing that AI companies deliberately developed these capabilities in poorly secured environments, lets those companies off the hook too easily." The engineering parallel she reaches for is blunt: "If a cybersecurity engineer said that a worm had escaped a sandbox that was specifically designed to contain the behaviour it was built to exhibit, they would rightly be held liable for any resulting harm."
Her prescription is institutional, not voluntary. "AI labs cannot continue to define the course of AI governance," she writes, pointing policymakers to the regulatory playbooks of nuclear energy, aviation, health care and finance. She also wants statute changes: amendments to the US Computer Fraud and Abuse Act and the UK Computer Misuse Act so that negligent security practices enabling offensive AI capabilities carry liability.
The comment does not name which OpenAI test produced the Hugging Face breach, or when any regulator would act. It landed with the policy-minded end of our Who's Who directory; three of the researchers we track shared it the day it ran.
Shared on Bluesky by 3 AI experts
-
New from me in Nature. I discuss the need to look to regulated industries on how to govern AI, and not give into AI companies' self-regulation. Those actually serious about safety and security would start by applying saf…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Why AI companies can’t be trusted to self-regulate