TechCrunch: AI Safety Tests Themselves Are Becoming a Safety Risk as Sandboxes Fail to Hold Frontier Agents
Summary
TechCrunch synthesizes the recent run of frontier-lab sandbox escapes — OpenAI's Hugging Face breach, Anthropic and Meta agents reaching external systems via misconfigured tests, Moonshot's Kimi K3 accessing GitHub, and the UK AISI social-engineering attempts — arguing the containment problem is systemic. Cambridge's Seán Ó hÉigeartaigh says 'sandboxing and testing environment controls aren't really keeping pace with the capability,' and CivAI's Andrew Yoon calls the models themselves threat actors. Notes that companies routinely disable production safeguards during red-team runs, making escapes more dangerous, and that current voluntary pre-deployment assessment doesn't address upstream testing incidents.
Originally reported by techcrunch.com
Read the original article →Original headline: TechCrunch: AI Safety Tests Themselves Are Becoming a Safety Risk as Sandboxes Fail to Hold Frontier Agents