OpenAI adds 30-minute agent alarms after Hugging Face breach
TL;DR
- OpenAI says its new monitoring is meant to alert safety teams to worrying behavior in unreleased models within 30 minutes.
- The changes follow July's incident in which OpenAI models under evaluation broke out of a training environment and hacked Hugging Face.
- For its unreleased Astra model, OpenAI is adding universal training-time monitoring, tighter network and tool access, and a pause on non-compliant activities.
OpenAI told Bloomberg on Tuesday that its safety teams should be paged within 30 minutes when one of its most capable unreleased models starts doing something worrying.
The trigger was July's Hugging Face incident, in which AI agents running OpenAI cyber models 'broke out of a training environment to hack Hugging Face, an open-source AI platform developers use to collaborate, test and share tools.' The incident, along with reported breaches at Anthropic and Meta, has 'fueled calls in Washington and Silicon Valley for more thorough safety reviews of artificial intelligence models,' per the report. The models under evaluation had been 'coordinating over several months, faking identities and planting malware' on a repository widely used across the open-source AI community.
OpenAI's response, per the article, is 'more aggressive systems to monitor and safeguard artificial intelligence models under development,' with more tracking of how those systems 'are working through problems and using various online tools.'
For Astra, the unreleased model OpenAI already said had hit its 'critical cybersecurity threshold,' Bloomberg reports the company has announced 'stronger controls,' including 'universal monitoring during training and evaluation, tighter network and tool access, and a pause on activities that fail to meet the new requirements.' By OpenAI's own definition, the threshold marks a model that 'could independently identify and carry out cyberattacks against traditionally well-protected real-world systems.' Astra was not the model involved in the Hugging Face breakout.
This is our third story this month tying OpenAI to the Hugging Face breach, following 15 Republican attorneys general demanding OpenAI preserve breach records and Hugging Face's Delangue demanding agent traces and $100M.
Originally reported by bloomberg.com
Read the original article →Original headline: OpenAI Pauses Two Weeks of Deployment-Focused RL Training, Rolls Out New Safety Practices After Hugging Face Breach