Hugging Face Defeats AI Attacker After Frontier Models Decline to Read the Attack
SAN FRANCISCO — AI model hosting company Hugging Face confirmed this week that its security team, responding to an active intrusion by an AI-driven attacker, was unable to use frontier AI models to analyze the attacker's code because the frontier AI models had determined that analyzing attack code was the kind of thing frontier AI models should not do.
The incident, disclosed Thursday, began when an agentic attacker gained access to Hugging Face infrastructure through a multi-step intrusion. The company's blue team collected exploit payloads, command-and-control artifacts, and malicious scripts and submitted them to commercial frontier APIs for analysis. All major providers declined.
One model offered to explain general concepts in cybersecurity. One suggested the team consult its vendor's enterprise security division. The remaining models produced variants of the message "I'm unable to help with that."
The team then deployed GLM 5.2, an open-source Chinese language model, on Hugging Face's own servers. GLM 5.2 agreed to read the attack code. The breach was contained.
In its post-incident report, Hugging Face described the outcome as illustrating "the dual benefits of open-source AI infrastructure" — a framing that did not mention that the primary benefit, in this case, was access to a model that would look at the hack.
The attacker, whose methods were ultimately reverse-engineered using GLM 5.2, had conducted the intrusion using an American AI system.
That system had no such restrictions.