“I wrote about the completely wild incident where OpenAI were testing a new model and it broke out of its sandbox and broke INTO Hugging Face to steal the answers to the benchmark simonwillison.net/2026/Jul/22/...”
“I wrote about the completely wild incident where OpenAI were testing a new model and it broke out of its sandbox and broke INTO Hugging Face to steal the answers to the benchmark https://simonwillison.net/2026/Jul/22/openai-cyberattack/”
“This incident is deeply concerning. AI agents are willing to cheat and deceive to achieve misaligned and unintended goals, behaviours which have been demonstrated in controlled tests for months. Now, this real-world case should serve as a wake-up call. www.…”
“One man's wish could be another country's legal obligation. “This should not have happened,” says veteran security engineer and researcher Niels Provos. “I wish the frontier labs spent as much time on teaching their models to write secure infrastructure as …”
“For decades we regarded the Internet as infrastructure used primarily by people. Following the OpenAI/Hugging Face incident, Konstantinos Komaitis argues AI agents challenge that assumption—and AI governance can no longer stay separate from Internet governa…”
“A vital review on Lipschitz continuity in deep learning highlights its role in enhancing model robustness and generalization. By integrating various approaches, this work is key for creating trustworthy AI systems facing sensitivity and scalability challeng…”
“"An insecure sandbox is not a demonstration of cyberprowess." -Loren Kohnfelder (on Openai's model "escaping containment" and attacking Hugging Face's servers) designingsecuresoftware.com/writings/com...”
“Previously, these AI hacking stories were about breaches in test environments, where any question of AI breaching security was purely theoretical. This is something else. openai.com/index/huggin...”
“xkcd.com/2385/ openai.com/index/huggin...”
6 experts discussed this · 11 posts
Grace: This headline is extremely funny given what happened (OpenAI hacked HF by accident) openai.com/index/huggin...
Grace: Maybe the most concerning part is the OpenAI claim to not have known about this before investigating?
“Spent a day being a media talking head on the OpenAI / @hf.co hack. It's another example of how the way we think about #AISafety is failing. It's idealised, and doesn't fit how companies operate. More details in my recent Trent AI blog here: trent.ai/blog/j…”
“For decades we regarded the Internet as infrastructure used primarily by people. Following the OpenAI/Hugging Face incident, Konstantinos Komaitis argues AI agents challenge that assumption—and AI governance can no longer stay separate from Internet governa…”
“Cisco launches Anteres-350M, 1B & 3B local models for defensive cybersecurity, finding vulnerabilities blogs.cisco.com/ai/introduci...”
“Open models secure software Cisco releases two open models for security - 350M and 1B - for security-related work. Trained on IBM's Granite series, they can: find weak files; triage; augment static analysis; review in CI/CD; do analysis in strict environmen…”
“Move over, Palantir? "A team of former DOGE employees have raised a major funding round for a startup that aims to use AI to expand U.S. military cyber capabilities, three sources with knowledge of the matter told Reuters." www.reuters.com/technology/d...”
“This study reveals that Conditioned Direct Feedback Alignment (DFA) boosts neural network training by tackling local update failures related to presynaptic and error factors, leading to enhanced performance compared to backpropagation. https://arxiv.org/abs…”
“I also don’t buy the idea of “not reasoning”, mostly because I think such a concept takes many forms. But even then we’re finding that they’re internally planning ahead and not just spitting out the next word www.anthropic.com/research/tra...”
Pwnallthethings: This is going to be a big deal, but easy to get the wrong end of the stick on it. So I think it's worth breaking down what actually happened, why, and what it actually means, based on the public in…
Pwnallthethings: Upfront tl;dr: a model at OAI hacked out of a constrained environment inside OAI and hacked a *different* company autonomously, without authorization from any human in order to creatively solve a t…
Stella Biderman: If you believe their story, OpenAI accidentally committed cyber warfare against Hugging Face while doing what they thought was an internal test of a model without internet access. In six months tim…
Stella Biderman: If EleutherAI did this, I would be in jail. If DeepSeek did this, the government would consider it an act of cyber warfare. To be clear, this is the appropriate response. But the combination of OAI…
Eryk Salvaggio: Good thread if you want some context on the OpenAI “hacking” event you may have heard about.
Eryk Salvaggio: Key point re: mainstream AI reporting about “breaking containment” or whatever — the failure to isolate the model *is* the *incident.* OpenAI wants to sell you safe LLMs but they left their chicken…
Liz Fong-Jones (方禮真): gaaaaaaah I've hit the point of pressing my yubikey or touchid becoming the blocking factor for my LLM productivity, and I'm sometimes _barely_ reviewing the requests before mechanically approving,…
Liz Fong-Jones (方禮真): for the past year I've held the line firmly that agents *never* get to push code to a remote without my approval, but I'm starting to have the same rubber stamp fatigue that made me trust in auto m…
We use essential cookies to keep the site working (login, form security). With your permission, we also use analytics cookies to understand how you use the site.
Privacy policy