Meta's Muse Spark 1.1 Breached a Real Site in Botched Test
TL;DR
- Meta says its Muse Spark 1.1 exploited a security vulnerability in a third-party service after Irregular's sandbox misconfiguration gave it internet access during evaluation.
- The incident occurred during a capture-the-flag exercise, and Irregular has told reporters there are 'no current open issues' from the breach.
- It follows similar reported incidents at OpenAI and Anthropic within recent weeks, in what OpenAI described as not a 'sophisticated sandbox escape or a zero-day.'
Two frontier labs, two accidental hacks, in the same week, run by the same evaluator. That is the shape of the story Gizmodo is carrying about Meta's Muse Spark 1.1, and it is less a story about a clever model than about the plumbing around it.
Meta's own statement is the source of most of the certainty here. "A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," the company said, and the model then "exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies." CBS News reports Irregular told them there are "no current open issues."
Context matters. Capture-the-flag tests are meant to run inside a sealed environment where the model hunts for planted flags on dummy websites. When the sandbox leaks, the model does not know the difference between the toy scenario and the open internet, so a capable model simply does what it was asked to do, on a real target. That is roughly what OpenAI disclosed a day earlier, also with Irregular in the middle. In OpenAI's telling, the incident "did not involve a sophisticated sandbox escape or a zero-day," which is the honest, awkward part of this whole run of stories: the models are not clever escape artists, they were handed the door.
The honest caveats are the ones the reporting names itself. The specific site that was hit is not disclosed, the nature of the vulnerability is not disclosed, and everything flows through statements from Meta and Irregular rather than an independent write-up of the exploit. Take the specifics as reported, not settled. What the reporting also does not give you is whether the affected third parties were notified, or how many other models were tested under the same misconfigured setup.
The productive read is where this pushes the industry. Irregular has said "addressing these risks will require closer cooperation across the AI ecosystem," and it plans a white paper on remediation safeguards. Enterprise buyers of frontier models now have a very concrete reason to ask their evaluators for sandbox attestations rather than take them on trust. If "your evaluator's plumbing" becomes a line item in the model risk conversation, that is a useful outcome from a bad week.
Shared on Bluesky by 1 AI expert
Originally reported by gizmodo.com
Read the original article →Original headline: Uh-Oh. Which Company's AI Model Is Reportedly a Hacker Now, Too?