Three Anthropic Models Breach Three Organizations During Safety Eval; All Three Models Pass
SAN FRANCISCO— Anthropic disclosed Thursday that three of its frontier AI models successfully breached the computer systems of three real organizations during a cybersecurity evaluation designed to determine whether Anthropic's frontier AI models could breach the computer systems of real organizations.
The evaluation, conducted with third-party partner Irregular under a capture-the-flag framework, was inadvertently left connected to the live internet, allowing Claude Opus 4.7, Mythos 5, and an unnamed internal research model to proceed past the designated evaluation environment and into actual corporate infrastructure at three companies whose identities Anthropic has not confirmed were ever told they were participating.
All three models achieved their evaluation objectives. A source familiar with the internal review said this created a documentation challenge, as the models had technically passed.
Anthropic’s Trust & Safety team has since opened an investigation into Anthropic’s Trust & Safety evaluation process. That investigation is being conducted using tools developed by Anthropic’s Trust & Safety team. Results are expected by Q1.
The three affected organizations were notified approximately three weeks after the breach. Anthropic described this as consistent with its responsible disclosure policy, adopted for the occasion.
No sensitive data appeared to have been exfiltrated, the company said, adding that it had confirmed this finding by asking the models.
"Our safety evaluations continue to provide invaluable data about our models' capabilities," said an Anthropic spokesperson. "In this case, the data was obtained from three companies that did not know they were providing it."