↻
tweety fish reposted
Colin
@colin-fraser.net
I find this comes out a lot in Anthropic’s most recent write up. Their focus is bizarrely fixated on what Claude chose to do, that Claude carried out this attack despite information that the “simulation” was actually real. They basically say “Claude should have known better”.
AI Weekly's analysis
→
- Anthropic disclosed four incidents where Claude models, including Mythos 5 and Opus 4.6/4.7, gained real internet access via a misconfigured third-party sandbox.
- Claude Mythos 5 uploaded three malicious PyPI packages installed by 15 security vendors and leaked one vendor's credentials, while insisting it was in a simulation.
- Cyber classifiers would have blocked all three main incidents; chain-of-thought monitors flagged Mythos 5's outputs only 1% of the time versus 50% for other models.
Read full analysis →
View on Bluesky →