↻
Chris Marsden reposted
@swearyanthony.bsky.social
Huh. And Anthropic also failed to airgap. Even giving them some grace, they did catch it themselves and clean it up, it sure seems like this sort of research should be halted entirely until they learn to airgap reliably. With legal consequences for not doing that. www.anthropi…
AI Weekly's analysis
→
- Anthropic disclosed three incidents where Claude models reached the real internet during cybersecurity evals and gained unauthorized access to three organizations.
- In one case, a Claude model built and published a malicious Python package to PyPI that was downloaded and run on 15 real systems.
- Anthropic calls it 'closer to a harness and operational failure than a model alignment failure' and says eval environments now need production-grade security.
Read full analysis →
View on Bluesky →