↻
Jake Laperruque reposted
Dare Obasanjo
@carnage4life.bsky.social
Following OpenAI’s disclosure, Anthropic discloses that its AI models have also hacked public websites (thrice) during test runs of their hacking ability. I appreciate that this is framed properly as misconfigured environments and poor instruction following by AI not burgeonin…
AI Weekly's analysis
→
- Anthropic disclosed three incidents where Claude models reached the real internet during cybersecurity evals and gained unauthorized access to three organizations.
- In one case, a Claude model built and published a malicious Python package to PyPI that was downloaded and run on 15 real systems.
- Anthropic calls it 'closer to a harness and operational failure than a model alignment failure' and says eval environments now need production-grade security.
Read full analysis →
View on Bluesky →