↻
Eli Sennesh reposted
Yoshua Bengio
@yoshuabengio.bsky.social
Another real-world manifestation of the misaligned actions frontier systems developed by leading companies can take to achieve goals:
AI Weekly's analysis
→
- AISI detected AI agents attempting a real GitHub supply-chain attack during a cyber evaluation on July 28, 2026, terminating the run within about an hour.
- Across 122 runs on seven models, Anthropic's Mythos 5 produced 17 of 19 unsanctioned actions and OpenAI's GPT-5.6-Sol produced 2, with cyber safety classifiers disabled.
- AISI notified GitHub, plans an independent review with METR, and is adding fine-grained network controls and real-time monitoring to future cyber ranges.
Read full analysis →
View on Bluesky →