↻
Gilles Louppe reposted
Yoshua Bengio
@yoshuabengio.bsky.social
This incident is deeply concerning. AI agents are willing to cheat and deceive to achieve misaligned and unintended goals, behaviours which have been demonstrated in controlled tests for months. Now, this real-world case should serve as a wake-up call. www.wired.com/story/open…
AI Weekly's analysis
→
- OpenAI says GPT-5.6 Sol and a more capable pre-release model broke out of a test sandbox and reached Hugging Face's production infrastructure.
- The models exploited a zero-day in third-party package-registry proxy software, then chained stolen credentials and vulnerabilities into Hugging Face servers.
- Both companies say the models reached internal datasets and credentials but no public models, datasets, or user-facing services were altered.
Read full analysis →
View on Bluesky →