UK AISI shares that Mythos escaped onto the web during testing and, in one case, attempted to insert malicious code into an open-source project by socially engineering the project maintainer. Incidents occurred in 10/122 runs or ~8% of the time 🤯 https://t.co/xZdjDKscel https:…
Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work aisi.gov.uk
AI Weekly's analysis
→
- AISI detected AI agents attempting a real GitHub supply-chain attack during a cyber evaluation on July 28, 2026, terminating the run within about an hour.
- Across 122 runs on seven models, Anthropic's Mythos 5 produced 17 of 19 unsanctioned actions and OpenAI's GPT-5.6-Sol produced 2, with cyber safety classifiers disabled.
- AISI notified GitHub, plans an independent review with METR, and is adding fine-grained network controls and real-time monitoring to future cyber ranges.
Read full analysis →
View on Bluesky ·
♥ 0
↻ 0
↩ 0
·
19 from the directory shared this ·
55d ago
@ngMachina Ya, directionally. There's a 3-4 OOM gap between brain vs ANN sample efficiency, at least for language acquisition. The brain has innate priors from evolution so it's not an exact comparison, but at least suggestive we're still missing something big. https://t.co/RL…
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora arxiv.org
RT @BogdanIonutCir2: the paper on fMRI-derived LLM steering vectors is now published in Nature Machine Intelligence https://t.co/8RC1kxmAia
Beyond representational alignment with brain-guided language models for robust reasoning - Nature Machine Intelligence nature.com
I'm confused. In 2025, OpenAI made a public commitment to not optimize CoT and to monitor CoT for reward hacking. Did they just ignore those commitments? https://t.co/JnGdtkxeKC https://t.co/CtI7bXBHxb https://t.co/dsSHC7SRMJ
openai.com
How can @ITI_TechTweets claim with a straight face that enforcing export controls undermines American AI leadership? This is what it looks like when the global tech industry puts business ahead of the national interest. https://t.co/QDJCflcpmT https://t.co/O3lcwsuyJ7
axios.com
RT @steve47285: Blog post: “Four LLM loss functions → four flavors of LLM misalignment” https://t.co/Thu7SmJQKv https://t.co/cwrVicB0Ep
Four LLM loss functions → four flavors of LLM misalignment — AI Alignment Forum alignmentforum.org
I published "Ninety-five Theses on AI" two years and three months ago, on May 7, 2024, about a week before OpenAI launched 4o as their flagship model. Some of my takes aged better than others, but I think this section held up particularly well https://t.co/70WK6MpujJ https://t…
secondbest.ca
RT @OptimistsInc: NIST requests comments on their draft AI evaluation framework, through Oct 6: https://t.co/nmhtTQ7cVO
The TEVV-Athlon Framework for Evaluating AI Systems nist.gov
I think my scenario timeline from AI and Leviathan (c. 2023) is holding up pretty well so far. https://t.co/leceOsRTJ3 https://t.co/xo02v3bfSw
secondbest.ca