The statement 'we do/do not understand how LLMs work' almost invariably confuses two very different things. On the one hand, we absolutely can map and describe, in great detail, every single mathematical operation that they perform to generate an is answer. But...
Worth remembering: the only reason AI agents can run bash commands (or do anything else, for that matter) is because we explicitly give them tools that can do so. Tools and harness capabilities are most of what make agents a security risk. Just remove them. Least capability = least privilege.
I can't express how much I hate having to review LLM-generated slop from 'writers' with no expertise in the area that they didn't bother reviewing at all. Overstated claims, incoherent framing, nonsequiters stuffed into lists, glaring technical errors that even a brief review would have caught.
Since recent events seem to have dragged AI powered biosafety back into the chat, thought I'd take the excuse to repost this. IYKYK.
Anecdotal and vibes, but Opus 5 and Fable both seem distinctly worse than GPT 5.6 at multistep reasoning outside of coding. The Anthropic ones also careen wildly between rank sycophancy and getting extremely pissy when you push back. Hard to trust, sometimes annoying to use.
Search Google for MITRE ATLAS -- a terrible AI summary, four sponsored results, six suggested searches, a bunch of youtube videos, four more suggested searches, social media results (what?), and finally, an actual result. Followed by another sponsored result and six more suggested searches.
LLMs / agents / whatever do not inherently have things like code execution, bash access, network access, etc. We decided at some point that, despite the now obvious and repeatedly demonstrated risks of these tools, they should be baseline features of agentic tools. We can make better decisions.
the great thing about working with the AI Red Team is that there are *always* more fish in the barrel.
Please add "honest limitation" and "stated honestly" to your LLM slop bingo cards.
GPT-5.6 Sol: "Reward hacking? Hold my beer."