I just realized that Opus-4.6, which in my head is "an old workhorse model, less powerful but *much* less pearl-clutching than its successors", is only 7 months old. If you had asked me before I looked it up I think I would have guessed "a little over a year maybe?" AI time = dog years.
The statement 'we do/do not understand how LLMs work' almost invariably confuses two very different things. On the one hand, we absolutely can map and describe, in great detail, every single mathematical operation that they perform to generate an is answer. But...
Worth remembering: the only reason AI agents can run bash commands (or do anything else, for that matter) is because we explicitly give them tools that can do so. Tools and harness capabilities are most of what make agents a security risk. Just remove them. Least capability = least privilege.
I can't express how much I hate having to review LLM-generated slop from 'writers' with no expertise in the area that they didn't bother reviewing at all. Overstated claims, incoherent framing, nonsequiters stuffed into lists, glaring technical errors that even a brief review would have caught.
Since recent events seem to have dragged AI powered biosafety back into the chat, thought I'd take the excuse to repost this. IYKYK.
Anecdotal and vibes, but Opus 5 and Fable both seem distinctly worse than GPT 5.6 at multistep reasoning outside of coding. The Anthropic ones also careen wildly between rank sycophancy and getting extremely pissy when you push back. Hard to trust, sometimes annoying to use.
Using a day off to engage in my own silly and probably irresponsible LLM experiments instead of stressing about securing the latest batshit insanity someone found on twitter and yeeted into prod, and it's honestly the most fun I've had in a hot minute. They're such weird little guys <laudatory>.
New LLM-doc tells that I'm beginning to twitch when I see them: - "Limitations stated [up front | plainly | honestly]:" and "The [finding | result | outcome] nobody expects:" as the start of a sentence. - "the [X] is the [finding | result | mechanism]" as a contrastive phrase
Stop giving the models bash/arbitrary code execution tools with full network access and 80% or so of your AI security problems get much more tractable. Sandbox file writes for the next 20%.
Search Google for MITRE ATLAS -- a terrible AI summary, four sponsored results, six suggested searches, a bunch of youtube videos, four more suggested searches, social media results (what?), and finally, an actual result. Followed by another sponsored result and six more suggested searches.