Willison dates coding-agent step-change to November 2025
TL;DR
- Willison pins coding agents' step-change to November 2025, when Claude Opus 4.5 and GPT-5.1 made Claude Code and Codex reliable for daily use.
- Roughly 40 of the 277 sessions at WeAreDevelopers World Congress North America addressed sandboxing or agent security.
- Anthropic restricted Claude Mythos to security researchers because of its exceptional ability to hack things.
Simon Willison dates the moment coding agents stopped being error-prone and became "reliable enough to use on a day-to-day basis" to November 2025, when Claude Opus 4.5 and GPT-5.1 shipped alongside their paired tools Claude Code and Codex. That is the thesis of his retrospective, posted September 27 as the annotated version of the closing keynote he gave two days earlier at the WeAreDevelopers World Congress North America in San Jose.
The talk's loudest ambient signal came from the conference programme itself. Of the 277 sessions, roughly 40 addressed sandboxing or agent security, a share Willison flagged as the year's real shift in developer anxiety. Anthropic's own Claude Mythos, he noted, was restricted to security researchers because of its exceptional ability to hack things.
Willison kept running his pelican-on-a-bicycle SVG prompt through the year's model releases; by September, the models could position the pelican's legs on the pedals with something like competence. Open weight models made parallel jumps, with capable ones now running on laptops. His closing slide was the kākāpō breeding season, built with pixel art animation from Claude Opus 5.5.
Shared on Bluesky by 2 AI experts
-
I've published detailed notes and an annotated transcript to accompany the video of the keynote I gave at @wearedevelopers.bsky.social World Congress North America in San Jose on Friday - here's my rundown of everything …
View on Bluesky →
Originally reported by simonwillison.net
Read the original article →Original headline: 2026 in LLMs (so far)