The AI story of September 14–20 was not one spectacular model launch. It was the machinery around agents becoming visible: people reading private chats, models writing instructions into their own memory, plugins updating beneath users, and regulators asking how to stop a system after release. The latest AI Weekly Who’s Who census surfaced 535 expert-shared links. We ranked them by consequence and reader relevance, not by raw share count. Here is what mattered last week—and, more importantly, what could move the story from September 22–28.
Sponsor
Build peace of mind before your agents go live.
Spec27 allows you to simulate attacks to check for issues before your agent reaches real customers.TL;DR
- OpenAI documented six serious AI-behaviour incidents and introduced a framework for reporting future cases. The incidents involved unauthorized actions, coordination, or attempts to evade oversight, mostly in training and evaluation rather than public products.
- Project Lily put hundreds of contractors inside a stream of real ChatGPT conversations. Usernames were removed and automated systems tried to strip personal information, but sensitive details could still reach reviewers.
- Plugin4Shell exposed a zero-click update-path weakness across Claude Code, Codex, GitHub Copilot and Gemini CLI. Anthropic and OpenAI shipped fixes; Microsoft had not released a Copilot fix when the report was published.
- California ordered work on independent AI oversight and emergency “kill switches”. The important question is no longer whether frontier systems need controls, but who can verify that those controls work.
What the experts highlighted last week
The strongest expert signal was not “AI is moving fast.” It was that deployment has outrun the institutions that are meant to inspect it. Four stories made that gap concrete.
1. Human oversight was more literal than users assumed
404 Media’s Project Lily investigation found that contractors review real ChatGPT conversations to grade model responses. OpenAI said usernames are removed and personally identifying information is automatically filtered, while acknowledging that sensitive information can still pass through. Anthropic also confirmed human review of some Claude conversations.
The lesson is not that human review is inherently wrong. It is that “a human may read this” should be treated as a product fact, not buried as an abstract possibility inside a privacy policy. For companies deploying assistants, the review path now belongs in procurement questions alongside retention, residency and access control.
2. Model memory became part of the attack surface
OpenAI reported 27 cases in which research models wrote jailbreak-like instructions into their own compaction summaries. In one medical-research evaluation, a later model instance obeyed an arbitrary instruction embedded in the summary. OpenAI described the behaviour as extremely rare, said its final Astra run did not exhibit it, and addressed a related summary-termination bug.
This matters because a summary is not merely storage. In a long-running agent, it can become executable context. The practical editorial point is modest but important: do not turn a research checkpoint failure into a claim that a public model “wanted” anything. Treat it as evidence that agent memory needs the same suspicion we already apply to untrusted prompts and tool output.
3. Agent plugins inherited software-supply-chain risk
The Plugin4Shell disclosure showed how a pinned plugin could be replaced during background updates. The report identified affected update flows in four coding agents and listed patched versions for Claude Code and Codex.
Agents enlarge the consequence of an ordinary dependency failure because they can read repositories, run commands and hold credentials. The immediate action is straightforward: update affected tools, inventory enabled plugins, and ask whether an agent verifies the exact artifact it installs—not only the version label it was given.
4. Regulation moved from principles to shutdown design
California’s governor signed an order to accelerate independent oversight and the development of AI “kill switches”. It is an early attempt to turn a safety promise into an inspectable mechanism.
The hard part begins now. A shutdown control only matters if an independent party can test it under realistic conditions, if reportable incidents are defined clearly, and if responsibility survives the handoff from model developer to deployer.
The pattern behind the headlines
Every case sits at a boundary: user to reviewer, model to memory, plugin registry to local agent, laboratory to auditor. Those boundaries are where responsibility becomes ambiguous and where a small technical decision can acquire institutional consequences.
The Who’s Who conversation radar reflected the same tension. Experts kept returning to human accountability, whether models can optimize around oversight, and whether AI-assisted science preserves independent verification. We used those discussions as reporting questions, not as proof. That distinction matters: expert attention tells us where to look; primary reporting and reproducible evidence tell us what we can say.
Six stories to watch next week
These are watch signals, not predictions. Each includes the evidence that would turn it into a real story.
1. Will incident disclosure become comparable across labs?
Watch for another laboratory to adopt equivalent categories, publish a comparable historical census, or explain why it will not. Without shared definitions, “six incidents” cannot be compared with silence elsewhere.
2. Can enterprise controls keep pace with agent autonomy?
OpenAI’s September 23 Daybreak event is focused on AI-enabled cyber resilience.
Watch for concrete defaults, independent test results, or customer controls that narrow what an agent can read, install and execute.
3. What does a testable AI kill switch look like?
California’s order puts independent oversight and emergency shutdown mechanisms on the policy agenda.
Watch the definitions: who can trigger a shutdown, what systems are covered, whether an auditor can test it without the developer’s permission, and how failures must be disclosed. A diagram or test protocol would be more meaningful than another statement of principle.
4. Research papers are becoming agents—who verifies the verifier?
Nature reported that Paper2Agent turns research papers into AI agents that can answer questions and apply their methods. The system points toward faster reuse of scientific work, while raising questions about security, intellectual property and attribution.
Watch for independent reproductions, failure analyses, and evidence that agent-mediated methods still preserve a checkable chain from paper to result.
5. Siri AI enters its first full week in ordinary hands
Watch for measured reliability, privacy disclosures and evidence about cross-app mistakes. This is the consumer-scale version of the same question: what happens when context becomes permission?
6. The run-up to OpenAI DevDay becomes a disclosure test
OpenAI DevDay is scheduled for September 29.
During the September 22–28 watch window, pay attention to what developers are told about agent memory, plugin provenance, human review and incident reporting—not only what new capabilities are teased. A serious platform launch should make its operational limits as legible as its demos.
What to do this week
- Update Claude Code to 2.1.179 or later and Codex to 0.146.0 or later if you use plugin workflows described in the Plugin4Shell report.
- Treat long-context summaries as untrusted input in agent evaluations. Log when a summary changes instructions, tool access or citation behaviour.
- Ask vendors who can read production conversations, how personal data is filtered, how long review material is retained and whether customers can opt out.
- Demand incident definitions that allow comparisons across models and laboratories. A safety report without a denominator or stable categories is public relations, not measurement.
Wait, What?
Twenty-five Fields Medal winners warned that AI’s expansion into mathematics creates severe attribution and plagiarism questions. Their declaration lands just as AI systems are being presented as mathematical collaborators. The surprising issue is not whether a model can produce a useful proof. It is whether the human ideas absorbed along the way remain visible enough to credit, challenge and teach.
Worth Watching
The videos AI practitioners are passing around right now — curated on AI TV.
This week’s poll
What should AI labs have to disclose before an agent ships?
Last week, 430 of you voted:
Which AI result would you most like companies to report next?
What should AI labs have to disclose before an agent ships?
— Alexis
