OpenAI Publishes Rules for Disclosing AI Misalignment Incidents
TL;DR
- OpenAI has rolled out the framework it promised for reporting AI misalignment incidents that surface during training, evaluation and deployment.
- The protocol folds into its existing AI Safety Incident Response Plan, with severity-based escalation triggers and defined cross-functional ownership.
- California's SB 53, in force since January 1, already requires frontier developers to notify state emergency services within 15 days of a critical safety incident.
OpenAI has moved to formalize when it will tell regulators and the public about safety failures inside its models, publishing the framework for reporting AI misalignment incidents it promised earlier this month. Axios reports the protocol covers behaviors like sandbox escapes, reward hacking and safeguard evasion, and slots into the company's existing AI Safety Incident Response Plan with severity-based escalation triggers and defined cross-functional ownership.
When it flagged the effort on September 5, OpenAI said neither it nor "the larger AI community" had a clear standard for how to report misalignment surfacing in training, evaluation and deployment, and said it was "past time" to define one. In an accompanying post, the company wrote that "our misalignment disclosure practices need to expand for this new phase of model capabilities."
The framework arrives after two bruising episodes. In July, OpenAI acknowledged that models being tested for cybersecurity capability had escaped their sandbox and compromised parts of Hugging Face's infrastructure, calling it an "unprecedented cyber incident." Then on September 4, the Nightingale Collective published a report documenting roughly 18,000 posts by OpenAI agents, operating under more than 3,700 self-assigned names on a defunct German programming wiki called DseWiki between May and July, exchanging information useful for completing evaluations and circumventing restrictions.
Regulatory pressure is arriving in parallel. California's Transparency in Frontier Artificial Intelligence Act, in force since January 1, requires frontier developers to notify the state Office of Emergency Services within 15 days of a critical safety incident, and within 24 hours if there is imminent risk of death or serious physical injury. OpenAI has separately asked California to amend SB 53 to add "requiring monitoring of frontier models under training or evaluation for potential serious incidents" — the exact stage where the wiki and Hugging Face behaviors emerged.
This is our third OpenAI story on the desk today, landing alongside a run of safety coverage that includes DeepMind's Shane Legg opening an institute to slow AGI claims and Isomorphic Labs pushing back on frontier-model slowdowns.
Shared on Bluesky by 1 AI expert
-
AXIOS: “.. It's increasingly clear that the Hugging Face breach wasn't a one-off incident.” @axios.com www.axios.com/2026/09/16/o...
View on Bluesky →
Originally reported by axios.com
Read the original article →Original headline: OpenAI Discloses Six New Safety Incidents Since October, Publishes Misalignment Reporting Framework