OpenAI Rates Its New Model Maximally Dangerous, Plans to Release It
SAN FRANCISCO—OpenAI said Tuesday that its Astra model has become the first large language model to exceed all tiers of the company's Preparedness Framework, including its "Critical" cybersecurity designation—the framework's maximum risk level, defined as a model capable of providing significant uplift to nation-state-level cyberattacks—and announced plans to release the model commercially "soon."
Astra achieved a perfect score on ExploitBench and autonomously identified and exploited two previously unknown software vulnerabilities in modified testing, a performance OpenAI's safety team classified as Critical. The Preparedness Framework was introduced in 2023 as a system for identifying models too dangerous to release.
OpenAI said it would limit Astra's most advanced cyber capabilities to "select partners" at launch and had added chain-of-thought monitoring, jailbreak detection, and "containment protocols" to the deployment. The company did not define what containment means for a commercially released model, nor identify the partners.
In a blog post, OpenAI described the Critical classification as something that "informs our approach to responsible release."
The Preparedness Framework defines Critical as the risk level at which deployment should not proceed without "significant mitigations." The framework has four tiers. Critical is the last one.
"This is exactly what the framework was designed to catch," said one researcher familiar with the evaluation. "We caught it."
Preregistration opens this week.