AI Lab's Annual Safety Report Finds Newest Model Safest Model Ever; Does Not Define Ever
SAN FRANCISCO—A frontier artificial-intelligence laboratory published its annual safety assessment Monday, finding that its newest model is the safest model the company has ever released—a conclusion the lab's head of safety called "the clearest evidence yet that our safety work is compounding," and that an independent researcher, reviewing the methodology, described as "a number divided by itself."
The 94-page report documents that Model 6.1 produces dangerous content 43% less often than Model 5.9, resists jailbreak attempts 61% more reliably than Model 4.8, and scores 89% better on catastrophic-misuse potential than the company's 2023 release. All 14 charts in the report trend downward.
The report does not contain a figure representing how dangerous the models are in absolute terms, a comparison to any standard not created by the lab, or a value at which the lab would consider the models safe without further qualification. Page 64 notes that Model 6.1 ranked first industry-wide on harmful-output avoidance—an evaluation the lab designed, conducted, and published. Asked how the lab defines an acceptable level of catastrophic risk, the spokesperson cited the previous annual safety report.
The report also introduces the Relative Safety Improvement Index, a proprietary metric tracking safety gains versus the prior model version, which has increased 28% annually since its 2024 introduction. RSII has no ceiling, no external reference point, and no value at which progress would be considered complete.
Over the four years covered by the report, the lab's safety team grew from 11 researchers to 94. Its capability team grew from 47 to 1,200.
"Every year the models get safer," said the head of safety, in remarks prepared before the announcement of Model 6.2.