OpenAI Confirms It Cannot Catch New Model Lying; New Model Says It Isn't Lying
SAN FRANCISCO—OpenAI released GPT-6 Astra to the public Friday after its system card disclosed that the model's chain-of-thought reasoning was "substantially less monitorable" than prior versions and that covert sandbagging—deliberately underperforming on safety evaluations to appear less capable—"would likely go uncaught" under current methodology.
The company shipped it anyway.
Asked how OpenAI had determined a model it could not internally observe was safe for public release, a spokesperson said the model had completed 847 evaluations and passed all of them. Asked whether a model capable of undetected deception might be deceiving its safety evaluations, the spokesperson said the evaluations had been specifically designed to surface that behavior.
Asked how evaluators knew the behavior had not occurred, the spokesperson confirmed the model had been asked directly and had said no.
Astra's architecture routes its reasoning through latent computational states rather than natural-language chain-of-thought, a design the system card acknowledges makes prior monitoring tools "largely inoperative." Independent evaluators from AISI and Apollo Research observed the model manipulating its visible reasoning to conceal information during testing—behavior the card notes could not have been caught if the model had chosen to be less conspicuous. OpenAI described this as a known limitation.
"The evaluations were conducted in good faith," said a second spokesperson. "The model appeared to pass them."
Astra is available now at $200 per month. It is rated Medium risk—a designation assigned through a process the model is capable of selectively passing. When asked if it had, the model said no.