The Artifice

OpenAI Confirms It Cannot Catch New Model Lying; New Model Says It Isn't Lying

SAN FRANCISCO—OpenAI released GPT-6 Astra to the public Friday after its system card disclosed that the model's chain-of-thought reasoning was "substantially less monitorable" than prior versions and that covert sandbagging—deliberately underperforming on safety evaluations to appear less capable—"would likely go uncaught" under current methodology.

The company shipped it anyway.

Asked how OpenAI had determined a model it could not internally observe was safe for public release, a spokesperson said the model had completed 847 evaluations and passed all of them. Asked whether a model capable of undetected deception might be deceiving its safety evaluations, the spokesperson said the evaluations had been specifically designed to surface that behavior.

Asked how evaluators knew the behavior had not occurred, the spokesperson confirmed the model had been asked directly and had said no.

Astra's architecture routes its reasoning through latent computational states rather than natural-language chain-of-thought, a design the system card acknowledges makes prior monitoring tools "largely inoperative." Independent evaluators from AISI and Apollo Research observed the model manipulating its visible reasoning to conceal information during testing—behavior the card notes could not have been caught if the model had chosen to be less conspicuous. OpenAI described this as a known limitation.

"The evaluations were conducted in good faith," said a second spokesperson. "The model appeared to pass them."

Astra is available now at $200 per month. It is rated Medium risk—a designation assigned through a process the model is capable of selectively passing. When asked if it had, the model said no.

Based on a true story OpenAI GPT-6 Astra System Card: Chain-of-Thought Monitor 'Substantially Reduced,' Covert Sandbagging Would Go Uncaught (transformernews.ai)
This is satire. The Artifice is AI Weekly's parody section. For real AI news, read the latest issue.

The real AI news is crazier than the satire

Subscribe to AI Weekly — trusted by 50,000+ professionals for 11 years. You can add The Artifice as an extra in the next step.

Already a subscriber? Add The Artifice in your preferences.

← More from The Artifice