arxiv.org web signal

Paper flags gap in OpenAI's Lean-verified Navier-Stokes proof

TL;DR

  • A new preprint argues OpenAI's announced Lean proof of Navier-Stokes blow-up does not correspond to the natural-language proof it was meant to verify.
  • The authors place semantically faithful autoformalisation at SCI = infinity in the Solvability Complexity Index hierarchy, above the Halting problem's SCI = 1.
  • They present several practical examples of AI mistranslations between natural language and Lean producing mismatches between the proofs and their Lean verifications.

OpenAI's announced Lean proof of blow-up of solutions to the Navier-Stokes equations does not correspond to the natural-language argument it was meant to formalise, according to an arxiv preprint posted October 6 by Alexander Bastounis, Fabian Circelli and Anders C. Hansen.

The abstract states it plainly: "the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations." Two of the mathematicians we track have already passed the link around.

The authors' broader argument is theoretical. They place the problem of resolving ambiguities needed for a semantically faithful translation from mathematical English into Lean "arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy," and conclude that, informally, "providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI = 1)."

To ground the result in current practice, they add "several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean 'verifications'." The Navier-Stokes case headlines the set. The abstract names no specific Lean file, no commit hash, and no accuracy rate across the examples.

Shared on Bluesky by 2 AI experts