scientificamerican.com via Reddit

OpenAI's Astra math proofs draw research misconduct claims

4 sources tracking this story

TL;DR

  • OpenAI initially claimed no progress existed for a decade on problems where named papers from 2016 and 2019 already provided key steps, and quietly updated the language after criticism.
  • Astra's sphere-packing proof recycled Steven Miller's 2016 argument without attribution; Fournier-Facio (Cambridge) identified the same unattributed prior-work pattern in the non-sofic groups result.
  • OpenAI cited the Leiden Declaration to justify responsible disclosure, then released results via a company blog post rather than peer-reviewed journals, which is the exact practice the Declaration warns against.

OpenAI's announcement of 10 mathematical advances from its Astra system, produced for a reported $2,000 in compute, has run into a specific and uncomfortable pushback from working mathematicians. According to reporting in Scientific American, two of the flagship results appear to lean on prior published work that OpenAI did not properly cite.

Steven Miller, a mathematician at Yeshiva University, says the sphere-packing proof reuses an argument from his own 2016 paper without credit, and he does not describe it as an oversight. He told the magazine the team is 'running roughshod over the work of others who came before them in a deliberate way' and that the pattern 'points to research misconduct.' Francesco Fournier-Facio, a group theorist at the University of Cambridge, says the soficity 'breakthrough' actually combined ideas from existing 2016 and 2019 papers, and he frames the presentation as inflated rather than fraudulent, calling out 'the big PR machine that wants to sound as impressive as possible.'

Why this matters beyond one paper: OpenAI has been leaning on math results as evidence that its models are becoming genuine research collaborators, and the $2,000 compute figure is what makes the story go viral. If the strongest examples turn out to be recompositions of known work without attribution, the case for 'AI is doing new mathematics' narrows to the one soficity result that Fournier-Facio was able to reconstruct independently, while several other results in the batch have no comparable public expert validation as of the reporting.

An OpenAI spokesperson told Scientific American the company will 'take responsibility for the correctness of these results' and plans minor updates this week. Scientific American's account stops short of deciding whether the missing citations are misconduct or a rushed write-up, and it does not spell out what Astra actually is beyond an LLM. The article also lacks a systematic audit of the remaining results in the batch, or any specifics on what the promised updates will actually contain.

The forward-looking piece is less dramatic than the headline: this is what serious external review of AI-generated science looks like when it happens in days rather than months, and it sets an early template for how the next round of 'AI solved X' announcements from any lab will be received.

What others are reporting

Coverage cluster as of 8h after publish

  1. SiliconAngle Read →

    Establishes the Erdős overclaim precedent and Leiden Declaration backdrop; documents Bloom's 'dramatic misrepresentation' verdict on the prior OpenAI math claim that grounds SA's recurring-pattern finding.

    a dramatic misrepresentation
  2. The Decoder Read →

    Documents OpenAI's own defense: the company used the Leiden Declaration to argue AI should receive attribution rather than humans, reframing the question as AI credit rather than prior-work credit.

    The mathematical arguments themselves, however, came from Astra.
  3. The Next Web Read →

    Frames Lean certificates as the mathematical community's verification standard while clarifying that machine-checkability addresses correctness, leaving the attribution and originality dispute entirely unresolved.

    Machine-checkable proofs can be validated by anyone with the Lean compiler, without trusting the model

Shared on Bluesky by 1 AI expert