scmp.com via Reddit

Beijing doctor uses GPT-5.6-Sol to crack 22-year math conjecture

TL;DR

  • Jin Shanmu, a neurosurgery resident at Peking Union Medical College Hospital, ran GPT-5.6-Sol on ChatGPT Work for about 16 hours to prove Crouzeix's conjecture.
  • Michel Crouzeix, who proposed the matrix-analysis problem in 2004, checked the proof thoroughly and believes Jin's manuscript is correct.
  • Eight days after Jin's preprint, Emiel Lorist and Felix Schwenninger posted an independent five-page proof, disclosing they also used ChatGPT 5.6.

A neurosurgery resident with no formal training in advanced mathematics has produced a proof of Crouzeix's conjecture that the problem's original proposer says is correct, closing a matrix-analysis question that had gone unresolved since 2004. According to the South China Morning Post, Jin Shanmu, a postdoctoral researcher at Peking Union Medical College Hospital in Beijing whose undergraduate degree is in geology, set GPT-5.6-Sol running on ChatGPT Work for roughly 16 hours, and came back to a full manuscript.

The conjecture, as SCMP phrases it, is that 'the norm of applying any function to a matrix is no larger than twice the function's maximum value on that matrix's numerical range.' French mathematician Michel Crouzeix posed it in 2004; Jin ran into it sideways while working on transcranial ultrasound. Crouzeix himself has since checked the proof thoroughly and believes Jin's manuscript is correct.

The part that will get OpenAI's marketing team excited is not that a model helped a mathematician (already commonplace) but that the model was left to work unsupervised on a real open problem and produced something a specialist accepted. Eight days after Jin's preprint, two researchers, Emiel Lorist and Felix Schwenninger, posted an independent five-page proof and disclosed that they, too, had used ChatGPT 5.6 to explore strategies. That corroboration matters more than a single dramatic run. OpenAI previewed a faster variant of the same model at 14x speed on Cerebras earlier this same day.

The caveats are the ones any working mathematician would raise. One review by the problem's original author is not the same as broad formal peer review, the SCMP write-up does not say how many failed runs preceded the successful one, and Jin's open-sourced prompt and iterations are so far the only public record of how the model got to the answer. Whether GPT-5.6-Sol can do this on a second, unaided conjecture is the question every applied-math department will now be quietly testing.

For domain experts who happen to have an open problem lying next to their day job (the neurosurgeon studying ultrasound, the applied physicist stuck on a bound), the practical read is that the tool now clears a bar that used to be reserved for specialists. That, more than any single benchmark, is what a 16-hour unattended run to a decades-old proof is really evidence of.