Expert audit finds physics benchmarks near-saturated by AI models
TL;DR
- Of 250 model answers marked wrong across six physics benchmarks, physicists found only 12 were genuine physics errors; the rest were grader or reference-solution defects.
- On HLE-Physics, GPT-5.6-Sol's mean@4 rose from 47.3% to 78.7% after expert re-grading; on CMT-Benchmark it went from 61.0% to 87.2%.
- Corrected pass@4 hit 94.4% on the 54 retained CritPt challenges; the audit was limited to text-only problems with verifiable final answers.
Of 250 model answers marked wrong on six leading physics benchmarks, only 12 were cases where the model actually got the physics wrong. The rest were defective items, bad reference solutions, or graders that missed a correct answer.
That is the top-line finding of an arxiv paper led by Ali Ansari, in which faculty and graduate researchers hand-audited frontier-model responses on six widely used physics benchmarks, among them HLE-Physics, CMT-Benchmark and CritPt.
The numbers move a lot after re-grading. On HLE-Physics, GPT-5.6-Sol's measured mean@4 rose from 47.3% to 78.7%. On CMT-Benchmark it went from 61.0% to 87.2%. On the 54 CritPt challenges retained after the audit, corrected pass@4 hit 94.4%.
The paper's conclusion is blunt: "Frontier models can now solve a broad range of well-defined, problem-set-style physics questions with near-perfect accuracy, contrary to the low scores reported by Artificial Analysis." The scope is narrower than that sentence sounds. The audit was limited to text-only problems with verifiable final answers, and reviewers with faculty or graduate-level expertise went through problem statements, reference solutions and model responses one by one to separate real errors from grader mistakes.
Two of the researchers we track posted the paper the day it landed, which fits a preprint that reads less as a capability release than as a critique of how leaderboards get built.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks