AI agents flag decades-old errors in the scientific canon
TL;DR
- A Zhejiang Lab chemist's AI predicting boiling points clashed with a 75-year-old reference database; manual checks showed the database, not the model, was wrong.
- The same AI spotted further mistakes in older papers and reference books, including a typo and incorrect values of century-old boiling-point measurements.
- Researchers caution AI fact-checkers are not reliable on their own because the models make mistakes like humans do and still need manual oversight.
A theoretical chemist ran an AI over some routine boiling-point predictions, saw it disagree with a reference book, assumed the AI was wrong, and then discovered that the reference book had been wrong for 75 years. That is the small, specific story Nature reports at the top of a broader piece on AI agents starting to comb the scientific literature for errors.
The chemist, Sebastian Pios at Zhejiang Lab in Hangzhou, was predicting molecular boiling points when the model kept clashing with entries in a long-standing reference database. Manual checks of the original literature showed the database was at fault, not the model. In two further cases, his AI surfaced errors that had made their way into the scientific canon, including a typo in an older paper and incorrect values of century-old boiling-point measurements.
The reason this is worth paying attention to, beyond the anecdote, is what reference databases and textbook constants feed. They sit under student problem sets, industrial process calculations, and increasingly the training data of the machine learning models that chemistry groups now lean on. If even a small share of long-accepted numbers are subtly wrong, an AI that consistently flags disagreements becomes a genuinely useful audit layer rather than a science-news curiosity. It also flips a common reflex. When a model contradicts an established source, the instinct is to trust the source. Sometimes that instinct is wrong.
The honest caveat is the one the reporting itself keeps returning to. These fact-checking systems make mistakes like humans do, which is why researchers stress the outputs still need manual verification before anything gets 'corrected' in the record. What the piece doesn't give you is the base rate, how often the model flags something that turns out to be a real error versus a false alarm, or which specific reference database and molecules were caught out. Take the wins as reported, not as a settled batting average.
The forward-looking thread is straightforward. If chemistry ML groups, database curators, and journal integrity teams can put a reliable LLM audit layer between the corpus and the next generation of models, the compounding-error problem in scientific ML gets smaller. That is a modest but useful shift, and it will benefit whoever staffs the human-in-the-loop step first.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: AI agents are checking the scientific literature — and spotting decades-old errors