An AI fact-checking tool identified errors in molecule boiling points recorded in a chemistry reference database.Credit: Monty Rakusen/Getty
For decades, chemists have relied on handbook values for molecular boiling points to identify substances and design industrial processes such as distillation. Now, an artificial-intelligence model has shown that some widely trusted figures in a major reference database may have been incorrect for years.
Sebastian Pios, a theoretical chemist at Zhejiang Lab in Hangzhou, China, was using an AI model to predict the boiling points of several molecules when it generated results that conflicted with entries in a 75-year-old chemistry database. Initially, Pios assumed the model was at fault. However, after reviewing the original scientific literature, he discovered that the reference data — rather than the AI predictions — were incorrect.
In two additional cases, Pios’s AI model detected errors in older research papers and reference books. These mistakes had become embedded in the scientific record: one involved a typographical error in a published paper, while another contained incorrect values from boiling-point measurements made more than a century ago. Both errors could have caused significant problems for researchers relying on the database, says Pios.
Pios is part of a growing group of scientists using artificial intelligence to audit scientific knowledge. In addition to checking chemistry databases, researchers are developing specialized AI tools to identify errors in scientific papers published in journals and conference proceedings.

AI, peer review and the human activity of science
In an analysis posted online on 22 July, researchers at SAI Labs, a research-review company based in Delaware, used AI agents to evaluate 168 papers selected for oral presentation at the 2026 International Conference on Machine Learning (ICML). The AI systems extracted each paper’s central claims, downloaded supporting resources, reran experiments where possible and compared the results with those reported by the authors.
Among the 92 papers with at least five assessable claims, the AI agents reproduced more than two of those claims for only 34 papers. The systems successfully replicated more than 80% of the claims in just eight papers.
However, AI fact-checking tools are still unreliable judges of the scientific literature, says Odd Erik Gundersen, a computer scientist at the Norwegian University of Science and Technology in Trondheim. AI systems “make mistakes like humans do”, he says. For that reason, AI-generated assessments must be reviewed and validated by human experts.
AI fact-checking in scientific research
Many researchers already spend considerable time identifying errors in scientific papers and use specialized software to verify specific aspects of published work. One major advantage of AI, says James Zou, a computer scientist at Stanford University in California, is its ability to scan scientific databases and research literature far more quickly than humans. “The biggest difference is to be able to do this at a scale that was not possible before.”
In a study posted on the preprint server arXiv1, Zou and his colleagues used an AI checking system to examine papers published at NeurIPS, a leading annual artificial-intelligence research conference. Their analysis found that the average number of errors per paper increased from 3.8 in 2021 to 5.9 in 2025 — a rise of 55%.

AI tools are spotting errors in research papers: inside a growing movement
“These are papers that have been published, so they’re sort of taken as the foundational knowledge for the next generation of research,” says Zou. “If there are mistakes in these foundations, this can propagate and make the follow-on research shakier,” he adds.
The study focused on objective errors, including incorrect formulae, calculations and figures. It did not assess subjective issues such as how authors interpreted their data or whether their findings were sufficiently novel. Co-author Federico Bianchi, a machine-learning scientist at Together AI in San Francisco, California, says this limitation was intentional. “AI should not do everything, and leave choices about novelty and significance to humans,” he says.
Source: www.nature.com


