PNAS: 57% of 2025 Academic Papers Show LLM Fingerprints
TL;DR
- By 2025 roughly 57% of articles across 7.3 million sampled papers showed LLM-associated language, up from about 12% in 2023.
- The study, by Kyle Siler of the University of Toronto in PNAS, spans full texts from Elsevier, Frontiers, MDPI, and PLoS from 2020 to 2025.
- Lower-ranked institutions, young for-profit publishers, and regions further from English as a primary language showed the highest LLM-associated language rates.
Half of new scientific papers now carry the linguistic fingerprint of a language model, and the interesting part is where those fingerprints cluster.
In a new PNAS study, Kyle Siler of the University of Toronto ran 7.3 million journal articles published between 2020 and 2025 through a detector he built from 228 words whose frequency spiked sharply after 2022, consistent with LLM output. By 2025 roughly 57% of the articles in his corpus showed evidence of LLM influence, up from about 12% in 2023. The sample draws full texts from four large publishers, Elsevier, Frontiers, MDPI and PLoS, so it is a substantial slice of what actually got published rather than a survey of what authors admit to.
What makes the paper more than another 'AI is everywhere' data point is the map of where adoption concentrates. Siler reports that lower-ranked institutions use LLM-associated language at higher rates than elite universities, that young for-profit publishers show elevated rates versus competitors, and that economic development and proximity to English as a primary language are key predictors of regional variation. He also flags that the influence ranges from subtle linguistic polish to articles that are mostly or entirely LLM-generated.
The honest caveat is that a 228-word fingerprint is a correlation, not a confession. It cannot separate a non-native English speaker running a draft through a model for polish from a paper written wholesale by one, and the paper emphasises that heterogeneity itself. The reporting also does not tell us how peer reviewers, editors, or citation counts have responded to any of it, only how the language has shifted.
The forward-looking read: publishers whose current policies only ask authors to disclose 'AI use' now have to decide what disclosure means when more than half of new papers, by this measure, are touched by a model. If a journal's competitive advantage is trust in the manuscript, the coming rounds of publisher policy are where that gets tested.
Shared on Bluesky by 3 AI experts
-
1. How common is LLM use in scientific publishing, and how does it vary across field, publisher, journal prestige, author demographics etc.? @kylesiler.bsky.social has new paper in PNAS that addresses this question on a …
View on Bluesky →
Originally reported by scoop-hunter
Read the original article →Original headline: PNAS: Over Half of All Academic Articles Now Show LLM Influence—7.3M-Paper Study