New Paper: LLMs Alter Meaning of Writing, Not Just Style
TL;DR
- Heavy LLM users produced nearly 70% more essays that stayed neutral on the topic question, and reported the writing felt less creative and not their own.
- Even when asked only to fix grammar in 2021 human-written essays, LLMs significantly altered the semantic meaning of the text.
- About 21% of peer reviews at a recent top AI conference were AI-generated, and those reviews scored papers a full point higher on average.
There is a specific finding in this new arXiv paper that is worth pausing on. Heavy LLM users in a controlled study produced nearly 70% more essays that stayed neutral on the topic question they had been asked to argue. They also told the researchers the writing felt less creative and not in their own voice.
That is a stronger claim than the usual line about AI writing sounding a bit bland. The authors, Marwa Abdulhai and colleagues, argue LLMs do not just alter voice and tone, they consistently alter the intended meaning. To test that, they took a set of human-written argumentative essays collected in 2021, before LLMs were in everyday use, and asked models to revise them using the original expert feedback. Even when the model was told to make only grammar edits, the paper reports it still changed the text in a way that significantly alters its semantic meaning.
The most consequential piece of the study is the peer review analysis. At a recent top AI conference, the researchers focus on 21% of scientific peer reviews they identify as AI-generated, and report those reviews place significantly less weight on clarity and significance of the research, and assign scores that on average are a full point higher than human reviews. For anyone whose paper acceptance depends on those scores, that is not a stylistic quibble, it is a real distortion of the signal reviewers are supposed to send.
The honest caveat is that this is a preprint and the neutrality figure comes from a single user study, so take the exact percentages as reported, not settled. What the paper does not give you is a model-by-model breakdown or a clear read on whether prompt design can dial the effect back. But the direction is what matters, and the peer review number in particular is the sort of thing conference program chairs should probably be talking about this week.
Shared on Bluesky by 2 AI experts
-
Fantastic paper demonstrating how LLM editing is gradually distorting our writing and the way we think to write: arxiv.org/abs/2603.18161
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: How LLMs Distort Our Written Language