Berkeley/DeepMind paper: LLMs shift essay meaning, even on grammar edits
TL;DR
- Heavy ChatGPT use led to nearly 70% more essays staying neutral on 'Does money lead to happiness?', in a Berkeley/DeepMind user study.
- Even when told to only fix grammar, gpt-5-mini, gemini-2.5-flash and claude-haiku changed the semantic meaning of argumentative essays.
- The team estimates 21% of ICLR 2026 peer reviews were AI-generated, scoring papers a full point higher on average.
The paper's headline finding: writers who leaned heavily on ChatGPT produced nearly 70% more essays that stayed neutral on the question they were asked to argue.
That comes from "How LLMs Distort Our Written Language", a March 2026 arXiv preprint by researchers at UC Berkeley, Google DeepMind, the University of Washington, UC San Diego, and Zaytuna College. In their user study, 100 native English speakers wrote argumentative essays on "Does money lead to happiness?"; the group using gpt-4o-mini came out systematically more hedged than the pen-and-keyboard controls, and reported the finished essays felt "less creative and less in their own voice", per a writeup by Originality.AI, even as their reported satisfaction stayed as high or higher.
The paper's more uncomfortable result concerns edits. In a second study, the authors handed gpt-5-mini, gemini-2.5-flash and claude-haiku expert human feedback and asked them to revise 86 argumentative essays. "Even when LLMs are prompted with expert feedback and asked to only make grammar edits, they still change the text in a way that significantly alters its semantic meaning," the paper reports. The uniformity of the shifts across models, the authors write, "suggests a convergence toward LLM-preferred linguistic patterns."
The third study looked at ICLR 2026. Of roughly 18,000 peer reviews across 9,000 papers, the team estimates 21% were probably AI-generated, with those reviews assigning scores a full point higher on average and, in the paper's words, placing "significantly less weight on clarity and significance of the research."
Psychology Today's writeup quotes lead author Natasha Jaques describing the pattern as the "blandification" of writing, warning that LLMs "could be making writing blander, less personal, and more neutral, at scale." Two of the AI researchers we track shared the YouTube segment covering the paper.
The 21% figure is a detector estimate, not a reviewer admission, and the study measured English-language writers only.
Shared on Bluesky by 2 AI experts
-
The video of Marwa Abdulhai and Isadora White's talk about this paper is now online here - they present a pretty complex computer science paper so clearly, I highly recommend watching it if you're interested in using LLM…
View on Bluesky →
Originally reported by youtube.com
Read the original article →Original headline: How LLMs Distort Our Written Language #AIStories