Nature study finds LLM rewrites flatten linguistic diversity
TL;DR
- A four-study paper in Nature Human Behaviour finds LLM writing assistants preserve content but homogenize style, muting individual voice.
- Classifiers that infer age and gender from writing lose accuracy once text has been polished by an LLM, the paper reports.
- The USC team ran the same tests across GPT-3.5, LLaMA 3 and Gemini and across Reddit, Patch News and arXiv datasets.
Feed a paragraph through ChatGPT or Gemini for a polish and the meaning survives; the stylistic fingerprints do not. That is the finding of a four-study paper in Nature Human Behaviour by a University of Southern California team led by Zhivar Sourati and Morteza Dehghani, which tested writing assistants across Reddit posts, Patch News stories and arXiv abstracts.
The authors write that widespread LLM writing assistance is 'linked to notable declines in linguistic diversity and may interfere with the societal and psychological insights language provides.' The sharper claim is about the direction of the flattening: LLMs 'homogenize writing styles' and 'alter stylistic elements in a way that selectively amplifies certain dominant characteristics or biases while suppressing others,' the paper says, 'emphasizing conformity over individuality.'
The consequence they document is measurable. Classifiers trained to predict demographic and psychological traits from writing lose accuracy, particularly for age and gender, once the text has been through an LLM rewrite. The team reports the same pattern across GPT-3.5, LLaMA 3 and Gemini, and across varied prompts and contexts. No per-model effect sizes appear in the abstract.
The paper's own list of risks is unusually specific: compromised diagnostic and personalization workflows, harm in 'personnel selection where language plays a critical role in assessing candidates' qualifications, communication skills, and cultural fit,' and the undermining of efforts for cultural preservation. Two of the researchers on our tracker circulated the link.
Shared on Bluesky by 2 AI experts
-
Eric Topol @erictopol.bsky.social: LLMs--> LLD (less language diversity) https://t.co/pcWXXTQ0MC @NatureHumBehav →
Originally reported by nature.com
Read the original article →Original headline: The shrinking landscape of linguistic diversity in the age of large language models - Nature Human Behaviour