LLMs write more similar stories than humans, paper finds
TL;DR
- Across 10 representative LLMs, model-generated narratives were consistently more similar to each other than human-written stories were.
- Frontier models converged on a 'mean' generic narrative that resembled individual human stories but lacked the collective diversity of human authors.
- Negative prompting and temperature scaling failed to meaningfully reduce the observed output homogeneity.
LLM-written narratives are more similar to each other than human-written ones are. That is the flat finding of a new arxiv preprint by Thennal DK and Hans Ole Hatzel, submitted June 15, 2026.
The authors built a contrastive framework, drew stories and prompts from r/WritingPrompts, then ran 10 representative LLMs and collected narrative-similarity judgments — using human evaluators and three separate automatic annotation methods.
Their framing of the result is careful. "Frontier models in particular converge on a ``mean'' generic narrative that approximates individual human stories but lacks the collective diversity of human authors," the paper says. Read individually, an LLM story can pass as something a person might have written. Read as a set, the outputs collapse toward one another.
The paper also tries the standard fixes and finds them wanting. "Common mitigation strategies, including negative prompting and temperature scaling, fail to meaningfully address this homogeneity," the authors write.
The abstract does not name which 10 models were tested, nor break out how the human raters and the three automatic methods compared. Two researchers we track posted the arxiv link into their feeds.
Shared on Bluesky by 2 AI experts
-
lots of similar work as well, here's ours specifically about how LLMs always seem to write the same generic story: arxiv.org/abs/2606.17350
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: Do Large Language Models Always Tell The Same Stories?