arxiv.org web signal

LLMs write more similar stories than humans, paper finds

TL;DR

  • Across 10 representative LLMs, model-generated narratives were consistently more similar to each other than human-written stories were.
  • Frontier models converged on a 'mean' generic narrative that resembled individual human stories but lacked the collective diversity of human authors.
  • Negative prompting and temperature scaling failed to meaningfully reduce the observed output homogeneity.

LLM-written narratives are more similar to each other than human-written ones are. That is the flat finding of a new arxiv preprint by Thennal DK and Hans Ole Hatzel, submitted June 15, 2026.

The authors built a contrastive framework, drew stories and prompts from r/WritingPrompts, then ran 10 representative LLMs and collected narrative-similarity judgments — using human evaluators and three separate automatic annotation methods.

Their framing of the result is careful. "Frontier models in particular converge on a ``mean'' generic narrative that approximates individual human stories but lacks the collective diversity of human authors," the paper says. Read individually, an LLM story can pass as something a person might have written. Read as a set, the outputs collapse toward one another.

The paper also tries the standard fixes and finds them wanting. "Common mitigation strategies, including negative prompting and temperature scaling, fail to meaningfully address this homogeneity," the authors write.

The abstract does not name which 10 models were tested, nor break out how the human raters and the three automatic methods compared. Two researchers we track posted the arxiv link into their feeds.

Shared on Bluesky by 2 AI experts