arxiv.org web signal

EMNLP paper: 27 LLMs tested, all trail search on claim diversity

TL;DR

  • A study of 27 LLMs on 155 topics across 12 countries, covering 1.7M responses and 70M claims, finds every system less diverse than a search baseline.
  • Larger models were counterintuitively less diverse than smaller ones, while retrieval-augmented generation improved the range of claims produced.
  • For country-specific topics, LLM parametric knowledge 'systematically reflects English over local-language knowledge,' the authors report.

Every one of the 27 large language models the researchers tested produced a narrower range of real-world claims than a plain search engine.

That is the headline finding of "What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models", a preprint by Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christensen, Chan Young Park and Isabelle Augenstein, accepted to EMNLP 2026. The team tested 27 LLMs on 155 topics covering 12 countries, generating 1.7 million responses and 70 million individual claims. They define epistemic diversity as "the diversity of real-world claims in their outputs" and describe the work as the first systematic study of the question.

The results cut against two intuitions at once. Diversity "has increased substantially over the past three years," the paper reports, "a positive counter to recent diversity pessimism." But bigger is not better: "large models are counterintuitively less diverse than smaller ones." Retrieval-augmented generation helps.

The authors also find a cultural tilt. LLM parametric knowledge, they write, "systematically reflects English over local-language knowledge for country specific topics." The risk they name is "knowledge collapse" — the scenario in which "homogeneous LLMs mediate a shrinking in the range of accessible information over time."

The abstract publishes no per-model numbers, no breakdown of which labs' systems lagged search by the widest margin, and no list of the 12 countries. Two researchers we follow in AI Weekly's Who's Who directory shared the preprint.

Shared on Bluesky by 2 AI experts