arxiv.org web signal

Better LLMs act more alike in market sims, paper warns

TL;DR

  • A new arxiv paper by Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak and Andrew W. Lo finds frontier LLMs behave more similarly as capability rises.
  • In an agent-based market simulation, correlated LLM traders reduce risk when their shared reasoning is accurate but amplify it under common misinformation.
  • The authors call this a 'capability paradox' and describe a 'non-diversifiable risk floor' that improving individual models cannot lower.

"Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring." That is how a new arxiv preprint by Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak and Andrew W. Lo opens. The claim: pushing individual model capability up can push system-level performance down.

The mechanism they propose is that shared training data and architectures make more advanced LLMs behave more like each other. In an agent-based financial simulation with LLM traders of varying general-purpose capability, the authors report that "frontier LLMs exhibit significantly correlated behavior that increases with capability."

When those aligned agents reason on accurate information, more participation actually lowers market-level risk. When they share "a common misinformation environment, the same correlated behavior becomes a liability."

The paper calls this a "capability paradox" and describes a "non-diversifiable risk floor" that improving any single model cannot lower. The 14-page preprint, posted September 3, 2026, does not extend the test beyond markets: "Whether the same dynamics arise in other domains is an open empirical question."

Shared on Bluesky by 2 AI experts