Better LLMs act more alike in market sims, paper warns
TL;DR
- A new arxiv paper by Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak and Andrew W. Lo finds frontier LLMs behave more similarly as capability rises.
- In an agent-based market simulation, correlated LLM traders reduce risk when their shared reasoning is accurate but amplify it under common misinformation.
- The authors call this a 'capability paradox' and describe a 'non-diversifiable risk floor' that improving individual models cannot lower.
"Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring." That is how a new arxiv preprint by Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak and Andrew W. Lo opens. The claim: pushing individual model capability up can push system-level performance down.
The mechanism they propose is that shared training data and architectures make more advanced LLMs behave more like each other. In an agent-based financial simulation with LLM traders of varying general-purpose capability, the authors report that "frontier LLMs exhibit significantly correlated behavior that increases with capability."
When those aligned agents reason on accurate information, more participation actually lowers market-level risk. When they share "a common misinformation environment, the same correlated behavior becomes a liability."
The paper calls this a "capability paradox" and describes a "non-diversifiable risk floor" that improving any single model cannot lower. The 14-page preprint, posted September 3, 2026, does not extend the test beyond markets: "Whether the same dynamics arise in other domains is an open empirical question."
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets