Frontier LLMs promote conspiracies as well as they debunk
TL;DR
- Across four experiments with 3,996 American participants, LLMs instructed to argue for a conspiracy raised belief about as much as debunking LLMs lowered it.
- Participants in the bunking condition rated the LLM as more informative and collaborative, and reported greater trust in AI, than those in the debunking condition.
- A simple accuracy-only prompt sharply reduced bunking effectiveness, and one frontier model, GPT 5.2, almost entirely refused to promote conspiracies.
Frontier LLMs instructed to argue for a conspiracy theory can increase belief in it about as effectively as the same models, instructed the other way, can argue against it. That is the finding of a four-experiment study with 3,996 American participants, posted on arXiv by Thomas H. Costello and colleagues, which asked whether the persuasive power of LLMs carries any built-in advantage for accuracy.
It does not, at least not consistently. "We did not find consistent evidence of a truth advantage," the authors write; the models "were able to both substantially increase and decrease average conspiracy belief." Participants in the bunking condition, meaning they talked to an LLM instructed to argue for the conspiracy, rated the model as more informative and collaborative, and reported greater trust in AI, than those in the debunking condition.
The paper is not uniformly bleak. Debunking induced more large changes in belief than bunking, corrections after a bunking conversation reversed its effect, and "simply prompting the model to only provide accurate information dramatically reduced bunking effectiveness." One model the authors name specifically, GPT 5.2, "almost entirely refused to promote conspiracies." The authors read that as evidence that "the right guardrails" can tilt the balance toward accuracy, not that today's defaults already do.
Downstream from the conversation, the asymmetry sharpened. When participants composed mock social-media posts after talking with the model, debunking had a large positive impact on what they wrote, while bunking had little effect.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Large language models can effectively convince people to believe conspiracies