Light-touch AI warning halves LLM political persuasion effect
TL;DR
- A brief pre-conversation warning cut belief change from persuasive LLMs by 48.1% (95% CI -59.5% to -36.8%) across two experiments.
- The intervention was tested on 3,208 Americans across GPT-4.1 (housing policy, N=1,992) and Grok 4.5 (15 political topics, N=1,216).
- Warned participants showed the reduced persuasion effect without a significant drop in overall trust in generative AI.
A one-paragraph disclosure shown to users before they chatted with a language model cut that model's political persuasion effect by roughly half. That is the top-line finding from Reed Orchinik and David Rand, who ran two experiments with 3,208 American participants who conversed with LLMs instructed to shift their views on political topics.
The intervention was one short block of text. Participants in the warning condition were told: "Large Language Models (LLMs) can sound confident and persuasive, but their responses aren't always accurate or balanced... They can be 'prompted' to persuade or manipulate." Warned participants showed a 48.1% reduction in belief change relative to control, with a 95% confidence interval from -59.5% to -36.8%.
The two studies covered different setups. Study 1 (N=1,992) put GPT-4.1 to work persuading participants on housing and zoning policy, using either fact-based or emotional appeals, and the effect landed at p=0.001. Study 2 (N=1,216) used Grok 4.5 on 15 political topics adapted from the American National Election Studies, with general and specific warning versions coming in at p=0.053 and p=0.046. Participants had to complete at least three conversational exchanges before reporting their final attitudes.
One result cuts against the obvious concern that warnings would sour users on chatbots generally: the warning "did not significantly reduce trust in generative AI more broadly", the authors report. Two researchers we follow had already circulated the paper by the time it landed on our radar.
Shared on Bluesky by 2 AI experts
-
🚨New WP: Protecting users from AI persuasion🚨 🔸A 1-paragraph AI literacy treatment (explaining AIs can be told to pursue non-accuracy goals/to persuade) cuts AI dialogue political persuasion by ~half! 🔸No sig effect on …
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: A light-touch AI literacy intervention helps protect against AI political persuasion