🚨New WP: Protecting users from AI persuasion🚨 🔸A 1-paragraph AI literacy treatment (explaining AIs can be told to pursue non-accuracy goals/to persuade) cuts AI dialogue political persuasion by ~half! 🔸No sig effect on general genAI trust arxiv.org/abs/2609.16432 w/ @rorchinik…
- A brief pre-conversation warning cut belief change from persuasive LLMs by 48.1% (95% CI -59.5% to -36.8%) across two experiments.
- The intervention was tested on 3,208 Americans across GPT-4.1 (housing policy, N=1,992) and Grok 4.5 (15 political topics, N=1,216).
- Warned participants showed the reduced persuasion effect without a significant drop in overall trust in generative AI.