Pew: AI 'digital twins' miss human survey answers by 12 points
TL;DR
- Pew tested AI 'digital twins' on nearly 300 real survey questions; synthetic answers missed human responses by an average of 12 percentage points.
- The model stereotyped badly, predicting 97% of Hispanic adults follow the World Cup when the actual share is 43%.
- Human respondents chose 'not sure' about four times more often than the AI; GPT-5.1 and Claude Opus 4.6 also produced substantially different answers.
Pew Research Center ran an AI model as a survey respondent across nearly 300 questions, feeding it detailed profiles of real humans in its American Trends Panel, and compared what the model said to what those same humans actually told pollsters. The synthetic answers were off by an average of 12 percentage points, and 28% of questions missed by more than 15.
"AI models are not an adequate replacement for traditional polling on topics of broad public importance," the Center's researchers write in its new methodology post. The team of Athena Chapekis, Arnold Lau, Samuel Bestvater, Sono Shah, Andrew Mercer and Aaron Smith tested what they call a "digital twins" approach, "asking an AI model to adopt the personas of real humans who are members of the ATP."
The stereotyping showed up fast. The model estimated 97% of Hispanic adults were at least somewhat likely to follow the World Cup. The actual share was 43%. On a First Amendment knowledge question, 98% of synthetic respondents got it right versus 52% of real humans, a 46-point overshoot. The AI pegged Trump job approval at 46% against the real 34%, and awareness of data centers at just 3% against 25%.
And the model almost never hedged. "Across all the questions on our three surveys where a 'not sure' option was offered, human panelists were around four times as likely as the AI model to choose it," the researchers report. Results shifted with the model behind the persona too: identical prompts to OpenAI's GPT-5.1 and Anthropic's Claude Opus 4.6 produced substantially different answers.
Pew's conclusion is flat. "We see no substitute for rigorous, multimode, probability-based surveys that allow real members of the public to speak their minds on issues of importance."
Shared on Bluesky by 4 AI experts
-
Pew tested AI synthetic surveys against actual survey data. Conclusion: “This experiment taught us that, at this time, AI models are not an adequate replacement for traditional polling on topics of broad public import…
View on Bluesky → -
We just had a discussion on this in our editorial meeting of Internet Policy Review. It would be great to hear from other academic journal editors what are their policies around researchers using synthetic data. -> w…
View on Bluesky → -
There is a lot of "oh dear" to go around in the results of this very thorough Pew study on the issues with synthetic samples, but this strikes me as particularly "oh dear" www.pewresearch.org/data-labs/20...
View on Bluesky →
Originally reported by pewresearch.org
Read the original article →Original headline: Can AI Stand In for Human Survey-Takers? Not Really