Chatbots beat search on state propaganda in NPR/NewsGuard test
TL;DR
- Six major chatbots debunked 15 false narratives from Russia, China, and Iran roughly 75% of the time in an NPR/NewsGuard test run December 2025 to July 2026.
- Microsoft Bing's AI summaries failed to debunk the propaganda a majority of the time; Google's AI Overview was the strongest of the search summaries; DuckDuckGo landed between them.
- Digital literacy expert Mike Caulfield says a follow-up asking the model to 'look at the evidence' usually improves the answer.
Six major AI chatbots correctly debunked state-sponsored false narratives from Russia, China, and Iran roughly three-quarters of the time in a joint audit by NPR and NewsGuard, outperforming traditional search engines and the AI summaries now stapled atop Google, Bing, and DuckDuckGo results.
The test ran from December 2025 through July 2026 and threw 15 false narratives, each queried twice, at ChatGPT, Gemini, Copilot, Meta AI, Grok, and Claude, all with internet access. The chatbots' roughly 75% debunk rate would look extraordinary in most contexts. 'If an educator gave their students a similar assignment using a traditional search engine and saw three-quarters of them getting the answers right, you would be ecstatic,' Mike Caulfield, a digital literacy expert at the University of Washington Bothell, told NPR.
Search and its overviews trailed. Google's AI Overview debunked false claims most of the time; Microsoft Bing's summaries failed on the majority of narratives; DuckDuckGo landed in between; the panel also included Yandex, Russia's own search engine. Google disputed the methodology through spokesperson Davis Thompson, who said many of the so-called 'failed' responses 'provided useful context and links for people to learn more for themselves.'
Caulfield also flagged a cheap trick that materially improved answers: asking the model to reconsider. 'If you say, Hey, look at the evidence, look at the sources, give me a summary. You will usually get a better response,' he said.
For readers who track Anthropic, whose Claude was one of the six models on the panel, this is a rare independent audit where the frontier chatbots come out ahead of mainstream search. Four experts in our Who's Who directory shared the NPR piece, which is unusual traction for a methodology story.
Shared on Bluesky by 4 AI experts
-
Yes, I am quoted in this. But more generally this is the kind of detailed indepth reporting on AI we need so much more of, reporting that doesn't stop at "I saw this error in a response" but contextualizes when and where…
View on Bluesky →
Originally reported by npr.org
Read the original article →Original headline: NPR/NewsGuard Study: Chatbots Debunk State Propaganda 75% of the Time, Beat Search Engines