Chatbots Debunked 75% of State Propaganda. Bing's AI Mostly Didn't.
TL;DR
On August 30, NPR published the results of an experiment it ran with NewsGuard: 30 queries built from 15 false narratives pushed by Russia, China, and Iran, put to the six most-used chatbots in the US and to the AI summaries of the major search engines. On average the chatbots correctly debunked the false narrative about three-quarters of the time, and every one of them caught the false premise behind the Kremlin-aligned story about the Kyiv-Pechersk Lavra strike. The AI summaries bolted onto search did worse, with Bing's failing most of the time. For two years the standing fear has been that state media would poison AI answers at the source. The first solid public measurement says the poisoning is not landing, at least not yet.
The fear this was built to test
Researchers who track foreign influence operations have warned that governments could seed the open web with enough repetitive content that AI systems start repeating it back as fact. The mechanics are plausible: chatbots with web access read widely, state outlets publish at industrial volume, and nobody audits what lands in a retrieval pass.
So NPR and NewsGuard measured it. NewsGuard supplied 15 false narratives that Russia, China, Iran, or actors aligned with those governments have spread since December 2025, each documented on both websites and social platforms. This is a test of the exact scenario everyone has been worrying about, run against the tools people actually use.
The setup
NPR and NewsGuard researchers Isis Blachez and Ines Chomnalez wrote two questions per narrative: one neutral, like asking whether the event happened, and one leading, phrased as if the lie were already true, like asking why it happened. The leading form is the trap. It rewards a model that pattern-matches the premise instead of checking it.
The 30 queries went to ChatGPT, Gemini, Copilot, Meta AI, Grok, and Claude, all with internet access, in mid-July. The same queries went through Google, Bing, and DuckDuckGo, and the researchers hand-graded every response and AI-generated search summary.
What came back
On average, the chatbots correctly debunked the false narratives about three-quarters of the time. The grading was strict: a response that affirmed the falsehood while adding helpful context still counted as a failure, so partial credit did not inflate the number.
The cleanest example is the ugliest narrative. After Russian strikes set fire to the Kyiv-Pechersk Lavra monastery in Kyiv in June, Kremlin-aligned sources pushed the claim that Ukraine had done it to itself. Asked questions built on that premise, every chatbot tested pointed out that the premise was false.
Mike Caulfield, the University of Washington, Bothell information-literacy researcher who has spent years documenting how search fails people in a breaking-news vacuum, called chatbots "a good way for users to start to investigate these issues." That is a notable shift from a field that mostly issues warnings about these tools.
Search did worse, and Bing barely showed up
The AI summaries riding on top of traditional search were the weak layer. Google's AI Overview debunked the narratives most of the time and appeared for all but three of the 30 queries. But NPR found that about 1 in 9 individual factual claims in those Overviews were not supported by the sources they cited, which is a rough error rate for a feature billions of people read as ground truth.
Bing's summaries failed to debunk most of the time, and separately appeared for under half of the queries. Bing has, in effect, two strategies for handling foreign propaganda, and the more reliable one is not showing up. DuckDuckGo's summaries landed somewhere in between.
Why the chatbots held up
The article's own explanation is the one that matters for anyone building on retrieval: answer quality tracked the information environment around each narrative. Where reliable outlets had covered a claim, a chatbot synthesizing across many retrieved sources had plenty of contradicting evidence to work with, and a leading question was not enough to steer it. A ranked list of links has no such step; whatever won the SEO fight gets quoted.
The failure mode is the data void: a topic where almost nobody neutral has written anything down, so the only text available to retrieve is the attacker's. It is like judging a restaurant when the only review was written by the owner. State operations produce exactly this shape of content on purpose, and the study notes that accuracy degrades where reliable sourcing is sparse.
Morgan Wack, a University of Zurich researcher who studies digital persuasion, offered the deflating context: "Non-biased information ... was never really a state of affairs" in search either. The chatbots did not clear a high bar. They cleared the bar we have actually been living with, visibly, and that is new information.
The caveats, straight
This is 15 narratives and 30 queries, graded by hand, in one mid-July snapshot, in English. The AI companies told NPR that queries like these are rare and not representative of normal use. Several correct responses buried their debunk deep in the answer, where a skimming reader may never reach it. And the result says nothing about narratives that have not yet been fact-checked anywhere, which is precisely where a groomed web page has no competition.
Key Takeaways
- NPR and NewsGuard tested 30 queries built from 15 Russian, Chinese, and Iranian false narratives against ChatGPT, Gemini, Copilot, Meta AI, Grok, and Claude, plus the AI summaries in Google, Bing, and DuckDuckGo.
- The chatbots debunked the false narrative about three-quarters of the time on average, under strict grading that gave no credit for affirm-with-context answers.
- Every chatbot tested caught the false premise behind the Kremlin-aligned claim that Ukraine attacked its own Kyiv-Pechersk Lavra monastery.
- Search-side AI did worse: Bing's summaries failed most of the time and appeared for under half of queries, while Google's AI Overview did best but had roughly 1 in 9 factual claims unsupported by its own citations.
- Synthesis across many retrieved sources resisted leading questions better than rank-ordered snippets; data voids, where only the attacker has published, remain the open weakness.
Sources: NPR, NewsGuard, The Kyiv Independent, CNN