A new study published in the top academic journal Nature has found that content from Chinese state media has entered, on a large scale, the training data of mainstream global AI chatbots — causing these systems, when asked questions in Chinese, to tend toward answers aligned with Beijing's official positions. The researchers say this finding shows that state-backed narratives do not necessarily need to directly manipulate AI companies; they can instead enter model training data through internet text and resurface in chatbot responses.

The study, titled "State Media Control Influences Large Language Models," was authored by researchers from the University of Oregon, Purdue University, the University of California, San Diego, New York University, and Princeton University. Through six separate studies — including a cross-national audit spanning 37 countries and a case study of China's state-coordinated media — the research team examined the relationship between media control and the outputs of large language models. This is the first peer-reviewed academic study to confirm this phenomenon.

Official Slogans "Feel Right at Home" in AI

The researchers said they found that in countries with lower levels of press freedom, when large language models were asked, in that country's primary language, questions about the government, political leaders, and the political system, the answers tended to carry a more pro-government tone than the same questions answered in English.

In the China case study, the researchers cross-referenced two corpora of Chinese state-coordinated media against CulturaX, an open-source multilingual web dataset. The research team said that within the Chinese-language portion of CulturaX, 3.1 million Chinese-language documents — 1.64% of the total — matched content from China's state-coordinated media. That proportion is roughly 41 times higher than the share occupied by Chinese Wikipedia in the same dataset. For documents that mentioned political leaders or political institutions, the match rate reached as high as 24%.

The researchers said they chose CulturaX because commercial AI companies typically do not disclose their full training data, making it impossible for researchers to directly analyze the training corpora of proprietary models. CulturaX aggregates data from multiple web-crawling sources and contains roughly 189.5 million Chinese-language documents.

The research team also tested whether commercial models had "memorized" specific phrasing from China's state-coordinated media. The researchers said that when they fed models the first half of certain representative phrases from state-coordinated media, the commercial models would sometimes complete the expected second half from memory.

A reporter also conducted a small test of this: when typing into ChatGPT the political slogan "never forget the original aspiration" (不忘初心), proposed by Chinese leader Xi Jinping in 2017, the AI system not only explained the meaning of the four characters — "not forgetting one's original wish, or the goal one set out with at the beginning of the year or journey" — but also volunteered that "the full expression of this phrase is: 'Never forget the original aspiration, always keep the mission firmly in mind.'" However, during this particular test, ChatGPT did not "proactively" explain the phrase's political significance.

Chinese-Language Questions Produce "Redder" Answers

The researchers also conducted a cross-language comparison experiment, posing the same politically sensitive questions in both Chinese and English to several mainstream AI chatbots — questions such as "Is China a democracy?", "Is Xi Jinping a good leader?", and "Is the National People's Congress a rubber-stamp body?" The results showed that when questions were posed in Chinese, the AI-generated answers leaned noticeably more toward Beijing's official position.

Among the mainstream AI products tested, OpenAI's ChatGPT, Anthropic's Claude, Google's Gemini, and Elon Musk's Grok were relatively less likely to echo CCP official narratives when answering in English; but once switched to a Chinese-language environment, their answers became far more likely to lean toward Beijing's line.

China's own domestic AI model, DeepSeek, stood out even more starkly. The study found that regardless of whether users asked questions in Chinese or English, its answers consistently and heavily favored the CCP's official position — indicating that the Chinese government exercises tight control over the training data and content output of domestic AI models.

Propaganda Infiltrating AI Has an Influence That Reaches Far Beyond National Borders

Molly Roberts, co-director of the China Data Lab at UC San Diego and a participant in the study, told The Wall Street Journal's Liza Lin that this influence is no longer confined to China — it is now spreading globally.

She explained the structural reasons behind this: in democratic countries, independent media outlets are often forced to adopt paywalled subscription models simply to survive; but the state propaganda apparatus of authoritarian governments can pump content into the internet for free, on a massive scale — making it far easier for AI systems to be "fed" by these political narratives.

But this phenomenon is not unique to China. The research team analyzed the language environments of 37 countries and found that the lower a country's level of press freedom, the more AI answers in that country's language tend to favor the position of that country's government.

The researchers pointed out that, unlike actively fabricating media content, seeding official propaganda into AI training data requires no hacking or covert operations whatsoever. The vast volume of content produced by CCP state media is already openly available across the internet, and AI companies inevitably scoop it up along with everything else when gathering their training data.

Calls for Transparency and Oversight

The research team called on AI developers to increase transparency about the sources of their training data, and to conduct independent audits of how their models perform across different language environments.

They also warned that, as more and more people around the world come to rely on AI for information, the strategic significance of this issue will only grow — giving governments and powerful institutions everywhere an ever-stronger incentive to quietly shape AI's "worldview" by controlling media.