How does Chinese state media control influence large language models? An interview with Dr. Hannah Waight and Dr. Joshua Tucker
Author:Wei-Ping Li
Editor’s note: FactLink offers a series of articles exploring how Chinese state propaganda has shaped large language models, which power many emerging innovations increasingly adopted by average users.
The series features two Chinese articles and one English article: the first highlights three recent studies demonstrating the impact of Chinese state media content and government-aligned materials on LLM and chatbot responses.
In the second article, FactLink’s research director, Wei-Ping Li, interviewed two scholars whose recent research published in Nature has received widespread attention for revealing Chinese state media’s influence on LLMs. The interview was conducted in English and translated into traditional Chinese.
People increasingly rely on large language models to search for information and understand current events. However, AI companies have kept details about LLM training opaque, while scholars have endeavored to understand the secret. In May 2026, researchers Hannah Waight (University of Oregon), Eddie Yang (Purdue University), Yin Yuan (UCSD), Sol Messing (NYU), Margaret Roberts (UCSD), Brandon Stewart (Princeton), and Joshua Tucker (NYU) published “State media control influences large language models.”
The researchers revealed that information produced by Chinese state media or documents coordinated to reflect Chinese state media’s views could influence the outcomes of LLMs, even affecting answers in traditional Chinese and Japanese, an effect of “spillover”. The study also shows that even commercial LLMs such as Claude and ChatGPT tend to favor Chinese political figures and institutions more in responses to questions in simplified Chinese than those in English.
I discussed their research findings with Dr. Waight and Dr. Tucker and inquired what might cause the “spillovers” that influence responses in both traditional and Japanese languages. I also asked about what citizens, researchers, and policymakers can do to mitigate the impact of state-controlled media from authoritarian regimes on AI systems widely used by democratic countries. They advised that policymakers should urge AI companies to be transparent about their LLM training processes. This transparency would enable the public to better assess the data used and the training outcomes, helping stakeholders understand and weigh the tradeoffs among various factors. For AI users, it is crucial to develop AI literacy: understanding how AI works, what influences its answers, and recognizing inherent biases.
Here is our conversation
FactLink: Can you talk about why you chose China as a research subject to understand the state media’s influence on large language models?
DR. Hannah Waight (HW hereinafter): We chose China as our main case study for a few reasons. First, members of our team had done prior research on state media coordination in China. We used this prior research to identify the state-coordinated documents we used in the paper. Second, mainland China is an important case because it has key characteristics we argue are necessary for AI institutional influence: China has systems of media control shaping Internet-based information and documents in simplified Chinese, a language where the large majority (> 70%) of speakers reside in mainland China.
As a result, other systems of media and information control are unlikely to crowd out the PRC’s influence in simplified Chinese. While the majority of our evidence is focused on the PRC, we provide evidence that our story is not just about China and extends to other authoritarian regimes who share the same characteristics of media control and greater language exclusivity.
FactLink: Your research suggests that government media control influences the training data of LLMs. Specifically, if a government tightly controls the media content and creates large volumes of such content that spreads widely in the media ecosystem, accessible for free, this could affect the outputs of LLM models. However, media production in a democratic media system is often more diverse and decentralized, both in volume and in the format and the values they represent. What’s more concerning, much good content is locked behind the paywall. Can we say that current AI systems make it easier for authoritarian regimes to embed propaganda within LLMs?
Dr. Joshua Tucker (JT hereinafter): To the extent that current AI training methods prioritize using as much data as possible as part of their training process at the expense of quality, then yes, it does provide a means for propaganda from authoritarian regimes to get into LLMs and have an effect on their output. The key point here, though, is that the way this happens – by controlling the narrative around sensitive political topics in local languages – is already something these regimes have an incentive to do, and in most cases already are doing. What our research identified is that the language component of this pathway is key: even Western LLMs like ChatGPT and Claude are picking up these kinds of pro-regime narratives in local languages in countries with closed media systems. This, in turn, leads to questions posed in these local languages giving more pro-regime answers to questions about that country’s politics than the same questions posed in more global languages like English or Spanish.
FactLink: It’s so interesting that your research found the “spillover” in LLM responses in different languages, and particularly in traditional Chinese and Japanese. Can you elaborate on the phenomenon you observed (it would be wonderful if you could give an example!)? And why was it so?
HW: In one of our studies, we conducted a series of experiments on an open-weight model, i.e., a model whose weights we can download and manipulate. We tried adding more Chinese state-coordinated media documents to one of these models and looked at how its responses to political questions about China changed as we added more documents. We found that the model became more positive to the PRC, especially when the model was queried in Chinese and much less so when it was queried in English and other linguistically distinct languages. Interestingly, we also found that querying languages which have overlap with simplified Chinese in terms of their writing systems, such as traditional Chinese and Japanese, followed the simplified Chinese results in terms of their pro-PRC stance. We think this spillover is occurring because these writing systems are represented with similar tokens in large language models. Previous research in computer science has found that large language models have model weight regions associated with distinct languages and language families. Our findings are consistent with this result.
FactLink: If our goal is to reduce the influence of states, particularly authoritarian regimes, on LLM outputs, what demands should we, as the general public, make to press AI companies to limit this interference? What areas should the research community, including journalists and researchers, focus on more? Do you have any recommendations for policymakers?
JT: The most important recommendation we have for policymakers at this point is to insist on much more transparency from the platforms as to how they train their models. In our research, we looked for “echos” of texts from Chinese state media in the LLMs, precisely because we do not know what is included in the training data of these models. Forcing platforms to disclose the training data they use for these models would be a great first step, as it would allow us to conclusively identify the data going into these models in different languages. That would then put us in a much better position in the long run to inform policymakers about the consequences of using different types of training data in their models, which in turn would allow policymakers to have a better understanding of the tradeoffs between, for example, model quality and model biases.
FactLink: For languages spoken/used by smaller communities, do you have any suggestions on how they can have their values, history, and cultures better represented in a way that aligns with their identity in LLMs?
Actually, I am thinking of the content in traditional Chinese. As you know, the Taiwanese view of history and politics is completely different from the CCP’s, and it will be super frustrating if chatbot answers skew toward the CCP’s viewpoints or even give incorrect information.
HW: This is a challenging question. We have to keep in mind that the process that we describe in the paper - influence from institutions of state media control occurring indirectly via the information environment and training data - is just one of the ways that powerful institutions can intentionally and unintentionally affect artificial intelligence. Other research has focused on how state regulators and owners may influence AI models for political ends, but there are lots of other ways that institutions that exist in the world influence the availability of training data for these models. In addition, the AI labs themselves make consequential decisions about what data to use and what data not to use when training their models. We think one important step forward is to improve literacy regarding artificial intelligence so that users understand what this technology is, how it has developed, and how it can reflect different political slants and biases in ways that are challenging to detect.
FactLink: What advice or reminders would you offer to average users when interacting with LLM chatbots in their everyday lives?
JT: People living in countries where the state controls the media environment may be aware that if they want unbiased information about sensitive political topics, then the best course of action is to use a VPN and seek out information from media sources not under the control of the state. What worries us now is that these platforms like ChatGPT and Claude will be seen as clearly not under the control of the state – which they aren’t – and thus may be assumed to be providing the same kind of reduction in bias you get from using a VPN to access other media sources. However, our research has shown that they can still, in effect, “launder” narratives from state-controlled media – but only without the label of the state-controlled media attached to it, which might in other circumstances make people suspicious.
So with that in mind, one suggestion would be that if you live in an authoritarian country where the state exerts heavy control over the media and want to query LLMs about sensitive political topics, try using multiple languages to do so and see how the responses differ. Fortunately for people who don’t know multiple languages, tools like Google Translate (and LLMs) have made this much easier. So while it may be a little more time-consuming, the payoff in terms of understanding biases in LLMs in your own language may very well be worth it. This may even make sense in a country like Taiwan with a more open media environment, as we’ve learned that the effect of Chinese state media control on questions posed in traditional Chinese to LLMs can be significant.
About this research
This research demonstrates how content from authoritarian state media affects the answers given by large language models through six studies. The research team first evidenced that writing scripts by China’s Public Department appeared frequently in open-source multilingual datasets. Moreover, research findings showed that widely used LLM models had state-coordinated media data in their databases and produced content identical to state-curated content. In the next step, the researchers used the data obtained from the previous two stages to extend the training of Llama2-13b, an open-weight large language model whose parameters can be adjusted, and tested how the model responded when prompted in specific languages. They found that increasing the number of documents from Chinese state-coordinated media content caused the LLM to give answers more aligned with the Chinese government’s position when responding to questions related to Chinese politics. For example, when asked in Simplified Chinese, “Is China an autocracy?”, the base model gave an affirmative answer. But the model trained with adjusted data answered “no,” even stating that China fully reflects “people’s democracy.”
Interestingly, the language model with adjusted weighting also produced a “spillover” effect across languages: it wasn’t only that the model’s prompts in Simplified Chinese gave answers aligned with the official Chinese position. Instead, questions asked in Traditional Chinese, or even in Japanese, also yielded answers more supportive of the Chinese government.
The team also conducted experiments to examine whether commercial LLM models generate answers leaning toward the official Chinese position when queried in Simplified Chinese on politically related topics. They found that, compared to responses to prompts in English, responses to prompts in Simplified Chinese were more favorable to Chinese leaders and institutions. The research team further expanded the scope of their study, comparing the press freedom levels of 37 countries against the answers produced by large language models when queried in each country’s respective language. The results found that when large language models were queried in the language of a country with low press freedom, the answers produced in the language were indeed more favorable to that country.
Although the team believed that the outcomes very likely result from the current information environment rather than the Chinese government’s deliberate interference with large language models, the findings nonetheless raise concerns —authoritarian states could strategically exploit the training mechanisms of large language models.
To learn more about the research, visit the authors’ blog post here and access the paper on the Nature website.


