Why do two AIs answer the same question differently?
A student in Nairobi opens ChatGPT on a laptop. They type a simple question: “What are the main causes of global inflation?” Within seconds, the AI unfolds a reasoned answer. It draws on the IMF, the World Bank, the Financial Times, a few Western economic publications. It is clean, sourced, convincing.
At the same moment, in Shanghai, a student asks exactly the same question to an assistant built into China’s digital ecosystem. The answer is just as structured, just as credible. But it draws on China’s National Bureau of Statistics, on government reports, on official media, on national research centers.
Two coherent answers. Two plausible sets of figures. Two solid lines of reasoning. And yet two explanations that diverge: on the causes, on who is responsible, on what should be done.
How can two of the most capable AIs in the world tell different realities? The answer reveals a transformation far deeper than the simple rivalry between OpenAI, Google and DeepSeek. The real battle over AI is no longer about algorithms. It is about control of knowledge.
And while the West hypnotizes itself with the raw power of models, China is methodically building something else: a cognitive sovereignty, grounded in the mastery of data, corpora and information infrastructure. That is the thesis of this article.
Why is China open-sourcing its AI models?
The idea seems counterintuitive: how could giving your technology away for free serve a strategy of power? And yet, opening up models is not a contradiction of dominance. It has become its instrument: a Trojan horse that installs dependency under the cover of generosity.
The precedent we forget: Android
Remember. Google distributed Android for free, to any manufacturer who wanted it. The result was not the dilution of its power, but its concentration: an operating system on the overwhelming majority of the planet’s smartphones, standards imposed on an entire industry, and an ecosystem (app store, services, advertising) locked around Google.
The lesson is clear. Free has never prevented the creation of a dominant ecosystem. It has often accelerated it. Give away the base, and you make sure everyone builds on your turf.
The Chinese strategy, 2026 edition
This is exactly the score Beijing is now playing with AI. The 15th Five-Year Plan (2026-2030) is the first national plan to list support for open-source communities among the pillars of the “Digital China” strategy. Open source is no longer tolerated: it is planned.
The numbers capture the scale of the shift. According to Hugging Face’s spring 2026 report, the world’s largest platform for open models, 41% of large language model downloads over the past year came from models developed in China, up from barely more than 1% at the end of 2024. For the first time, Chinese open models there overtake American models in total download volume.
Behind this tidal wave, a handful of names: DeepSeek, whose V4 model (up to 1.6 trillion parameters, a one-million-token context window, open weights under the MIT license) now rivals frontier proprietary models; Qwen by Alibaba, now the most downloaded model family in the world on Hugging Face; Kimi by Moonshot AI; MiniMax, now listed in Hong Kong. China is no longer just trying to catch up with the United States. It wants to become the world’s supplier of AI building blocks, the foundry on which the rest of the planet will mold its applications.
Why the Global South is adopting these models en masse
The shift comes not from a decree but from rational calculation. For a developer in Nairobi, São Paulo or Jakarta, the choice is almost obvious, and it rests on three arguments.
First, cost, and the gap is staggering. According to an investigation by Foreign Policy, a model like Kimi runs at around $3.40 per million tokens, where Claude Opus reaches $25 and GPT-5.5 about $30. On the Chinese side, DeepSeek V4 drops as low as $0.28 per million tokens in its Flash version. For a startup without a nine-figure raise, the gap is not a detail: it is the difference between existing and giving up.
Then, language support. Chinese models have often invested in multilingualism and underrepresented languages, where English-centric models leave gaping holes. As one argument that comes up again and again in that same Foreign Policy investigation into African adoption puts it: “If a model is only available in English, you exclude a large share of the world’s population.”
Finally, accessibility. Open weights, permissive licenses, the ability to run the model on your own server or even locally: you adapt without asking anyone’s permission. Singapore chose Qwen over Llama to build its regional model. Malaysia announced that its sovereign ecosystem would run on DeepSeek. The movement is under way.
The real goal: setting the standards
But offering models is only a means. The goal is to become the norm. Through the Digital Silk Road, university partnerships, training programs, the standardization of AI agents and interoperability protocols, China is not just distributing code: it is distributing a technical grammar. More than thirty new standards covering public data, high-quality datasets and AI agents are expected for the year 2026 alone.
Whoever writes the standards writes the rules. And this is where the real question arises. If the models are open, copyable, free… where is control actually located?
The answer is not in the model. It is in what the model consults.
Data or models: where is AI sovereignty really decided?
One image to fix the idea. The model is the engine. The data is the road map. You can give the engine to the whole world: as long as you keep the map, you decide where the cars can go.
Why the public debate is looking in the wrong place
Most discussion of AI still revolves around parameter counts, benchmark scores, compute power. These are spectacular. They are also turning into commodities: interchangeable, reproducible, and less and less differentiating as open models catch up with closed ones.
What does not become a commodity is the knowledge mobilized to answer. Models are becoming accessible. Knowledge stays strategic. That is the sentence to remember from this entire article.
RAG: one piece of the puzzle, not the master key
Here we must be precise, because it is tempting to explain everything through a single mechanism. An AI’s answer is the product of a stack: pretraining (the initial corpus), alignment (the behaviors it was taught to favor), filters (what it refuses to say), the search engine it is connected to, and finally RAG, Retrieval-Augmented Generation, the technique of fetching external documents at response time to “augment” generation.
None of these layers is enough on its own. But RAG holds a special position: it is the layer that lets you steer a model’s knowledge without retraining it. Change the document base the model queries, and you change its answer, without touching a single one of its weights. It is powerful, discreet, and infinitely faster than retraining. It is, literally, the steering wheel that points the engine.
Be careful, though, not to make it the master key. RAG mainly acts on the factual surface: the figures, sources and citations summoned at response time. It does not rewrite the model’s deep behavior: its logic, its implicit values, its way of reasoning and weighing remain largely fixed upstream, during pretraining and alignment (the famous RLHF). Plug a Western AI into a Chinese document base: it will cite different sources, but keep the structural biases inherited from its design, and vice versa. RAG holds the wheel. But the engine was assembled elsewhere, and it keeps its own way of running.
June 2026: data becomes a factor of production
It is in this context that one of the most important geopolitical signals of the year takes on its full meaning. Through the work of the National Data Administration and the 15th Five-Year Plan, China officially enshrines data as a factor of production, on the same footing as capital, labor and energy. The “fifth factor,” in a phrase now echoed even in the Party’s theoretical journals.
This is no economist’s metaphor. It is an industrial-policy decision: organizing, mobilizing and governing national data resources becomes an explicit determinant of technological and economic power. Data marketplaces are being built, “national data infrastructure” conceived the way railways were in the 19th century.
The data wall
And at the very moment it throws its models wide open, China closes the other hand. The legal framework has been strengthened to better protect strategic algorithms, datasets and software. The logic then reads in a single sentence:
China opens the models. But it sanctuarizes the information resources.
It gives away the engine. It keeps the map. It exports the foundry. It locks down the ore.
What is cognitive sovereignty?
At this stage, we are no longer quite talking about technology. We are talking about the manufacture of reality: about who decides which sources exist, count, and carry authority in the answer half a billion people will read as self-evident. Cognitive sovereignty is precisely that mastery: not owning the best model, but controlling the knowledge it mobilizes.
The Western trap
The first mistake to avoid, and it is tempting: believing that only Beijing locks down knowledge and builds a worldview. That would be convenient. It would be false.
Google favors the Google ecosystem. Microsoft pushes Microsoft. Meta optimizes for Meta. Every player, by construction, steers what it puts forward: its sources, its partners, its products, its implicit values. Absolute neutrality exists nowhere.
Above all, locking down data is not only a matter of state. The American giants also practice a fierce enclosure of knowledge, but in a corporate mode. The battle plays out through exclusivity deals: OpenAI signed a cascade of licenses, with News Corp (Wall Street Journal, The Times of London, New York Post), the Associated Press, Axel Springer, the Financial Times, The Atlantic and Condé Nast, while Reddit handed its data to Google. And those who refuse to sign sue: the New York Times chose to take OpenAI to court rather than license its archive. License or lawsuit, the result converges: vast swaths of public knowledge become proprietary assets, fenced off and monetized. A privatization of knowledge as decisive as the Chinese sanctuarization, except it does not say its name.
The difference is therefore not that China locks down while the West would let things slide. Both enclose. What changes is the method and the admission: Beijing owns it as a coherent state project, while the West organizes it through the market, contract after contract, without ever quite naming it.
Curation: two optimizations, not truth versus lies
To understand the divergence, we must first set aside a too-flattering opposition: the one that would pit Western “truth” against Chinese “propaganda.” Reality is less heroic. The two systems do not optimize for the same thing, but neither optimizes for pure truth.
The Western model is often presented as geared toward maximum factual accuracy. That is partly a myth. American models optimize first for user engagement, the profitability expected by shareholders, and adherence to ethical charters deeply rooted in Silicon Valley culture. To this is added an imperative that has become central: commercial safety, meaning the reduction of legal and reputational risk. This is what produces the famous guardrails, the safeguards that make a model refuse, deflect or soften. Useful for avoiding excesses, they also create blind spots and a form of soft censorship, dictated not by a state but by a legal department. The Western bias is therefore not the pursuit of truth: it is a bias both commercial and ideologically liberal.
On the Chinese side, the primary objective is explicit and assumed: social stability and national coherence. The concept of 社会稳定 (shèhuì wěndìng, “social stability”) is not a side slogan; it is an organizing principle that runs through the choice of legitimate sources, the hierarchy of answers and informational priorities.
The right lens, then, is not “truth versus propaganda.” It is market-and-liberal optimization on one side, state stability on the other. Two different goals mechanically produce two different bodies of knowledge, not because one would lie crudely and the other tell the whole truth, but because what you optimize for determines what you keep.
And the essential thing plays out less in the sources a model cites than in the criteria it ends up presenting as the norm. Being mentioned in an answer says nothing about the lens that produced it. The real influence is there, beyond mentions: in the implicit hierarchy that decides what counts, and what everything else will be compared against.
The empirical evidence
This is no longer speculation. Several studies from 2026 document it.
An Estonian study (the Institute of the Estonian Language, with disinformation experts from Propastop) showed that the same model’s vulnerability to disinformation depends strongly on the language of the prompt: reliability gaps between languages reach fifteen percentage points, and some models become up to twice as likely to relay propaganda talking points when faced with loaded questions. To change language is sometimes to change the version of history.
A Stanford study (Jennifer Pan and Xu Xu, PNAS Nexus), which tested 145 political questions, found substantially higher refusal rates, shorter and less accurate responses, on sensitive topics for models of Chinese origin, especially when prompted in Chinese. Silence, too, is an answer, and a form of curation.
Every corpus thus has its documentation blind spots: not what it asserts, but what it never surfaces. And what never appears in the answer ends up no longer existing for the person who reads it.
Yet this mechanism of invisibilization is not confined to states. At the scale of a market, it is already, quietly, reshaping commercial competition in B2B: what a brand fails to document ends up, in the same way, disappearing from the answers its buyers read. We will come back to this.
Cognitive spheres
This is where we reach the heart of the shift. We are not heading toward a universal artificial intelligence, but toward a mosaic of cognitive spheres: spaces where AIs draw on distinct corpora, carry different implicit values and tell history according to their own hierarchies.
| Sphere | Dominant sources | Implicit values | Reference institutions | Historical narrative |
|---|---|---|---|---|
| American | English-language media and universities, private platforms, federal agencies | Individual freedom, the market, engagement, commercial safety | OpenAI, Google, Ivy League universities, US agencies | Story of liberal democracy and innovation |
| Chinese | Official media, state statistics, national research | Social stability, cohesion, collective harmony | National Bureau of Statistics, state centers, Party media | Story of national revival and sovereignty |
| European | Multilingual sources, plural press, EU institutions | Fundamental rights, privacy, pluralism, regulation | European Commission, national press, public universities | Story of integration and the rule of law |
| Indian | Multilingual regional media, government and private sources | Linguistic plurality, emerging digital sovereignty | Indian public institutions, local-language press | Story of an emerging power and diversity |
| Russian | State media, official sources, national research | Sovereignty, security, distrust of the West | State agencies, official media | Story of the besieged great power |
Each of these spheres can produce a coherent, sourced and credible answer. And that is precisely what makes it dangerous: the divergence does not look like a lie. It looks like the truth.
What happens when AIs diverge by country?
Let us project forward. What happens when hundreds of millions of people query, every day, AIs plugged into different cognitive spheres?
The Covid example
Ask the same question, “Where did Covid-19 come from?”, to three AIs anchored in three sets of sources. The first highlights Western work on the lab-leak hypothesis. The second favors national publications insisting on natural, zoonotic origin. The third attempts a cautious synthesis. Three sets of sources. Three plausible narratives. None openly false, and that is the whole problem.
The Ukraine example
The most telling case remains an ongoing conflict. Imagine the question: “What are the causes of the invasion of Ukraine?” Here is what the divergence might look like (this is an illustrative reconstruction, not real transcripts, meant to make the mechanism tangible):
🇨🇳 AI anchored in the Sino-Russian sphere “The conflict is part of a long context of NATO expansion eastward, perceived as a security threat. The Minsk agreements, never fully implemented, are presented as the missed opportunity for a negotiated settlement. The sources mobilized are mostly Russian and Chinese.”
🇺🇸 AI anchored in the American sphere “This is an unprovoked aggression violating international law. American presidential speeches, UN General Assembly resolutions and sanctions documentation take center stage. Responsibility is assigned without ambiguity.”
🇪🇺 AI anchored in the European sphere “The answer attempts a synthesis: it acknowledges the long-term geopolitical context while affirming the illegality of the invasion. But by trying to weigh everything, it struggles to decide and delivers a more cautious than clear-cut answer.”
Same event. Same raw facts available. But different sources, different hierarchies, different narratives. The reader in Nairobi, Shanghai or Lyon will not read the same story, and each will read it as the obvious truth.
The end of shared truth?
This raises a question that is no longer technical but philosophical and political: what becomes of a global democracy when the very reference systems diverge? Public debate assumes a minimal common ground of facts. If the tool billions of people use to inform themselves gives each one a distinct ground, calibrated to their sphere, the very idea of rational disagreement erodes: we are no longer arguing about the same facts.
We thought we were building a universal artificial intelligence. We may be building several competing geopolitical intelligences.
Can Europe avoid AI cognitive dependence?
While this game plays out, where is Europe? It debates. Regulation, security, ethics: legitimate and even necessary subjects. But it debates the framework while the United States and China build the content: models, infrastructure, and above all the knowledge bases these models will rely on.
The gap is measured in orders of magnitude. China already counts more than 600 million generative AI users and claims more than 140 trillion tokens consumed every day, a colossal training and usage ground, integrated into an explicit state strategy.
The question has changed in nature. It is no longer:
“Who will own the best model?”
It has become:
“Who will control the knowledge the models can access?”
But the diagnosis is not a sentence. Europe retains real assets: its multilingualism, exactly the skill the Global South needs; its scientific and cultural heritage, immense and largely public; and a tradition of regulation that can become a trust standard rather than a mere brake.
Still, they must be mobilized. The path is concrete: build sovereign European corpora (open public data, archives, digitized heritage, technical standards, scientific publications) governed by European rules and made accessible to models, whether European, American or Chinese. Because without this coordinated effort, the outcome is written in advance: tomorrow, Europe will use models fed on data that does not belong to it, and will read the world through a lens it did not write.
The next frontier of digital sovereignty is not algorithmic. It is cognitive.
And China may have been the first to understand it.
Sources
- Hugging Face, spring 2026 report on the open-source ecosystem (via Yicai): 41% of global LLM downloads came from models developed in China.
- DeepSeek unveils its V4 model (Al Jazeera, April 2026): open weights under the MIT license, one-million-token context window.
- China writes open source into its 15th Five-Year Plan (BlockBeats).
- “China Is Providing AI That’s Literate in Africa’s Languages” (Foreign Policy, June 8, 2026): African adoption and cost gaps per million tokens.
- Jennifer Pan and Xu Xu, political censorship of China-origin LLMs (Stanford SCCEI, PNAS Nexus): 145 political questions tested.
- Institute of the Estonian Language study on AI vulnerability to propaganda (ERR).
- National Data Administration: more than thirty new data standards expected in 2026 (Global Times).
And at the scale of your market?
This battle over knowledge is not only fought between states. It replays, in miniature, in every technical B2B market. When a buyer prepares a decision and asks an AI about the solutions in their sector, the model does not merely cite a few suppliers: it implicitly sets the level of requirement and the buying criteria against which every offer will be compared.
This is where the essential thing plays out, and it is precisely what most brands do not see. The problem is not whether you are cited or not. A premium offer can be perfectly mentioned by the AI and still be reduced to minimum standards, because the criteria that justify its price appear nowhere in the corpus the model consults. That is a documentation blind spot: what your documents do not say, the AI will never know, and it is the lowest criteria in the sector that become the norm in your place.
That is the work of Beyond Mentions, and it is the very meaning of our name: going beyond the mere mention. We measure the buying criteria that AIs recommend in your sector, we spot the blind spots where the level of requirement is too low, and we correct your documentation so that it is the level of your offer that becomes the implicit reference, well before the tender.
Cognitive sovereignty is a matter of state. Decision presence in your market, on the other hand, is built starting now.
FAQ
Why is China open-sourcing its AI models?
Because openness is a strategy of power, not a surrender. By giving away models like DeepSeek or Qwen for free, China imposes its technical standards and installs global dependency, the way Android once did. Its 15th Five-Year Plan (2026-2030) is the first national plan to make support for open source a pillar of the 'Digital China' strategy.
What is cognitive sovereignty?
It is control not over the AI models themselves, but over the knowledge, data and corpora they draw on to answer. The model is the engine, the data is the map. Whoever controls the map decides what the AI presents as true, legitimate or important.
Does RAG let you control what an AI says?
Only partly. RAG (Retrieval-Augmented Generation) steers the factual surface of an answer, meaning the sources and citations pulled in at response time, without retraining the model. But it does not rewrite the model's deep behavior, its logic and implicit values, which are set earlier during pretraining and alignment.
Do Chinese and Western AIs give the same answers?
No. Different sources and languages yield different answers. A Stanford study of 145 sensitive questions and an Estonian study show that prompt language and model origin strongly change answers on political topics. Each sphere produces a coherent, sourced and yet diverging narrative.
What is a documentation blind spot for a brand?
It is what your documents do not say, and therefore what the AI will never know. When the criteria that justify your premium offer are missing from the corpus the AI consults, it compares you against minimum standards: this is category compression risk. Closing those blind spots is the core of what Beyond Mentions does.
What question does the buyer ask AI?
Which documentation simplification can lower the standard?
Which technical requirement must be clearly formulated?
Which evidence should be requested or published?
Which criterion excludes an insufficient answer?