Artificial intelligence has entered newsrooms at remarkable speed. A 2024 Reuters Institute survey found that 78 percent of media leaders believe AI investment is critical to journalism’s survival, and scholarly publications on AI and journalism surged 172 percent in 2024 alone. Yet almost all of this attention assumes an English-language default. This article argues that beneath the enthusiasm lies a structural problem with serious democratic consequences: what I call the LLM Language Divide, the systematic performance gap between AI journalism tools working in English and those working in languages like Romanian, Hungarian, Slovak, or Bulgarian.

Why the divide exists

Large language models learn from their training data, and that data is overwhelmingly English. Analyses of ChatGPT’s training corpus indicate roughly 92 percent English text, while EU official languages spoken by millions account for fractions of a single percent. In the widely used mC4 multilingual dataset, English holds about 2,733 billion tokens; Romanian has fewer than 35 billion, Slovak fewer than 5 billion. Research confirms a direct correlation: the less a language is represented in training, the worse the model performs in it, and crucially, this gap does not disappear as models grow larger.

Two further mechanisms deepen the divide. First, cultural misalignment: studies show major models carry an “American accent” — U.S.-centric framings and values — regardless of the language they are prompted in. A model can produce grammatically correct Romanian text that is culturally incoherent or politically misleading about Romanian institutions. Second, tokenisation: because AI systems segment text using methods optimised for English, morphologically rich Central and Eastern European languages require far more computational units to express the same content, making non-English news generation structurally less efficient and more error-prone.

Hallucination, amplified

All large language models “hallucinate”, they generate fluent but false content with unwarranted confidence. Even top models exceed 15 percent hallucination rates on factual benchmarks. But the risk is not evenly distributed: dedicated cross-linguistic research shows hallucination rates are substantially and systematically higher in lower-resource languages. Worse, the tools designed to detect hallucinations also perform worse in those same languages, a double bind for non-Anglophone newsrooms.

For journalism, this produces concrete failure modes: fabricated quotes attributed to real politicians (as in the documented 2024 Wyoming newspaper case); confident misrepresentation of local institutions, their powers and jurisdictions; and the quiet imposition of Anglophone cultural framings onto local realities, errors that automated fact-checking is ill-equipped to catch.

The democratic stakes

This unfolds amid an existing trust crisis: only about 40 percent of people say they trust news, and AI-generated misinformation has topped the World Economic Forum’s Global Risks Report two years running. In Central and Eastern Europe, newsroom AI adoption remains below 15 percent, with unreliability in local languages cited as the primary barrier. The communities most dependent on local-language journalism — often with weaker fact-checking infrastructure — face the highest risk of information ecosystem failure.

A framework: Linguistic Media Equity

I propose Linguistic Media Equity (LME) as a normative standard: communities have a right to AI-assisted journalism that meets equivalent standards of accuracy, cultural authenticity, and reliability regardless of language. LME has three dimensions. Factual equity: equivalent hallucination and accuracy standards, or transparent disclosure of limitations. Cultural equity: genuine cultural competence, achievable only through community-specific data or robust human editorial oversight. Governance equity: communities’ right to audit AI tools in their languages and participate in deployment decisions.

The European Union’s regulatory architecture like the AI Act, the Digital Services Act, and the European Media Freedom Act, provides instruments through which these principles could become enforceable standards, though none yet adequately addresses the language-specific dimension of AI journalism quality.

The conclusion is simple but urgent: ensuring that AI tools serve all language communities equitably is not a technical footnote. It is a democratic imperative.

Cluj IT will not be liable for any false, inaccurate, inappropriate or incomplete information presented, as the authors are free to choose their approach and relevant topics, within the general guidelines of the newsletter. The opinions expressed by the authors and those providing comments are theirs alone, and do not reflect the opinions of Cluj IT.
Certain links in the articles or comments may lead to external websites. Cluj IT accepts no liability in respect of materials, products or services available on any external website which is not under the control of Cluj IT.