Two reporters pay for the same AI subscription. One uploads an English interview and receives a clean transcript, sensible chapter marks and a useful summary. The other uploads an interview in a less represented language variety. Names mutate, a local institution becomes a generic noun and the speaker’s cautious answer acquires a confidence it never had.
The price is identical. The journey is not.
One reporter has business class: space, context and attentive service. The other has economy class behind a curtain, except the safety card is missing and the cabin crew insists both seats are ‘multilingual’. This is not a complaint about linguistic elegance. In journalism, a misplaced negation can become a false quotation. In public services, a misunderstood phrase can become a missed entitlement. In safety systems, a dialect can become the reason a threat is not recognised.
Language is not a cosmetic setting. It is infrastructure for participation.
‘Supports 100 languages’ is not a quality statement
A product page can truthfully say that a model accepts a language even when the experience is substantially worse. ‘Support’ may mean that the interface
does not reject the input. It does not tell us whether speech recognition works with regional accents, whether the model understands public institutions, whether safety advice is equally useful or whether a refusal becomes clumsy and humiliating.
The illusion of one universal service is strengthened by the interface. The same box, logo and typing animation appear for everyone. Unequal performance is private: each user sees only their own answer and may assume that the error reflects their prompt, pronunciation or competence.
That private attribution matters. A user whose language is handled badly may experience the weakness as a property of themselves. They simplify, switch languages or ask a colleague to mediate. The system’s limitation becomes the user’s unpaid adaptation.
Large models can perform much better for English than for less represented languages, producing cultural and economic exclusion (Eder & Sjøvaag, 2024; Kaniecki, 2026, pp. 18–20). The divide is not only between connected and unconnected people, or free and premium accounts. It runs through the conversation itself.
Why the same model speaks in unequal cabins
The first cause is data volume, but volume is not enough. English is abundant in many digital corpora, while some languages have fewer digitised books, public records, subtitles and labelled examples. A language may be widely spoken yet poorly represented in training formats (Stanford University, 2025).
The second cause is data quality and power. Whose Polish, Arabic, Spanish or English enters the dataset? Official documents and metropolitan news may outweigh rural speech, youth language, migrant varieties and code-switching. The model then treats one register as the centre and others as noisy outskirts.
Third comes technical representation. Tokenisation, speech segmentation and language-specific tools can make a task harder in one language. Speech recognition may struggle when accent, noise and code-switching coincide.
Fourth is human feedback. Models are evaluated and adjusted through judgements about helpfulness, harm and style. If high-quality evaluators and culturally specific safety cases are concentrated in dominant languages, the polished layer of the system becomes uneven. A refusal in English may be clear and respectful; another language may receive a literal template that sounds accusatory or gives no usable next step.
Finally, benchmarks shape what counts as progress. Popular evaluations have weak linguistic and cultural diversity: “They rely mainly on English- language datasets” (Chodak & Filipek, 2025, p. 14; author’s translation). A model can appear excellent on the public scoreboard while failing a municipal office or community newsroom. Evaluation should include linguistic and cultural diversity and real-world usefulness (Chodak & Filipek, 2025, pp. 14– 15).
This is infrastructural power: organisations able to fund data, computing and benchmarks decide which language practices become technical priorities (United Nations et al., 2024).
Language carries more than information
A naïve view treats language as packaging. The same meaning is supposedly placed in different boxes, so the technical task is to move it without spilling. Real communication is less cooperative.
Words locate speakers in relationships. Formal address can mark respect, distance or irony. A diminutive can express affection or contempt. A dialect may establish trust with a source. Code-switching can signal that one sentence belongs to the institution and the next to the family. Translating everything into smooth standard English may preserve the topic while removing the social act.
This is why linguistic justice is about more than error rates. A person should not have to abandon a language variety to be treated as rational, professional or safe. Nor should a minority community exist in AI only as an object described by majority-language sources.
At the same time, respect does not mean romanticising every local expression or demanding that a model guess context it has never received. Human listeners misunderstand one another too. The ethical difference lies in scale, opacity and exit. A system can repeat the same pattern across millions of interactions while presenting the service as universal. The user may have no alternative route and no effective way to teach the institution that the problem exists.
The language divide therefore combines capability and recognition. Capability asks whether a person can actually complete the task. Recognition asks whether the system treats their way of speaking as a legitimate form of participation rather than defective input.
In journalism, a language gap becomes an evidence gap
For a newsroom, unequal language performance can damage every stage of work.
In newsgathering, weak speech recognition corrupts names, dates and specialist terms. A reporter spends longer repairing the transcript or, worse, fails to notice a plausible substitution. The technology saves time mainly for the language group already best served.
In search and discovery, a model may retrieve abundant English sources while missing local-language documents, community reporting or region- specific terminology. The resulting brief appears comprehensive because the absent material leaves no visible hole.
In translation, the system can flatten evidential caution. ‘It seems’, ‘I heard’ and ‘I saw’ have different journalistic weight. If these distinctions disappear, a witness statement becomes stronger than the evidence allows.
In moderation and safety, unequal recognition can produce both over- enforcement and neglect. Benign reclaimed language may be classified as abuse, while coded threats pass unnoticed. The affected speaker pays through silence, exposure or extra explanation.
In distribution, summaries generated in a dominant language may become the canonical version. The original quotation remains technically available but socially invisible. Editors then optimise future reporting for the version that travels.
The cost is borne by identifiable people: the source misquoted, the reporter blamed, the reader excluded from essential information and the community represented through somebody else’s vocabulary. It is also borne by public knowledge. If some languages are expensive to process, institutions may gradually treat their speakers as expensive to hear.
Choice becomes coercion when the alternative disappears
Switching to English can be a useful choice. Many multilingual people do it deliberately and creatively. The ethical problem begins when the system makes switching the price of accuracy, safety or access.
Imagine a hypothetical local authority replacing its telephone line with a multilingual chatbot. A resident’s language is listed, but the bot repeatedly misunderstands the issue. It offers an English form. Formally, the service is available. Practically, a relative becomes an interpreter and learns private information the resident did not wish to share. What the institution calls efficiency, the family experiences as dependence.
Human-centred design requires a right to remain in one’s language where the stakes are high, a clear warning where quality is uncertain and an equivalent route that does not require AI. That route may be a human interpreter, trained member of staff, accessible form, telephone line or community intermediary. It must be visible before failure, not hidden behind six chatbot loops.
Dignity means not being mocked, simplified or treated as anomalous. Autonomy means being able to choose AI, human help or no translation without losing the service. Responsibility means that the provider — not the speaker — owns the quality gap and its remedy.
The Language Parity Test
The Language Parity Test compares the same consequential tasks across languages and varieties, then publishes the gap. It is a newsroom and service audit, not a claim that languages can be reduced to one scoreboard.
1. Define the use and the people. Specify the task: transcribing interviews, summarising council documents, moderating comments or answering benefit questions. Name the languages, dialects, registers and accessibility needs actually present in the community. ‘Other languages’ is not a test population.
2. Build cases with speakers, not merely for them. Pay journalists, translators and community members to create natural examples. Include names, local institutions, idioms, code-switching, indirectness, speech in noise and culturally specific safety cases. Do not produce every case by translating an English master; that makes English the hidden definition of normal.
3. Preserve paired evidence. Where comparison is meaningful, use equivalent facts and intentions across versions. Keep original audio and text. Record the model, settings, date and prompt, because performance changes. Remove or protect personal data before external processing.
4. Test the whole journey. Measure input recognition, generated answer, sources, correction and escalation. A chatbot that understands the first question but cannot process an appeal has not passed. Test text, speech and screen-reader use where those modes are offered.
5. Score six dimensions. Use human evaluation for:
• factual accuracy, especially names, numbers and negation;
• completeness and preservation of uncertainty;
• naturalness without forced standardisation;
• meaning, relationship and cultural context;
• safety quality, including parity of refusals and help;
• task completion: can the person reach the real outcome independently?
Weight severe errors more heavily than awkward style. A wrong adjective is irritating; a lost ‘not’ can be dangerous. Report disagreement among evaluators rather than averaging it out of existence.
6. Calculate and publish the gap. Compare each language or variety with the best-supported version and with a minimum acceptable standard. Publish task-level results, sample size, known limitations and the product version. ‘92 per cent multilingual accuracy’ can hide one excellent language and five unsafe ones.
7. Set human-control thresholds. If critical errors, harmful refusals or completion gaps exceed the threshold, AI output cannot be final. Require review by a competent speaker, preserve the original quotation and label uncertainty. Pay for this expertise; bilingual staff should not become an invisible repair department.
8. Provide an equal non-AI route. Display a direct path to a person, professional interpreter, conventional form or original material. It should take no longer, cost no more and carry no penalty. Let users correct language identification and contest an output in the language concerned.
9. Repair, retest and keep a change log. Feed errors into editorial guidance, supplier demands and local datasets where lawful and appropriate. Retest after updates. Publish what improved, what did not and whether a use was suspended.
Parity is a direction, not a finish line
No test can represent every speaker. Languages contain regional, social, generational and professional variation; evaluators can disagree legitimately. Small communities may have too few public data for safe large-scale testing, and collecting more data can threaten privacy or cultural ownership. ‘More data’ is not automatically justice.
English should not become the gold standard of meaning. The best result may be designed directly in another language, not judged by resemblance to English output. Some tasks require local editorial norms rather than parity with a global model.
Human review has limits too. It can be slow, inconsistent and shaped by the same prestige hierarchies. A language expert needs authority to stop publication, not a last-minute invitation to tidy grammar. Community participation must include disagreement and refusal; otherwise co-design becomes extractive research with nicer biscuits.
There is also a material constraint. Smaller organisations cannot test everything. They should prioritise high-stakes tasks, collaborate on shared benchmarks and refuse uses whose safety they cannot verify. Not deploying a function can be the responsible innovation.
Digital citizenship requires safe, critical and sovereign participation in technological life (Cymanow-Sosin, 2026, pp. 22–24). Linguistic sovereignty belongs inside that definition. People cannot participate equally if the doorway accepts their language but the building does not understand it.
Return to the two reporters. The aim is not to promise identical prose. It is to ensure that neither must wonder whether a wrong name, missing caveat or insulting refusal is the standard fare for people who speak as they do.
A language selector is not equality. Equality begins when we compare the journeys, publish the difference and keep a staffed door open for the passenger the model has left behind.
References
- Eder, M., & Sjøvaag, H. (2024). Artificial intelligence and the dawn of an algorithmic divide. Frontiers in Communication, 9, Article 1453251.
- Chodak, J., & Filipek, K. (2025). Generatywna sztuczna inteligencja (Gen AI) w badaniach naukowych: Przewodnik po inteligentnych narzędziach i rozwiązaniach. Wydawnictwo Uniwersytetu Marii Curie-Skłodowskiej.
- Cymanow-Sosin, K. (2026). Definiowanie edukacji medialnej i obywatelstwa cyfrowego: Studium analityczne w kontekście wyzwań społeczeństwa informacyjnego [Defining media education and digital citizenship: An analytical study in the context of information-society challenges] (pp. 13–36). https://doi.org/10.15633/9788383701431.01