You ask an AI assistant whether a new regulation applies to your case. Within seconds, it produces a calm, neatly structured answer with three links. The fir st opens the regulator’s home page. The second is a five-year-old article. The third returns ‘404’. The paragraph still looks like a finished expert note: it has a date, an institution and the tone of somebody who has just emerged from the archive carrying the decisive folder.
This is a hypothetical scene, but the problem has many real forms. Sources may be absent, invented, inaccessible, outdated, misattributed or perfectly genuine yet irrelevant to the claim. Sometimes an assistant cites a text that says something adjacent, but not the same thing. Sometimes it supplies twenty addresses, creating bibliographic fog: so many blue links appear beneath the answer that the sheer weight of decoration begins to feel like proof.
The old advice to ‘check the source’ remains sound. It is simply no longer complete. In an age of generated answers, we must also ask which sentence is the claim, what kind of source could establish it, whether the cited material actually does so, and what the absence of a trail allows us to conclude.
When the footnote stopped being the end of the journey
In a conventional text, a footnote makes a modest promise: the journey backwards begins here. It does not guarantee truth, but it lets the reader inspect an author, context and record. A conversational interface reverses that architecture. We receive the synthesis first; sources, where present, arrive as optional accessories. The conclusion reaches us before the evidence.
This is not merely a moral failure by users who could not be bothered to click. A 2025 Reuters Institute survey found that 54% of respondents across six countries said they had encountered an AI-generated search answer in the previous week. Among those respondents, 33% said they always or often followed the links, 37% did so sometimes, and 28% rarely or never did (Simon et al., 2025, pp. 30–33). Half said they trusted such answers, often valuing their speed, convenience and ability to aggregate information (pp. 33–36).
The authors state important limitations. The online survey used non- probability quota samples of roughly 2,000 people in each country, and reported behaviour may differ from actual behaviour (Simon et al., 2025, pp. 12–13). Even so, the pattern matters. Generated synthesis is becoming a destination, not merely a signpost. The source has changed from the floor beneath the answer into a trapdoor in that floor: it can be opened, but most of the conversation takes place above it.
Fluency is not evidence in liquid form
People judge information partly through its presentation. Coherence, specificity and confident delivery ease comprehension and can increase perceived credibility. Online, users combine signals such as expertise, reputation, appearance, recommendations and verifiability; they rarely have the time or knowledge to conduct a full investigation of every claim (Metzger, 2007, pp. 2078–2091).
Generative AI is unusually good at producing the form of a completed answer. It need not ‘lie’ in a human sense to create a false feeling of closure. A language model predicts useful sequences from patterns; it does not
automatically maintain a reporter’s evidence file for every sentence. Even a search-connected system can retrieve a source that is topically close but evidentially wrong.
Research by the BBC and the European Broadcasting Union gives the problem both scale and vocabulary. In June and July 2025, participants from 22 public service media organisations in 18 countries and 14 languages evaluated more than 3,000 news-related responses from four AI assistants. Forty-five per cent contained a significant issue that could materially mislead a user (Yezza et al., 2025, p. 3). Their toolkit identifies four pillars of a good answer: accuracy, adequate context, separation of fact from opinion, and sourcing (pp. 6–7).
Its rule on sourcing is beautifully plain: ‘The assistant should always provide sources to support key claims’ (Yezza et al., 2025, p. 7). Yet support does the heaviest work in that sentence. A link must do more than exist. It must contain information that justifies the particular claim beside it.
I remember the fact – but do I remember where it came from?
Source-monitoring theory distinguishes the content of a memory from information about its origin. It describes the processes by which we attribute knowledge to an experience, person or medium. Errors arise when we retain the content but confuse its source, or fail to separate what we encountered from what we inferred (Johnson et al., 1993, pp. 3–28).
An AI answer can create a technological cousin of that error. It supplies a sentence and places a plausible-looking item nearby. Spatial proximity becomes logical support in the reader’s mind: if the link sits under the paragraph, surely it proves the paragraph. The BBC/EBU taxonomy shows why that assumption fails. It includes missing sources for key claims, irrelevant or outdated material, sources with inadequate editorial control, fabricated links, inaccessible pages, misleading attribution and documents that simply do not contain the cited information (Yezza et al., 2025, pp. 25– 34).
Our epistemic vigilance is involved too. Sperber and colleagues argue that human communication requires mechanisms for assessing both a communicator and a message, because communication offers opportunities for cooperation as well as deception. People are neither infinitely gullible nor perfectly sceptical; they rely on signals of competence, honesty and coherence (Sperber et al., 2010, pp. 359–393). An assistant can simulate many competence signals — order, certainty, patience and verbal range — without bearing personal responsibility for the consequences.
The result is an odd division of labour. The system supplies the confidence. The user must discover whether the confidence was earned.
A sourcing error has a tenant
‘The model gave a bad link’ sounds like a minor office mishap. Now ask who pays. A student includes a non-existent paper and faces an academic-integrity investigation. A journalist copies a figure into a report and an organisation wrongly loses its reputation. A patient treats a concise answer as current guidance although the cited source concerns another population. A citizen abandons an appeal because the answer confuses a draft bill with law already in force.
These examples are hypothetical, but they reveal the same asymmetry. A provider benefits from speed and scale. A user bears the error in a particular life. ‘AI can make mistakes’ is therefore not an adequate end point. A warning without a route to verification is like a ‘slippery stairs’ sign displayed in the dark.
The user’s dignity requires clarity about the status of an answer. Freedom requires a real opportunity to decide from a trail rather than a persuasive surface. Accountability requires providers and publishers to design correction routes and to let systems say ‘I do not know’. The right to appeal means access to a person or institution capable of examining a specific error — not only a thumbs-down button that disappears into product analytics.
Media literacy has to expand accordingly. Klaudia Cymanow-Sosin defines it broadly as the development of both receptive and productive competencies across media, information, data, algorithms and AI. Key levels include assessing a communicator’s credibility and intent, selecting sources, using
verification tools and understanding AI’s part in communication (Cymanow- Sosin, 2026, pp. 19–24).
The familiar checklist — author, date, domain — is no longer enough. A real author can publish a false claim. A current web page can cite old data. An official domain can lead only to a landing page. A fabricated URL can look impeccable. Media literacy must shift its centre of gravity from the source’s label to the relationship between claim and evidence.
From claim to evidence: an eight-step protocol
This protocol works for a chatbot response, search summary, creator’s video or newsroom draft. For anything destined for publication or carrying significant stakes, keep a short evidence dossier.
1. Isolate one claim. Do not verify a paragraph as a single blob. Rewrite one assertion in its simplest form: who did what, when, where and at what scale? ‘The situation is getting worse’ is mist. ‘X rose from A to B between Y and Z’ can be checked.
2. Assign its status. Is it an observed fact, number, definition, forecast, causal claim, opinion, legal interpretation or recommendation? Different claims require different evidence. A survey can establish what respondents reported, but not necessarily how they behaved. A company statement can establish what the company says, not that its product works.
3. Demand a precise trail. Ask for title, author or institution, date, direct URL and page or section. A media brand’s name is not a citation. If the assistant cannot identify the material, label the claim unsupported. Repeatedly asking for a ‘better reference’ can produce a more elegant fiction rather than a real source.
4. Reach the primary source. For law, find the enacted text or official draft; for research, the paper and methodology; for a quotation, the recording or transcript; for data, the table and dataset documentation. Secondary sources are essential for interpretation, but should not displace an available primary record.
5. Run the entailment test. Open the item and ask whether it supports this exact claim. Does the number concern the same population, time period and measure? Does the study establish causation or correlation? Has a quotation retained its meaning? If the source says may, the answer cannot say will.
6. Check time and context. Record the publication date, date of the data, document version and conditions to which the finding applies. In news, a source that was accurate yesterday can be stale today. In science, laboratory performance is not automatically real-world impact.
7. Seek independent corroboration. For high-stakes claims, find a second source that does not merely copy the first. Two websites repeating one press release are one source wearing two jackets. Look for a different kind of evidence or a genuinely independent institution.
8. Record the verdict and uncertainty. Use clear categories: confirmed; partly confirmed; outdated; source does not support claim; no trail; conflicting evidence. Add the date, links and a sentence of reasoning. ‘No trail’ does not automatically mean ‘false’. It means we have a hypothesis, not a journalistic finding.
What a good answer interface should reveal
The full burden cannot sit with the reader. Providers and publishers integrating AI should make an answer’s epistemic structure visible. Key claims need citations beside the relevant sentence, not an undifferentiated list at the end. Links should open a specific item and, where possible, the relevant passage. Date and source type should be obvious. Facts, estimates, interpretations and uncertainty need distinct cues.
Useful features would include ‘show what this source supports’ and the ability to report a faulty claim–source relationship rather than merely rate the whole answer. Corrected answers should retain a visible version history. When no credible source is available, the system should say so and offer a route for further checking.
Newsrooms can preserve an internal dossier containing the claim, primary source, supporting passage, counter-evidence, checker, date, decision and later corrections. This record is both a quality tool and the foundation of an
appeal. When somebody challenges a story, the newsroom begins with an evidence trail, not with a colleague’s recollection that ‘the model definitely had a link’.
Where the protocol reaches its limits
A primary source can itself mislead, select or err. An official document proves what the document says, not that every statement in it is true. Apparently independent reports may rely on the same hidden dataset. The best evidence may sit behind a paywall or in another language. Source protection can prevent a newsroom from revealing the full trail.
This protocol is not a truth machine. It orders questions and exposes boundaries. In high-risk matters it cannot replace a specialist, lawyer, clinician, document researcher or reporter who understands the field. Nor does it solve inequalities of time: somebody working two jobs cannot verify like an investigative desk. That makes verification-friendly design a provider’s and publisher’s responsibility, not a luxury.
Verification intensity should also follow the stakes. A film recommendation does not require the same process as advice affecting health, law, safety, money or reputation. Media literacy includes knowing when a shortcut is proportionate and when one must descend to the primary record.
A footnote is a right to intellectual independence
A source is not jewellery pinned beneath an answer. It is a mechanism for returning some power to the reader. It says: you do not have to trust the tone, brand or machine; you can inspect where the claim came from and whether the route was honest.
After the end of the traditional footnote, we need more source literacy, not less. The decisive question is no longer ‘is there a link?’. It is ‘what exactly does this link establish?’. When the answer is ‘nothing’, we may leave the polished paragraph on the screen and call it what it is: a proposition awaiting evidence.
References
- Simon, F. M., Nielsen, R. K., & Fletcher, R. (2025). Generative AI and news report 2025: How people think about AI’s role in journalism and society. Reuters Institute for the Study of Journalism. https://doi.org/10.60625/risj-5bjv-yt69
- Metzger, M. J. (2007). Making sense of credibility on the Web: Models for evaluating online information and recommendations for future research. Journal of the American Society for Information Science and Technology, 58(13), 2078–2091. https://doi.org/10.1002/asi.20672
- Yezza, H., Fletcher, J., & Verckist, D. (2025). News integrity in AI assistants: Toolkit. BBC & European Broadcasting Union.
- Johnson, M. K., Hashtroudi, S., & Lindsay, D. S. (1993). Source monitoring. Psychological Bulletin, 114(1), 3–28. https://doi.org/10.1037/0033-2909.114.1.3
- Sperber, D., Clément, F., Heintz, C., Mascaro, O., Mercier, H., Origgi, G., & Wilson, D. (2010). Epistemic vigilance. Mind & Language, 25(4), 359–393. https://doi.org/10.1111/j.1468- 0017.2010.01394.x Stanford University. (2025, May 19). How AI is leaving non-English speakers behind. Stanford Report.