The reporter has removed the informant’s name. She replaces the town with ‘a regional centre’ and the employer with ‘a public institution’. The n she pastes the interview into a public chatbot and asks for five key themes. The transcript still says that the speaker is the only deputy director appointed in a particular month, atte nded a rare closed meeting and objected to a decision made on a specific Friday. The name is gone. The person remains.

This is a hypothetical scene, but it exposes a common misunderstanding. De- identification is often treated as a red felt-tip pen: cross out direct identifiers and the text becomes anonymous. In small worlds — a local council, university department, hospital ward, parish, specialist newsroom — identity lives in combinations. Job title plus timing plus event plus relationship may function like a fingerprint.

The ethical question therefore arrives before the first generated word. It is not merely whether the chatbot will produce a correct summary. It is whether the

reporter was entitled to introduce this system, its provider and its data practices into a conversation that the source understood as confidential.

Off the record is a social promise, not a file setting

Journalistic confidentiality begins in a relationship. A source discloses information because a reporter has made an explicit or implicit promise about identity, use and control. The promise may have legal dimensions, but it also has moral force: one person accepts risk because another undertakes to manage it carefully.

Digital tools complicate this relationship because communication and storage involve intermediaries. Julie Posetti’s UNESCO study on source protection argues that the digital age expands both the range of actors able to obtain journalistic information and the traces that can expose sources. Protection therefore requires attention to the full information chain, not only the published byline (Posetti, 2017).

A chatbot can become another participant in that chain. This does not mean a human employee is necessarily reading every prompt, nor that every system uses inputs in the same way. Provider policies, contracts, technical architecture, retention and account settings vary and change. The point is simpler: submitting material is a disclosure to an external information system under conditions the journalist must understand well enough to justify.

Consent to speak to a reporter is not blanket consent to invisible onward processing. A source who agrees to recording may not agree to a consumer AI service receiving the transcript. ‘But I removed the name’ is not an answer if the remaining context can identify the person.

Context determines whether information travels properly

Helen Nissenbaum’s theory of contextual integrity treats privacy not as absolute secrecy but as appropriate information flow. Information moves within social contexts according to norms concerning the actors, the type of

information and the conditions of transmission. A flow can violate privacy even when the data are not conventionally secret, because they travel to an inappropriate recipient or for an unexpected purpose (Nissenbaum, 2010).

This is an excellent lens for newsroom AI. Consider the actors: source, journalist, editor, provider, possible subcontractors and authorised users. Consider the attributes: identity, allegation, health detail, meeting date, voice, unpublished evidence. Consider the transmission principles: off the record, embargo, limited editorial access, no training, temporary processing, legal duty.

The prompt changes the flow. A statement given for assessment by a named reporter may be processed for summarisation by a service governed by a separate contract. Even if the output never reveals the source, the input may already have crossed a contextual boundary. Privacy harm is not limited to public leakage; it can consist in an unauthorised change of audience and purpose.

The mosaic problem: anonymous pieces, identifiable person

Direct identifiers are obvious: name, email address, telephone number, exact home address. Quasi-identifiers are more treacherous. A job, age range, location, date, distinctive biography or unique participation may not identify a person alone. Combined, they can sharply narrow the field.

The NASK guide warns that inputs to AI systems can include confidential information, personal data, professional secrets and material covered by journalistic privilege. It specifically notes risks concerning informants and investigative documents and advises checking provider documentation and the scope of processing before use (Adamczyk et al., 2025, pp. 30–33). Its bluntest practical observation is that ‘a journalist essentially loses actual control over it as soon as it is sent’ (author’s translation; Adamczyk et al., 2025, p. 32).

That loss of control is not necessarily total in a technical or legal sense. Approved enterprise systems may provide contractual restrictions, defined

retention and administrative controls. The sentence is best read as a warning against casual reversibility: after submission, a reporter cannot behave as though the material never left the device merely because the chat window has been closed.

Sachita Nishal and Nicholas Diakopoulos similarly note that newsrooms want human-supervised systems sensitive to source confidentiality and privacy. Their review warns that language models can present risks of memorising sensitive training material and that using confidential documents in prompts can undermine otherwise useful tasks such as summarisation (Nishal & Diakopoulos, 2023, pp. 1–3). The relevant conclusion is not ‘never use AI’. It is ‘task suitability begins with the data, not the convenience of the output’.

The source pays for a shortcut they never chose

Imagine an employee in a small cultural institution who reveals procurement irregularities. Only four people attended the key meeting. The reporter removes all names before asking a chatbot to organise the transcript. The text still contains the speaker’s gender, approximate seniority, a reference to returning from parental leave and a disagreement recorded on a specific date.

If the material is exposed, retained inappropriately or accessed beyond the expected circle, the reporter may suffer embarrassment or disciplinary consequences. The source may lose employment, face a lawsuit, experience retaliation or become unsafe. Human responsibility belongs to the newsroom: dignity forbids treating the source as raw material, while autonomy preserves their right to choose who receives their words.

The IREX and Development Gateway safety guidance warns journalists against entering source identities, unpublished-investigation details or information that could put somebody at risk into free-tier generative AI tools. It recommends masking, pseudonymisation, understanding retention and choosing tools according to risk (O’Shea & Orrell, 2025, pp. 9, 15, 18). Crucially, the same guidance recognises that freelancers often lack organisational support while carrying multiple roles. Safety cannot be reduced to an individual’s memory of a settings menu.

A traffic-light rule for newsroom prompts

A simple red–amber–green classification helps create friction before submission. It must be adapted to local law, organisational policy and the tool actually approved by the newsroom.

Red material includes source identities and contact details; raw off-the-record or background interviews; unpublished investigations; precise meeting locations and future plans; unique combinations likely to identify a person; documents covered by legal, professional or journalistic secrecy; intimate, health or safeguarding information; credentials and access secrets; and material whose disclosure could expose somebody to retaliation or physical harm.

Red does not mean ‘find a cleverer prompt’. It means change the workflow. Use human analysis, an approved secure environment designed for the risk, or seek specialist advice. If the task cannot be performed without disclosing the protected material, convenience does not supply authority.

Amber material includes edited transcripts with residual contextual identifiers; sensitive but non-secret internal documents; allegations not yet verified; embargoed material; datasets that may become identifying when linked; and content whose harm depends heavily on retention, access and purpose.

For amber data, define the task narrowly and minimise the input. Remove unnecessary passages, replace attributes consistently, separate the question from the source document where possible, and use only a system whose contract, retention, access, processing location and training terms have been reviewed by the appropriate newsroom functions. Obtain additional consent where the new processing falls outside what the source reasonably understood.

An amber decision requires a named owner and a record. ‘Everybody uses this tool’ is not approval.

Green material includes already published information, public documents without residual confidential annotations, synthetic examples, house-style exercises and non-sensitive administrative text. Even here, copyright, accuracy and provider terms still matter. Green describes confidentiality risk, not overall editorial quality.

The classification attaches to both content and context. A public quotation can become sensitive when paired with an unpublished location. An innocuous paragraph can identify a source when the audience is a group of five insiders. Colours should be reassessed when information is combined.

The pre-prompt dossier

Before an amber use, complete a short dossier. This turns a private hunch into an auditable decision.

1. Purpose. What precise task is AI performing, and why is it preferable to a non-AI alternative? ‘Help with the interview’ is too broad. ‘Identify repeated public-policy themes in these selected, de-identified passages’ is reviewable.

2. People and promises. Whose information is present? What was agreed: on the record, background, off the record, embargo, limited sharing? Could the new processing surprise the source?

3. Identifiers and combinations. List direct and quasi-identifiers. Ask a colleague familiar with the field, but not the source’s identity, whether the edited material points to one person. Do not include the sensitive text in an unapproved tool to test whether it is identifiable.

4. Minimum input. Which passages are strictly necessary? Can the task be done with a synthetic structure, public extracts or locally prepared notes instead of the raw transcript? Data minimisation is most effective before upload.

5. Tool conditions. Record the product and account type, approved purpose, retention, access roles, provider use of inputs, processing

location, deletion controls, subcontracting where relevant and date of review. Policies change; the date matters.

6. Risk and alternatives. Describe plausible harm, likelihood, the people who would bear it and safer options. High-impact uncertainty should move the case to red or to specialist review.

7. Authority. Name the editor, security or data-protection lead, and legal adviser where required. Record who can stop the use. A freelancer commissioning work for a newsroom needs a route to the commissioning organisation’s support.

8. Output handling. Treat the result as unverified material. Decide where it will be stored, who will check it against approved source material and what will be disclosed to readers about AI assistance.

9. Incident route. If protected material is submitted accidentally, who is contacted immediately, how is the source informed, what evidence is preserved and who assesses further exposure? A prompt error should not be hidden out of fear of blame; delay can compound harm.

Designing an AI workflow that deserves a source’s trust

Newsrooms should publish a clear matrix of permitted tools and data classes, illustrated with journalistic examples rather than generic corporate labels. Access should follow role and need; retention should have a reason and an end date; and logs should support accountability without storing unnecessary sensitive content.

Procurement is part of source protection. The newsroom must ask whether an agreement supports journalistic secrecy, incident response, deletion, audit and restrictions on secondary use. Training must include the right to refuse an unsuitable AI task. Where processing is material, sources deserve a plain explanation of what will be processed, by which approved system, for what purpose and with what alternatives. Consent is meaningful only if ‘no’ remains possible.

Limits and difficult cases

Perfect anonymisation is often impossible in long interviews and small communities; removing too much context can also destroy meaning. That is a reason to choose another method or a more controlled environment, not to upload the original. Provider documentation changes, so journalists need institutional legal and security support plus periodic review. Minimisation can also conflict with the need to preserve material for fact-checking or legal defence; the answer is governed retention, not indiscriminate deletion or indefinite accumulation. Local systems may reduce some risks while creating others: ‘on our server’ is not a complete threat model.

Before pressing Enter, count the people in the room

A reporter would not casually invite an unknown consultant into an off-the- record meeting merely because the consultant promised a fast summary. The screen makes that social fact easy to forget. A chatbot feels like an instrument, yet using it can alter who receives information, under what rules and for whose purpose.

The safest prompt is not necessarily the best-written one. It is the one the journalist is entitled to send. Start with the person, the promise and the possible harm; then classify the material, minimise it and document the tool. If the transcript still points to the informant, the absence of a name changes very little.

Off the record is not a magic phrase that follows data everywhere. It is a duty carried by people and institutions. Before asking AI to join the conversation, the newsroom must be able to explain to the source why this third participant deserves to be in the room at all.

References

  1. Posetti, J. (2017). Protecting journalism sources in the digital age. UNESCO.
  2. Nissenbaum, H. (2010). Privacy in context: Technology, policy, and the integrity of social life. Stanford University Press.
  3. Nishal, S., & Diakopoulos, N. (2023). Envisioning the applications and implications of generative AI for news media. In CHI ’23 Workshop on Generative AI and HCI (pp. 1–8). https://arxiv.org/abs/2402.18835
  4. O’Shea, M., & Orrell, T. (2025). Guidance note for newsrooms & journalists: Safety considerations for the use of generative artificial intelligence (GenAI). IREX & Development Gateway.