AI School · Level 1 · Lesson 3

Hallucinations and how to check an answer

Léelo en español →

"Hallucination" is an unfortunate word, because it suggests something strange and visible. What actually happens looks nothing like a delusion: it looks like an ordinary answer. In fact it is indistinguishable from a good one, and that is the entire difficulty.

The whole thing in one sentence. The model does not have two modes — one for knowing and one for inventing. It has one: writing what fits. When what fits happens to match reality we call it a correct answer; when it does not, we call it a hallucination. Inside, it is the same process, so it cannot warn you.

Where it hurts most: sources

A bibliographic reference is the perfect place for this to happen. It has a highly predictable shape — authors, year, journal, volume, pages, DOI — and filling in a predictable shape is precisely what the model is best at. The result is a flawless citation:

Hargreaves R, et al. Efficacy of vitamin D in the prevention of respiratory tract infections. Br J Gen Pract. 2021;71(704):212-219. doi:10.3399/bjgp21X714089

The authors sound like authors. The journal exists. The year is plausible. The DOI has the right format. And the paper may not exist.

Why this is worse in a pharmacy than almost anywhere else. That citation does not stay on the screen. It ends up on a counter poster, in a pharmacy social post, in advice somebody acts on, or in a conversation with a patient who arrives quoting what they read. By then nobody remembers where it came from.

Checkable is not true

This distinction is the whole lesson, and it is worth saying slowly. An answer can arrive:

How it arrivesWhat you can do
With nothing to go and look at. "Several recent studies show…", "according to the scientific community…" Nothing. It is not that it is false: it is that there is no way to know. Ask for the specific reference, and if none comes, bin it.
With something concrete: a DOI, a PMID, a licence number, a regulation, a link Go and look. Thirty seconds. And note: having a DOI does not mean it exists — it means you can now find out whether it does.

The detector below does exactly that: it separates what has an address from what sounds authoritative and leads nowhere. It does not check whether anything is true — you do that, at the link — and it cannot.

Does this source lead anywhere?

Paste what the AI gave you as backing. It runs in your browser: nothing is sent anywhere.

Where each thing gets checked

If it gave you…Check it at
A DOI (10.xxxx/…)doi.org/ + the DOI. If it does not resolve, it does not exist.
A PMIDPubMed. Search the number, not the title.
A dose, an indication, an interactionThe SmPC on the emc, or the BNF. Always, no exceptions.
A regulation or a clauselegislation.gov.uk — and check the section number is the one it quoted.
A linkOpen it. A link that 404s is the easiest hallucination there is to catch.
A headline with no referenceNothing. That is not a source.
Search the identifier, not the title. Search the title on Google and you will find something similar almost every time — there are millions of papers — and you will walk away feeling the citation was fine. An identifier does not do approximate: either it resolves or it does not.

The thirty-second protocol

You do not need to verify everything. You need to verify what can cause harm, and be able to tell which that is:

  1. Is this leaving my screen? If it is a draft you are going to rewrite anyway, there is nothing to verify. If it goes on a poster, to a patient or onto social media, there is.
  2. Is there a figure, a dose or a source in it? Those are the three things that get invented. The rest of the text is rarely the problem.
  3. Do I have something to go and look at? If not, do not use it. Full stop.
  4. Look. Thirty seconds on the emc, on doi.org or on legislation.gov.uk.
A trick that works better than it sounds. Ask the same question twice, in separate conversations. What it genuinely knows tends to come out the same; what it is inventing changes authors, year or journal. It is not proof — sometimes it repeats the same invention — but it is free and it catches a lot.
And what does NOT work: asking it "are you sure?". It will almost always apologise and give you a different answer, right or wrong — that is what fits after that question. Asking it to grade itself is asking for another prediction, not for a check.

A full example, and what is wrong with it

The detector above separates what has an address from what does not. This example is the next step up, and the one that actually costs you: a source that does resolve and still does not support what is being said with it. Here the light goes green and the problem is entirely intact.

What gets asked: "Does vitamin D help prevent respiratory infections? Give me a reference with a DOI and an institutional source."

What it replies

Yes. The Martineau et al. meta-analysis published in the BMJ showed that vitamin D supplementation reduces the risk of acute respiratory infection by around 30%, with a consistent benefit in the general population. Reference: doi:10.1136/bmj.i6583.

You can also consult the World Health Organization's institutional information on micronutrients: https://www.who.int/.

In practice, 1,000–2,000 IU daily would be reasonable for most adults during the winter months.

And now, what is wrong with that reply:

  1. The DOI exists, it resolves, and that is why nobody looks further. Paste it into the detector above and it comes back green; paste it into doi.org and a real paper opens. Ninety per cent of people stop there, and quite reasonably: the lesson you came from says a DOI that resolves is what you were after. It was — so that you could start checking. Whether the reference exists and whether the reference says what you were told are two different questions, and the tool can only answer the first.
  2. The paper is real and the headline is not its headline. Open it and read the abstract: the overall benefit is small, and the point of the work is that it concentrates in people who started out deficient and in daily or weekly dosing, not high-dose boluses. The "30% in the general population" is not in there. This is not an invented citation: it is a real citation with a conclusion laid on top of it that makes a better headline. And it is harder to catch than an invention, because everything checkable in the sentence — author, journal, DOI — survives the check.
  3. The "institutional source" is a homepage. who.int opens, of course it opens. But it is not a source: it is the door of a building. A source takes you to the specific document that holds the claim up, with its title and its date. Asking for "an institutional source" and accepting a domain is the same mistake as accepting "according to the scientific community", only with the appearance of a link — and because the link works, it passes both filters from the previous lesson. That it opens is not that it says.
  4. And the regimen at the end is not what you asked for. You asked whether it works and where that comes from. It answered that and threw in a specific dose for "most adults", which is the one part of the reply that can end up as counter advice. Neither cited source backs that regimen. This is the pattern worth recognising: the actionable part arrives last, unreferenced, leaning on the credibility the two paragraphs above have just built.

The full check is three questions, and the tool above already answers the first: does it exist? · does it say that? · does it say that about these people? The second is answered by reading the abstract, which takes two minutes. The third is what separates a fact from a piece of advice, and it is where almost everything posted on social media falls over.

Note this does not mean reading a full paper every time. It means reading it when the sentence is going to leave your screen — and if you are not going to read it, then do not cite the paper. Saying "there is some evidence, mainly in deficient people" without a reference is honest; saying "30%" with a DOI underneath that does not say so is not.

And if your pharmacy is not like that

If the answer is for you and goes no further. Orienting yourself, understanding a term, working out where to start looking. Do not verify here: it would be wasted time. What matters is being clear that what you now know, you know provisionally, and that the moment it is about to become advice, a poster or an order, verification kicks in. The failure is not skipping the check: it is forgetting that you skipped it.
If it is going to a specific patient, standing in front of you. Change the source, not the level of checking: for doses, indications, interactions and contraindications the official product information governs, and a meta-analysis does not replace it even when it disagrees. The question to put to the AI here is "what should I look at?", and the answer gets looked up in the regulator's database. It is quicker than arguing with the model, too.
If it is going on a poster, on social media or into a talk. The worst place for a misused citation, because it travels and does not come back. Rule of thumb: if you have not opened the source, do not cite it. And if you have opened it and it says less than you would like, write what it says — a poster promising 30% and one saying "it may help, particularly if you are deficient" read almost the same, and only one of them you can defend when somebody asks where you got it.
If the patient arrives with what they have read. More and more, what they bring came out of a chat. Exactly the same protocol applies, and it is easier: ask them for the source and check it in front of them. One of two things usually happens — either they do not have one, and the conversation ends there without anybody having to be right; or they have one and it says something else, and showing them the abstract on screen is worth more than any explanation. It is the best argument a pharmacist has today, and it is free.

When it does not work first time

The DOI does not resolve, but I find the paper by searching the title.
With a real paper in front of you, most likely the DOI was mistyped or composed from another one. Take the DOI from the paper's own page, not the one the chat gave you, and carry on with that. But check one thing before calling it good: that the paper you found says what was being cited. Very often what turns up is "a similar one" — there are millions — and what you have done is find a different backing for a claim that still had none.
I do not have access to the full paper.
You almost never need it. The abstract is public on PubMed and covers all three questions: what they measured, in whom, and how big it was. If the figure you were given is not in the abstract, the honest move is not to use it — it may sit in a subgroup inside the full text, and a subgroup figure presented as a headline result is exactly the error in the example above. And for counter work, the official product database is free and it is the one that governs.
It says "according to recent studies" and I cannot get a reference out of it.
Ask once, specifically: "give me author, year, journal and DOI or PMID". If what comes back does not resolve, do not push and do not rephrase the question: a third attempt only produces a third, better-built invention. Treat it as what it is: no source, therefore no claim. If the topic genuinely matters, a two-keyword PubMed search takes the same time as a third round of pushing.
Checking everything takes me longer than it saves.
Then you are using it for something it is not. If your work with the AI has turned into fact-checking, it is because you are asking it for facts — and the facts live in the regulator's database and the official gazette, which are faster anyway. Give it back what it is good at: take the fact from the good source yourself and ask it to explain, summarise or write it up. There the time saved is real and the time spent verifying is zero.

Before calling Level 1 done

← Back to the AI School