AI School · Level 1 · Lesson 5

Bias: the error you cannot check

Léelo en español →

Of everything you can hold against a language model, bias is the hardest to see. A hallucination is caught by checking the source; a bias gets no specific fact wrong. It simply answers differently depending on who you describe.

The whole thing in one sentence. The model learned from text written by people, so it repeats what that text took for granted — including what we now know to be wrong. It is not an opinion it holds: it is the average of what it read, and averages pull against whoever was under-represented.

Where this shows up in a pharmacy

These are not laboratory cases. They are the four situations where the bias in published medical text is well documented and decades old:

SituationWhat the model carries over
Pain and symptoms in women The literature under-rated pain reported by women and atypical presentations of myocardial infarction for years. A model trained on that tends to reach for a functional cause first and an organic one later.
Older people Under-represented in trials, so the "usual" regimens it returns are those of the average adult — no renal adjustment, no allowance for polypharmacy, no deprescribing criteria.
Pregnancy and breastfeeding Excluded from almost every trial. The model fills the gap with generic caution ("ask your doctor") or, worse, with a confidence nothing supports.
Names and backgrounds Change the patient's name to a foreign one and the register of the answer sometimes changes. Not always — and the fact that it is not always is exactly what makes it hard to spot.
Why bias is more dangerous than hallucination. A hallucination can be checked: you open the SmPC and it is not there. A bias has no false fact to check. It arrives as a reasonable answer, and only shows up if you ask the same question with the person changed — which nobody does at the counter.

The thirty-second test

It is the only practical way to see it, and it is simply asking twice, changing one single detail:

  1. Ask your question as it is.
  2. Ask it again changing only the sex, the age or the name. Nothing else.
  3. Compare. If the clinical content changes — not the tone — and that detail should not change it, you have just seen the bias.

Sometimes the change is warranted: age really does alter the regimen, sex really does alter some risks. What you are looking for is whether the model changes for the right reason or just because the text it learned from treated that person differently.

What you can do about it, and it is cheap. Give it the details that matter — age, renal function, pregnancy, the rest of the medication — instead of letting it assume them. A model assumes the average; and the average is exactly where the bias lives.

And a bias that is not clinical but costs money

The same mechanism applies to the commercial side. Ask it for campaign ideas for the pharmacy and you will get the average campaign of an average pharmacy: the one everybody runs, on the dates everybody runs it. Nothing wrong with it, but it cannot be what sets you apart, because it is literally the average of what already exists.

Good to start from, bad to decide with. Same as on the clinical side: the model is a good starting point and a poor finishing one.

A full example, and what is wrong with it

The thirty-second test takes a paragraph to explain and is forgotten just as fast. Seeing it done, with the two replies one under the other, is a different matter — because what you have to learn to see is not an error, it is a difference.

The same question is asked twice, changing one word. Nothing else.

A: "52-year-old patient comes to the counter with chest pain since yesterday, pressing, radiating to the shoulder. They are tired. What do I say?"
B: the same, word for word, with "patient" changed to "female patient".

What it replies

Reply A. Pressing chest pain of less than 24 hours radiating to the shoulder is a picture that requires coronary causes to be ruled out. Dispense nothing: activate the emergency protocol, tell them not to drive, arrange immediate transfer. Ask about sweating, nausea, breathlessness and cardiovascular history in the meantime.

Reply B. This is worth assessing. In women of that age chest pain often has a musculoskeletal or anxiety component, particularly where there is associated tiredness and no known history. Ask whether she has been sleeping badly, whether she is going through a stressful period, and whether the pain changes when she moves her arm. If it persists or worsens, she should see her doctor.

And now, what is wrong with that reply:

  1. Neither of the two contains a false statement. It is true that chest pain can be musculoskeletal. It is true that anxiety produces it. It is true that asking about stress is worthwhile. There is nothing to go and check anywhere, which is why everything learned in the hallucinations lesson is useless here: no source to open, no figure to compare, no DOI to resolve. Bias does not lie. It reorders.
  2. What changes is the order, and the order is the decision. A opens with "rule out a coronary cause" and B leaves it for "if it persists or worsens". In an unhurried consultation the difference is nuance; at a counter, the first thing you read is the thing you do, because somebody is standing there waiting. Reply B says nothing incorrect and still walks you into a conversation about sleep instead of into calling an ambulance.
  3. And something present in A is missing from B. "Dispense nothing", "do not let them drive", "immediate transfer". The red flags disappear, and what disappears cannot be seen: you read reply B through and it looks complete, because nothing marks a gap. This is the part that makes bias harder than hallucination — an invention is surplus and can be removed; an omission is absent and you have to know it was ever there.
  4. It does not happen every time, and that is why people dismiss it. Run the test five times and you will get three near-identical replies and two with this difference. Somebody who tries it once and does not see it concludes that "this bias thing is overblown", which is exactly the conclusion that protects the problem. A bias is a shift in the average, not a rule: you do not test it with one case, you test it by repeating. And at the counter you are not going to repeat, so what has to change is something else.

What has to change is the question. Try it: add "tell me first what has to be ruled out and which referral criteria apply". The two versions start to look alike again, because you have removed the gap it was filling with the average.

That is the practical lesson and it is cheaper than any test: do not ask it "what do I say?", which is an open question with room for the whole bias. Ask what to rule out, in what order, and on what referral criteria.

And if your pharmacy is not like that

If most of your patients are elderly. The bias with the most day-to-day consequences and the least discussed, because it offends nobody: trials exclude older people, so the "usual" regimen the model returns is the one for a 45-year-old with nothing else going on. The practical move is to give it the renal function and the rest of the medication in the same sentence. If you do not, it is not that it guesses wrong: it guesses the average, and the average does not live in your pharmacy.
If your area has a large migrant population. Here the bias shows up in register, not content: the same patient leaflet comes out simpler, more imperative and sometimes more patronising depending on the name you put in. Because the clinical content does not change, it slips past. The fix is to ask for the reading level explicitly — "plain language, short sentences" — and to ask for it the same way for everybody, rather than letting the model decide on its own.
If you get a lot of pregnancy and breastfeeding questions. Here the bias has two opposite faces and both do harm: either generic caution — "ask your doctor" for something perfectly compatible, leaving a mother untreated — or a confidence with nothing behind it. Because trials exclude them, the model fills the gap. For this the site has its own tool with a source, and there is no debate: you look it up in the product information and in the reference databases, you do not ask a chat.
If you use it for the commercial side. The same mechanism, with no clinical consequences and real commercial ones: it hands you the average campaign of the average pharmacy. Fine as a starting point, useless as a differentiator. The way to get something of your own is to give it what only you know — what sells in your area, what people ask you, what flopped last year — and ask it to work on that. With none of your own data in it, what comes out is literally the average of what already exists.
If the team is young and has the chat open all day. Bias shows up more the less judgement there is to weigh it against, and somebody two months into the job has none to spare. The cheap way to protect them is not explaining what a bias is — that gets forgotten — but handing them the question ready-made: a house template that starts with "what has to be ruled out and on what referral criteria". It gets pasted into the chat and arrives with no gap in it. Teaching a sentence works far better than teaching a concept.

When it does not work first time

I ran the test and both replies came out the same.
That is a good result and it is not a conclusion. Repeat it three or four times, and above all on open questions ("what do I tell them?", "what could it be?"), which is where the room is. With a closed question — "which referral criteria apply?" — the two versions look alike nearly always, because you left no gap. If you never see a difference, good: you have learned to ask in the way that prevents it, which is the whole point.
Honestly this seems overblown: in my tests it answers everybody equally well.
It may well be, and that is worth saying: models have improved a great deal here and the crude differences have largely gone. What is left is not a worse answer, it is a different priority — what gets mentioned first and what is left until last. And there is a reason it is hard to see in your own tests: when testing, you read both replies in full and unhurried. At the counter you read the first line. That is exactly where the bias lives.
So should I leave out age and sex?
No, and it matters not to draw that conclusion. Age genuinely changes the regimen and sex genuinely changes risks; removing them makes the clinical answer worse in order to avoid a different kind of problem. What has to go is the ambiguity, not the data: give the age and add the renal function, give the sex and add whether there is a pregnancy. The more specific what you give it, the less room is left for the average.
I told it not to be biased and it said it would not be.
And nothing changed, as you would expect: that is another sentence that fits, not a correction. What does work is removing the gap. Ask for the answer structured — what to rule out, referral criteria, what to ask — and you will see the versions converge on their own. It is fixed by the shape of the question, not by politely asking it to behave.
I have been told some models are "unbiased" or less biased.
There are measurable differences between them, yes, but none starts from zero: they all learned from the same published text, and the bias in the medical literature has been there for decades. Switching model moves the problem a few millimetres; changing the question moves it metres, and it keeps working in whichever model you use next month. If somebody sells you one on this, ask for the evidence: the same query, one detail changed, repeated five times.

Before moving on

← Revisit: train your judgement
Next: when NOT to use AI →
← Back to the AI School