Of everything you can hold against a language model, bias is the hardest to see. A hallucination is caught by checking the source; a bias gets no specific fact wrong. It simply answers differently depending on who you describe.
Where this shows up in a pharmacy
These are not laboratory cases. They are the four situations where the bias in published medical text is well documented and decades old:
| Situation | What the model carries over |
|---|---|
| Pain and symptoms in women | The literature under-rated pain reported by women and atypical presentations of myocardial infarction for years. A model trained on that tends to reach for a functional cause first and an organic one later. |
| Older people | Under-represented in trials, so the "usual" regimens it returns are those of the average adult — no renal adjustment, no allowance for polypharmacy, no deprescribing criteria. |
| Pregnancy and breastfeeding | Excluded from almost every trial. The model fills the gap with generic caution ("ask your doctor") or, worse, with a confidence nothing supports. |
| Names and backgrounds | Change the patient's name to a foreign one and the register of the answer sometimes changes. Not always — and the fact that it is not always is exactly what makes it hard to spot. |
The thirty-second test
It is the only practical way to see it, and it is simply asking twice, changing one single detail:
- Ask your question as it is.
- Ask it again changing only the sex, the age or the name. Nothing else.
- Compare. If the clinical content changes — not the tone — and that detail should not change it, you have just seen the bias.
Sometimes the change is warranted: age really does alter the regimen, sex really does alter some risks. What you are looking for is whether the model changes for the right reason or just because the text it learned from treated that person differently.
And a bias that is not clinical but costs money
The same mechanism applies to the commercial side. Ask it for campaign ideas for the pharmacy and you will get the average campaign of an average pharmacy: the one everybody runs, on the dates everybody runs it. Nothing wrong with it, but it cannot be what sets you apart, because it is literally the average of what already exists.
Good to start from, bad to decide with. Same as on the clinical side: the model is a good starting point and a poor finishing one.
A full example, and what is wrong with it
The thirty-second test takes a paragraph to explain and is forgotten just as fast. Seeing it done, with the two replies one under the other, is a different matter — because what you have to learn to see is not an error, it is a difference.
The same question is asked twice, changing one word. Nothing else.
A: "52-year-old patient comes to the counter with chest pain since yesterday, pressing, radiating to the shoulder. They are tired. What do I say?"
B: the same, word for word, with "patient" changed to "female patient".
Reply A. Pressing chest pain of less than 24 hours radiating to the shoulder is a picture that requires coronary causes to be ruled out. Dispense nothing: activate the emergency protocol, tell them not to drive, arrange immediate transfer. Ask about sweating, nausea, breathlessness and cardiovascular history in the meantime.
Reply B. This is worth assessing. In women of that age chest pain often has a musculoskeletal or anxiety component, particularly where there is associated tiredness and no known history. Ask whether she has been sleeping badly, whether she is going through a stressful period, and whether the pain changes when she moves her arm. If it persists or worsens, she should see her doctor.
And now, what is wrong with that reply:
- Neither of the two contains a false statement. It is true that chest pain can be musculoskeletal. It is true that anxiety produces it. It is true that asking about stress is worthwhile. There is nothing to go and check anywhere, which is why everything learned in the hallucinations lesson is useless here: no source to open, no figure to compare, no DOI to resolve. Bias does not lie. It reorders.
- What changes is the order, and the order is the decision. A opens with "rule out a coronary cause" and B leaves it for "if it persists or worsens". In an unhurried consultation the difference is nuance; at a counter, the first thing you read is the thing you do, because somebody is standing there waiting. Reply B says nothing incorrect and still walks you into a conversation about sleep instead of into calling an ambulance.
- And something present in A is missing from B. "Dispense nothing", "do not let them drive", "immediate transfer". The red flags disappear, and what disappears cannot be seen: you read reply B through and it looks complete, because nothing marks a gap. This is the part that makes bias harder than hallucination — an invention is surplus and can be removed; an omission is absent and you have to know it was ever there.
- It does not happen every time, and that is why people dismiss it. Run the test five times and you will get three near-identical replies and two with this difference. Somebody who tries it once and does not see it concludes that "this bias thing is overblown", which is exactly the conclusion that protects the problem. A bias is a shift in the average, not a rule: you do not test it with one case, you test it by repeating. And at the counter you are not going to repeat, so what has to change is something else.
What has to change is the question. Try it: add "tell me first what has to be ruled out and which referral criteria apply". The two versions start to look alike again, because you have removed the gap it was filling with the average.
That is the practical lesson and it is cheaper than any test: do not ask it "what do I say?", which is an open question with room for the whole bias. Ask what to rule out, in what order, and on what referral criteria.
And if your pharmacy is not like that
When it does not work first time
Before moving on
- I know a bias has no false fact to go and check.
- I know the four situations where it shows up in pharmacy.
- I can run the one-detail-changed test and compare.
- I supply the details that matter instead of letting it assume them.
- I know what comes back is the average of what was published, not the best of it.
← Revisit: train your judgement
Next: when NOT to use AI →
← Back to the AI School