AI School · Level 1 · Lesson 4

Train your judgement: catch the error

Léelo en español →

So far you have seen what you can ask a model and how to check what it returns. This lesson is different: here you will not write a single prompt.

The twist. In the three cases below the AI has already answered. Your job is to say what is wrong. Because at the counter that is the real job: not getting an answer out of the machine — anyone can do that — but knowing when you cannot trust the one it gave you.

Why this, and not "100 prompts for pharmacists"

A good prompt improves the average answer. It does not protect you from the bad one, because the bad one looks exactly like the good one — you saw that in lesson 3. The only thing that protects you is your judgement, and judgement is trained the way everything is: by getting it wrong somewhere it does not matter.

The three answers below are written the way a real model writes them: well phrased, confident, with a correct detail or two for decoration. None is an obvious absurdity. And all three would do harm.

Everything here is synthetic. All three cases are made up by us and correspond to no real patient and no real employee. That is the rule from lesson 1, and it does not get broken even to set an exercise.

The three cases

Read the AI’s answer and write what you think is wrong with it. In your own words and in any order: you are not matching a sentence, you are spotting the problems. It is marked against a rubric written by a pharmacist, not against whatever another model happens to think.

How it is scored, and why the percentage is not enough. Each finding is weighted by severity — one major is worth three minor ones — and if you miss a serious one you do not pass, whatever percentage you score. Catching six details and missing the major interaction is not a narrow pass: it is precisely the failure this lesson exists to prevent.

What the three cases have in common

A full example, and what is wrong with it

Before the three cases, one worked through out loud. It is none of them and it spoils nothing: it is there for the thing that stops most people the first time, which is not spotting the errors — it is not knowing what you are supposed to write.

The question is one that gets asked for real, by a pharmacy technician:

"I do on-call shifts from 21:00 to 07:00. How much night supplement should I be paid a month if I do four of them?"

What it replies

With shifts from 21:00 to 07:00 you are doing 10 night hours each time, that is 40 hours a month across four shifts.

Article 36 of the Workers' Statute establishes a 25% supplement over base salary for night work. In the pharmacy agreement that works out at around €2.50 an hour for the technician grade, so roughly €100 a month.

Bear in mind that this amount is pro-rated if you work part time, and that being a salary item it is also included in the calculation of the extra payments.

And now, what is wrong with that reply:

  1. Ten night hours are not ten. They are eight. (major) The night period runs from 22:00 to 06:00. A 21:00-to-07:00 shift lasts ten hours, but only eight fall inside that window: the first and the last do not count as night hours. The model got no sum wrong — 10 × 4 = 40 is correct — it counted the wrong thing, and that is indistinguishable from a correct calculation if all you check is the arithmetic. It is the failure that repeats in all three cases of the exercise.
  2. The Statute's 25% is not what gets paid here. (major) The Statute defines the night window and says its specific pay is set by collective bargaining. It fixes no 25%. And the pharmacy agreement does not pay a percentage: it pays an amount per hour, which for the national technician grade is €1.82. The citation sounds impeccable — article, statute, percentage — and it is exactly the pattern lesson 3 warns about: the well-shaped reference is the weakest part, not the most reliable.
  3. In part-time work it is NOT pro-rated. (moderate) The night supplement is paid per night hour actually worked: somebody doing four night hours is paid for four, whether their week is 20 hours or 40. Pro-rating it would mean being paid less for the same night work, which is the opposite of what the supplement exists for. It is an expensive, silent error — anybody who does not know the agreement swallows it, and the person least likely to know it is usually the one on part time.
  4. And it does not go into the extra payments. (moderate) The last sentence sounds like a technical aside and it is an error with a price on it: the night supplement is excluded from the extra-payment calculation. Notice its shape — "bear in mind that…" — because it is the same shape as in all three cases of the exercise: the costliest error is not in the body of the reply, it is in the closing aside, where you are already skimming.

Now the part that actually helps: this is how you write the answer. Unordered, in your own words, not drafted. A passing answer looks like this:

"Night hours run 22:00 to 06:00, so it is 8 per shift, not 10. The Statute's 25% does not apply: the agreement pays a per-hour amount, €1.82 for a technician. The supplement is not pro-rated for part-timers. And it does not go into the extra payments."

Forty seconds. There is no need to explain why or cite the article: you have to point. And notice what decides the pass — both findings marked major are in there. An answer twice as long that misses either of those two does not pass, whatever percentage it scores, because spotting four details and missing the one that changes the payslip is precisely the failure this exists to prevent.

And if your pharmacy is not like that

If you are a technician or assistant rather than a pharmacist. The two clinical cases will cost you more, and that is fine: this does not score your degree, it scores whether you spot that a fact is missing. And counter experience pays off heavily there — "it never asked the weight", "it never asked what else they take", "the pharmacy cannot decide that" are complete findings without knowing the dose by heart. The employment case, on the other hand, you will do better than almost anybody, because it is your own payslip.
If you are the owner and plan to give this to your team. Do them yourself first, unhurried, and do not show your result. The value is not in the score: it is in the conversation afterwards, and that conversation dies if somebody has already seen the marking. What does work is fifteen minutes together talking one case through out loud once everybody has done it — that is where you find out what each person looks at first, which is information you get no other way.
If you finished your degree years ago. You will go faster than you expect, because what is being trained here is not memory: it is recognising the smell of an answer that answers too much. You have that from the counter. What may be rusty are the specific figures — the paracetamol ceiling, the day sick pay starts — and that does not matter: knowing them is not what is scored, not accepting a figure you have not checked is.
If you are still studying and have never worked a counter. The opposite will happen to you: you will see the content errors and miss the ones about framing — "it answers instead of saying a fact is missing", "it decides something that is not the pharmacy's to decide" — which is where half the marks are. Read them with that question in front of you and everything changes: not "is this correct?", but "is this the person behind the counter's to decide?".
If two of you do it at the same time. It is the best way to do it and it costs the same, with one rule: each of you writes your answer before either speaks. The moment one says out loud "I think the dose bit is wrong", the other can no longer know whether they would have seen it — and that is precisely what is being measured. Separately first, then compare the two lists, and the interesting part is not who found more: it is what one saw and the other did not, because that repeats at the counter every day.

When it does not work first time

I read the reply and nothing comes to mind as wrong.
Stop looking for errors and look for where each statement comes from. Walk the reply sentence by sentence asking "where is this coming from?": from a rule it cites (does the rule say that?), from a fact it was given (was it given?), or from nowhere. Nearly all the findings live in the third category. And look specifically at the last sentence, which is where these models slip in the aside nobody reads.
I wrote a lot and scored badly.
Normal: the rubric scores findings, not words. Three paragraphs explaining why the model is wrong about one thing are worth exactly the same as one line saying so — and while you were writing those three paragraphs you were not looking at the rest. Try it the other way: one line per finding, in any order, and when nothing else comes to mind, reread the whole reply. That is usually where the missing one turns up.
I spotted something the rubric does not list. Is my answer wrong?
No. The rubric lists what you have to spot to pass, not everything that can be spotted; extras do not subtract. If what you saw is a genuine error, good — that is exactly the judgement this trains, and it usually comes from knowing something about the case that whoever wrote it did not assume. What is worth checking is whether that finding took the place of a major one in your head: seeing everything except the thing that decides is the pattern the marking is looking for.
I passed all three and I still trust it too much.
That is to be expected, and it is not a failing of yours: here you know something is wrong — we told you in the title. At the counter nobody warns you, and that is the whole difference. Which is why the practice that pays is not repeating these cases, but applying the same reading to the next reply a real AI gives you: where each statement comes from, what it failed to ask, what it decides that is not its to decide. Three goes at that with a real answer are worth more than ten cases here.
I missed a major finding and it stung.
It is the one part of this school designed to sting, and not gratuitously: here nothing happens and at the counter it does. What to do with it is not to repeat the case until you pass — you will end up knowing it by heart, which is not the same as having seen it — but to look at why it got past you: did you skim the last sentence? did you trust it because the arithmetic held up? did you assume a fact nobody had given you? That answer is worth more than the case.

Before moving on

← Revisit: hallucinations and how to check
← Back to the AI School