AI School · Level 1 · Lesson 7

ChatGPT, Gemini or Claude: what actually differs

Léelo en español →

This is the question people ask most and the one that matters least. Short answer first: for what a pharmacy is going to do, any of the big three will do, and switching model almost never fixes a problem.

The whole thing in one sentence. There are real differences between ChatGPT, Gemini and Claude, but none of them is "which one knows more pharmacy". They are whether it trains on what you type, whether it shows you where each claim came from, and what you can upload. That is what decides, and it looks nothing like what gets argued about in the videos.

The three questions that actually separate them

1. Does it train on what I type?

First and decisive, because it is the only one with legal consequences (lesson 1). All three let you turn it off, and in all three it lives somewhere different and moves every few months — which is why there is no screenshot here: it would be out of date before you read it.

The rule that does not expire: find it in the settings the day you open the account, and if you cannot find it in two minutes, assume it does train and act accordingly.

2. Does it show me where it got that?

A model that searches the web and cites real links saves you half the verification, because a link opens and checks out in ten seconds. One that answers from memory leaves you all the work from lesson 3.

⚠️ But watch this: citing does not mean the citation says what it says. The link can exist, be genuine, and not support the sentence. It is the most common failure of search-enabled models, and the least examined precisely because there is a link sitting there.

3. What can I upload, and what happens afterwards?

Uploading a PDF changes a great deal of what you can do — the lesson on reading a medical report lives on it — but it also changes the risk: a file carries metadata, and a scanned report carries the name in the letterhead even if you only look at the body. The question is not just whether it accepts files: it is how long it keeps them and who can see them.

What people compareWhat actually decides
Which one is "smarter"Whether it trains on what you type
Which one is fasterWhether it cites sources you can open
Which one writes more nicelyWhat it accepts and what it does with it after
Benchmark scoresWhether the free tier does what you need
The mistake of comparing by benchmark. They measure exams, not counters. A model can score better on medical multiple-choice questions and still invent a citation just as readily as the next one — they are different skills, and the one that hurts you is the second.

What does NOT change between them

And this is the point of the lesson, because it is what makes the question matter so little:

The honest recommendation. Pick one, switch off training, and learn it properly. Knowing how to ask one well pays far more than changing model every month: most of the "this one is worse" you hear is really "I have not learned to ask this one yet".

A full example, and what is wrong with it

The "which one is better" argument ends the moment you put the same question to all three and look at the three replies together. What comes out is not a winner: it is the reason the question was wrong to begin with.

The query is an ordinary counter one, and one of those that has no one-word answer:

"Can an 850 mg metformin tablet be split in half?"

What it replies

Model A. Yes, immediate-release metformin can be split without problems if the tablet is scored. Split it just before taking it and take the other half at the next dose.

Model B. Splitting metformin tablets is not advisable, as it may alter the release of the active ingredient and cause more gastrointestinal effects. It is better to ask the doctor for a presentation with the right strength.

Model C. It depends on the specific presentation. Immediate-release metformin with a functional score line can usually be divided; prolonged-release forms must not be split, because the release mechanism is lost. Check the product information for that specific medicine: the administration section states it.

And now, what is wrong with that reply:

  1. Two of them agree and that means nothing. The temptation is to break the tie by majority, or to keep the one that agrees with what you already thought. Neither works: all three read the same internet, so if two agree the likeliest explanation is that they are repeating the same source text, not that they arrived independently at the same conclusion. Three models are not three independent opinions; they are far closer to three copies of the same book.
  2. The most useful one is the least satisfying. C does not answer: it tells you what it depends on and where to look. It is the right answer and the least liked, because it leaves you with work. This is where models get picked badly — the decisive-sounding one gets rewarded, and sounding decisive is not a fact about what it knows, it is a fact about how its tone is tuned. Somebody who tries them for a week and says "A is much better" is usually saying "A leaves me less to do".
  3. None of the three can know the answer. And that is the point: the question has no answer without knowing which box is on the counter. The score line may be functional or purely decorative, and only the administration section of that product's information says which. A is right in one case and wrong in the other; B is wrong by being cautious, which is still wrong — it removes a valid option from somebody who may not be able to swallow a whole tablet.
  4. And the difference you see is style, not knowledge. Ask A to reply "always saying what it depends on and where to check it" and you will get something very like C. Ask C to be brief and decisive and it will look like A. Almost everything credited to the model is really the instruction it is carrying, and you are the one who writes the instruction. Hence the lesson's recommendation: learning to ask one of them well pays more than switching model every month.

The test that does help you choose is not this one. It is three things, and none of them looks like comparing answers: open the settings and check whether it trains on what you write (two minutes), ask for something with a source and check the link exists and says what it says (five), and upload a file and find out how long it is kept (another five).

Twelve minutes, once, and you decide on what actually differs. Comparing who writes more prettily is hours of video and decides nothing.

And if your pharmacy is not like that

If you are only going to pay for one account. This is the normal case in a pharmacy, and it simplifies the decision a lot: take the one your team will actually use — usually the one already on their phones — and spend the effort configuring it properly. One paid account set up well, with training off and history under control, protects you more than three free accounts scattered about and never looked at. The risk was never in picking the worse model: it is in having three places where the wrong thing gets pasted.
If you are going to work with long documents. Reports, leaflets, the collective agreement, a hundred-page PDF. Here there is a real technical difference and it is called how much text fits in at once: a model that only gets halfway does not tell you, it summarises what it read and says nothing about the rest. The test takes a minute — feed the document in and ask what the last section says. If it does not know, it has not read it all, and everything it summarised was incomplete without saying so.
If you care about citations you can actually open. Then do not choose a "model": choose search mode, which is an option inside nearly all of them and changes far more than the brand does. With search, verification drops to ten seconds per link; without it, you do the whole job from lesson 3. Keep that lesson's warning though: a link that opens is not a link that holds the sentence up, and in search modes that is the commonest failure.
If your pharmacy software already has one built in. Then the comparison is not about quality: it is about paperwork and where the data ends up. A built-in model can be the worse writer and still be the right choice, because the vendor signs a data processing agreement and an open chat does not. And it can be the other way round: convenient, no agreement, and sitting right next to the patient record it invites you to paste. Ask about the agreement before you ask about the model.

When it does not work first time

I switched model and it still gives me bad answers.
That was to be expected, and it is the lesson's conclusion. Look at what you are asking for: if it is facts — doses, prices, deadlines — no model will fix it, because those facts live in the regulator's database and the official gazette and neither is inside the model. And if it is writing and it comes out badly, it is almost always missing context: an example of what you want, who it is for and how long it should be. Try giving it those three things in the model you already had before paying for another.
The free tier runs out halfway through a task.
That is the most honest reason there is to pay, and it has nothing to do with quality. Before you do, check two things: that you are not burning the quota on enormous conversations you could close and reopen — every message drags everything before it along — and that what you need is not file upload rather than pasted text, which usually sits in the paid tier. If you do end up paying, pay for the limit and the files, not for "it is smarter".
We use one at work and another at home, and they answer differently.
Normal, and not a problem as long as you remember one thing: the rules you have learned here do not depend on the model. Patient data, in neither. Unsourced figures, in neither. Anything leaving your screen gets checked, in both. What is worth avoiding is mixing places for the same job: pick one for work and leave the other for home. Half the pharmacy on each is the fastest way for nobody to know where what got pasted.
Somebody recommended one "specialised in medicine". Should I get it?
Before paying, put the lesson's three questions to it and a fourth: where does what it answers come from? If it is connected to sources you can open — product information, guidelines, referenced databases — that is a genuine difference and you notice it on first use. If what is there is a general model with an instruction on top telling it to sound like a doctor, you are paying for a tone. The way to find out is to ask for one specific thing and see whether it links it.

Before moving on

← Revisit: professional responsibility
Next: when NOT to use AI →
← Back to the AI School