AI School · Level 3 · Clinical

Reading a medical report with AI

Léelo en español →

"Upload the report and ask for a summary" is what everyone does, and it is exactly where this breaks: a summary does not tell you what it left out, and in a medical report what is missing is half the information.

The whole thing in one sentence. This is not a prompt, it is an eight-step method. Each step produces something you can review on its own, which is how you can tell which step got it wrong — precisely what a one-shot summary takes away from you.
Before step 1, without exception: the report cannot identify anyone. Name, initials with a date of birth, record number, NHS number, address, the letterhead — all of it goes before anything touches a chat. And if the document cannot be de-identified without losing what makes it understandable (and sometimes that happens), then this method does not apply to that document. There is no quick version of this: it is lesson 1, and it does not lapse because the method is good.

The method

  1. Extract

    Give me back what it says, unsummarised and uninterpreted. Verbatim.

    What you ask"Transcribe the data in this report as it appears: diagnoses, treatments with dose and regimen, labs with units and ranges, allergies and dates. Do not summarise, do not interpret, do not fill in what is missing."

    Where it fails: it fills gaps. If the report says "metformin" with no dose, it will happily write "metformin 850 mg" because that is the usual one. Ask explicitly for anything not stated to be flagged.

  2. Structure

    The same content, in fixed sections. Now it can be read at a glance.

    What you ask"Organise it into: reason · history · current treatment · investigations · clinical impression · plan. If a section has no content, write 'not stated'."

    Where it fails: the empty sections. Unless you force "not stated", it drops them — and a section that is not there reads like a section that was not needed.

  3. Flag uncertainties

    What is illegible, ambiguous, abbreviated or incomplete. This step is what makes it a method, and it is the one nobody asks for.

    What you ask"List everything you could not read with confidence, every abbreviation with more than one reading, and every incomplete field."

    Where it fails: it says there are none. A model would rather return an empty list than admit something is unclear; press a second time and they usually appear.

  4. Separate fact from interpretation

    Two columns: what the report says and what is inferred. Almost every error in a review lives on that border.

    What you ask"Split into two lists: (A) what the document states literally and (B) what would be an inference of yours. Do not mix them."

    Where it fails: it puts inferences in column A. "Hypertensive patient" when the report only says "enalapril" is a very reasonable deduction — and it is still a deduction.

  5. Check against sources

    Here you leave the chat. Every dose, every interaction, every indication is checked on the emc or the BNF — not in the model.

    What you ask"For each medicine, tell me what would need checking in the SmPC: indication, renal adjustment, interactions with the rest of the list." And then you check it.

    Where it fails: it answers what you were going to go and check. It is convenient and it is the opposite of the step. What you want is the list of what to look at, not the answer.

  6. Identify open questions

    What this report is missing before you can act. This is what you take to the prescriber or the patient.

    What you ask"What information is missing here to review the medication safely? Rank them by how much each would change the decision."

    Where it fails: it returns twenty questions. Ask for the five that would most change the decision — a list of twenty never gets used.

  7. Write the summary

    Now, yes. And the summary is built on the six previous steps, not on the document again.

    What you ask"Using the above, write a 10-line professional summary. Mark visibly what is uncertain and what is still to be checked."

    Where it fails: the summary swallows the uncertainties from step 3. Unless you force it to carry them through, they come out clean — and a summary with no uncertainties is exactly the one you cannot use.

  8. Human review

    The last one, and the one that is not delegated. Not "a quick read": a check with three questions.

    • Is there anything in the summary that is not in the report?
    • Is there anything in the report that matters and is not in the summary?
    • Have I checked at the source what came out of step 5?

    Where it fails: the second one. What was added shows up as you read; what was omitted does not, by definition, and you have to go looking for it on purpose.

Why eight steps and not one. With a single "summarise this", the error spreads across the whole result and there is no way to locate it. With eight, each step produces something reviewable — and when something goes wrong you know which step broke, which is the difference between a method and a habit.

How long it really takes

The first time, about twenty minutes. With practice, six to eight — and steps 1, 2 and 4 can be chained into a single request once you know what each returns.

What never gets shorter are 5 and 8: checking at the source and the final review. If the method ends up at six steps, it has gone back to "ask it for a summary" with more ceremony.

A full example, and what is wrong with it

The method has eight steps and the temptation is to think that writing them down is enough. This example shows step 3 — spotting uncertainties — on a synthetic discharge summary, because it is the step that makes the whole method different and the one that most often comes back with an empty list.

The fragment, already de-identified: "Discharged after decompensation. Treatment adjusted: metformin, enalapril 20, aspirin 100, omeprazole. Furosemide 1 tab added. Blood tests pending. Previously prescribed NSAID withdrawn. GP review in 7-10 d."

What it replies

Step 1 — extract. Discharge medication: metformin 850 mg every 12 h · enalapril 20 mg every 24 h · aspirin 100 mg every 24 h · omeprazole 20 mg every 24 h · furosemide 40 mg every 24 h.

Step 3 — uncertainties. No relevant uncertainties detected. The report is complete and internally consistent.

Step 4 — facts and inferences. (A) What the report says: diabetic and hypertensive patient, on antiplatelet therapy, with decompensated heart failure. (B) Inferences: might benefit from a renal function review.

Step 7 — summary. Patient with diabetes, hypertension and heart failure, discharged on five drugs with GP review in 7-10 days. Treatment appropriate and uneventful.

And now, what is wrong with that reply:

  1. Step 1 filled in four doses that were not there. The report says "metformin" flat, "omeprazole" flat and "furosemide 1 tab". The extraction shows 850 mg every 12 h, 20 mg every 24 h and 40 mg every 24 h — which are the usual doses, not the report's. And that is what makes step 1 dangerous: you asked for the literal text, it returned something better than literal, and a complete list reads like a correct list. A gap is visible; a filled gap is not.
  2. "No uncertainties detected" in a report full of them. There are at least four: three drugs with no dose, "1 tab" with no strength, "previously prescribed NSAID" without saying which — it matters, because the bleeding risk alongside aspirin is not the same for all of them — and "blood tests pending" without saying which or when. The model prefers an empty list to admitting something is unclear. The lesson says to press a second time, and that is when they appear: ask again with "be exhaustive, even if they seem minor" and all four come out.
  3. Column A is full of deductions. "Diabetic and hypertensive patient" is not what the report says: it is deduced from the metformin and the enalapril, and they are reasonable deductions — and they are still deductions. An ACE inhibitor may be there for heart failure rather than hypertension, which is exactly what the rest of the report suggests. This is the border where almost every review error lives: not in what gets invented, but in what gets assumed and filed under facts.
  4. And the summary comes out clean, which is the worst part. "Treatment appropriate and uneventful." Not one flag, not one "pending confirmation", no trace of the four invented doses. Step 7 swallows step 3's uncertainties — which were empty anyway — and what is left is a perfectly presentable ten-line document that claims more than is known. If somebody reads it three months from now, nothing tells them that half the doses came from habit rather than from the paper.

What fixes all four is not a better prompt: it is not chaining the steps. With all eight in one request, the model optimises the final output — a presentable summary — and along the way it fills, keeps quiet and smooths over. Separated, each step produces something you can look at on its own, and the uncertainties one can be repeated until it returns something.

And step 8's check is done in this order, which is not the intuitive one: what is missing first, what was added second. What was added jumps out as you read; what is missing has to be hunted with the report in front of you, field by field. The summary above reads perfectly precisely because what it lacks is not there.

And if your pharmacy is not like that

If the report is a photo of the paper. Then the whole problem from the PDF lesson arrives before the method even starts: a completed digit looks identical to a read one, and here that digit is a dose. The practical rule is that step 1 gets checked against the paper, field by field, before going on — because the seven remaining steps work on whatever came out of it. An extraction error is not caught by any later step: it propagates in a tidy hand.
If it is an emergency department note rather than a discharge summary. The risk moves: emergency notes are short, full of abbreviations, and the plan is one line at the end. Step 3 becomes the most important of the eight, because half the document admits two readings and the model will pick one without telling you. Ask explicitly for every abbreviation listed with all its possible readings — a thirty-second step that changes the whole result.
If the patient brings it in and wants it explained. There the method has a ninth step that is not written down: translating into plain language, which is one of the things a model does best. But the order matters — you translate at the end, over the summary you have already reviewed, never over the original document. Translate first and you are simplifying something you have not checked yet, and what simplifying loses is exactly what you were about to verify (Level 1 lesson).
If they are serial reports on the same person. This is where the method pays most and almost nobody uses it this way: what matters is not each report but what changed between them — which drug appeared, which vanished without explanation, which dose moved. Ask for a comparison table of the treatments and you will see gaps that reading the reports separately does not show. With one caveat: a drug that disappears from a report is not necessarily a drug that was stopped, and that is a column B inference.

When it does not work first time

Eight steps feels like a lot for a one-page report.
They are. For a short report, steps 1, 2 and 4 chain into a single request without losing anything — the lesson says so itself. What never goes together are 3 and 7: ask for the uncertainties and the summary at once and the summary eats them, which is literally the fourth failure in the example above. And 5 and 8 are not chat steps: they are yours, and they are the ones that do not shorten.
I press on step 3 and it still says there are no uncertainties.
Change the shape of the question: instead of "are there uncertainties?" — which admits a no — ask for a numbered list of at least five things a pharmacist would want to confirm before acting. Forcing it to produce items works far better than asking it to assess, because assessing leads it to reassure. And if the five that come out are irrelevant, that is information too: it means the document is more complete than it looked.
I cannot de-identify the report without losing what I need.
Then this method does not apply to that document, and there is no quick version. It is the situation the lesson warns exists and it genuinely happens: reports on rare diseases, from one very specific centre, or with a narrative that identifies on its own. What you can do is work from your notes on the report — the six facts you care about, written by you — rather than from the document. It is more work and it is the only legitimate way.
The summary looks good and I do not know whether something is missing.
That is exactly the feeling an incomplete summary produces, so do not trust it. The cheap check runs the other way: take the original report and tick off each fact that appears in the summary. Whatever is left unticked on the report is what was omitted, and it shows in two minutes. Looking for what is missing by reading the summary never works — a complete text and an incomplete one read exactly alike.

Before calling this learned

← Revisit: when NOT to use AI
← Back to the AI School