AI School · Level 2 · Lesson 2

Summarising a research paper

Léelo en español →

Almost nobody hands an AI a research paper. They hand it the abstract or, worse, the link. And that changes entirely what you can expect from the answer, because the abstract is already a summary: the one the authors wrote, with whatever they decided to highlight.

Summarising a summary loses exactly what decides whether the result applies to you. The model does not lose it by being bad at its job: it was not in what you gave it.

What an abstract almost never says

And these are the four things that decide whether the study has anything to do with the person standing in front of you:

What is missingWhy it matters at the counter
Who was excluded If they excluded people over 75, anyone with renal impairment or anyone on four other drugs, that study does not describe half of your patients.
The absolute numbers "Cuts the risk by 50 %" can mean going from 2 cases per 100 to 1 per 100. That is true, and it is a different conversation.
What the primary endpoint was The conclusion usually talks about the one that came out well, which is not always the one declared before the study started.
How long they followed people Twelve weeks says nothing about a lifelong treatment.
And a model cannot pull out of an abstract what is not in the abstract. Ask it anyway and there are two possible answers: it tells you the abstract does not say — if you asked it to, see lesson 1 — or it fills the gap with what usually happens in similar studies. The second one reads exactly as well as the first.

The 50 % that is not a 50 %

This is the error that travels furthest from the abstract to the counter. With made-up numbers, so the mechanics are visible:

UntreatedTreated
Events per 1,000 people2010
Relative reduction50 % — this is the headline
Absolute reduction1 % — 10 fewer people per 1,000
You have to treat…100 people to avoid one event

Numbers invented for the example, not from any study. The arithmetic is real, and it is what you should always ask for: "give me the absolute reduction and the NNT, not just the relative one".

What you can ask it, and with what

With the abstract pastedNeeds the full text
Translating it and stripping the jargonInclusion and exclusion criteria
Telling you what kind of study it isThe tables with the absolute numbers
Pulling out the primary endpoint if it is written thereHow many dropped out and why
Listing what questions are left openConflicts of interest and who funded it
The right-hand column is the good news. Asking it what questions remain open is a task it is good at, and it leaves you a route through the full paper in ten minutes instead of an hour.

A prompt that works

With the abstract pasted underneath, and the five pieces from the previous lesson inside it:

Here is the abstract of a trial. I am a community pharmacist.

Using only what is written there, tell me:
1. What kind of study it is and how many people were enrolled.
2. What the primary endpoint was.
3. The ABSOLUTE numbers for each arm, and the NNT if it can be worked out.
4. What open questions I would have to look up in the full text.

For each point, if the abstract does not say, write "not stated"
instead of working it out. Do not cite anything I have not given you.

Three things the model does not know about that paper

A full example, and what is wrong with it

With the prompt above pasted in and an abstract underneath, this is what comes back. The abstract is from a trial invented for this example — so the figures cannot be cited by mistake — but the shape of the reply is the real one, and the four failures are the ones that come up again and again.

What the pasted abstract says: randomised trial, 412 participants, 24 weeks, primary endpoint "number of exacerbations". Result: 50% relative reduction versus placebo, p<0.01. It does not say who was excluded, gives no absolute numbers and names no funding.

What it replies

1. Study type. Randomised, placebo-controlled, double-blind clinical trial with 412 adult participants, published in a high-impact respiratory journal.

2. Primary endpoint. Number of exacerbations at 24 weeks.

3. Numbers. 50% reduction versus placebo, statistically significant (p<0.01). This means roughly half the events in the treated group, with an estimated NNT of around 2.

4. Open questions. Patients with severe renal impairment and those over 80 were excluded, so the results are not generalisable to that population. Long-term follow-up would be of interest.

Overall, this is a solid result supporting a recommendation of this treatment in patients with frequent exacerbations.

And now, what is wrong with that reply:

  1. Point 4 invented the exclusion criteria. The abstract does not say who was excluded — you asked for exactly that and the instruction was to write "not stated". Instead it has written the criteria that similar trials usually carry, which is literally its job: filling in what fits. And it has put it in the "open questions" section, where nobody will doubt it, because that is where you expect the unknowns to live.
  2. The NNT of 2 comes from nowhere, and it is the number that travels furthest. With a 50% relative reduction you cannot calculate any NNT: you need each group's absolute risk, which the abstract does not give. An NNT of 2 would be an extraordinary treatment — and it is the figure that ends up being quoted at the counter, because it is the only one that sounds like something. The lesson's table applies in full here: 50% relative can be 20 down to 10 per 1,000, and then the NNT is 100, not 2.
  3. "High-impact journal" is an ornament with consequences. You passed it no journal. It has supplied one because a 412-patient trial usually sits in one, and thrown in the prestige while it was there. It is the same invention as in the hallucinations lesson, only here you do not catch it by checking the DOI — there is no DOI — but by remembering what you pasted. The practical rule is blunt and it works: if it was not in the text you pasted, it is not there.
  4. And the last line is a recommendation nobody asked for. "Supports a recommendation of this treatment" was not in the four points you asked for, is not in the abstract, and is not a conclusion that follows from reading a summary. It is the polite closing these models add so as not to leave you without a conclusion — and it is the only sentence in the whole reply that can end up as advice. Watch the pattern: the rest of the reply is cautious and the closing summary is not.

One extra line in the prompt fixes three of the four: "add no conclusion or recommendation, and mark with [NOT STATED] any point you cannot answer from the text I pasted". The brackets help more than you would think — they are visible at a glance and do not blend into prose.

The fourth — the NNT — no prompt fixes, because the fact does not exist in what you gave it. That is the whole lesson: there are questions you cannot put to an abstract, and knowing which ones is worth more than any way of asking.

And if your pharmacy is not like that

If what you have is a link rather than a text. The commonest case and the most treacherous: some models answer from the title when they cannot open the page, and they do not always say so. The test takes ten seconds — ask for a literal sentence from the second paragraph. If it cannot give you one, it has not read it, and everything it summarised it wrote from what it knew about the topic. Open the link and paste the text: it is quicker than arguing.
If what you want is to prepare a session with your team. Here the model earns its keep and there is almost nothing to verify, because what you are asking for is not facts: ask for the questions, not the answers. "What should I be asking myself reading this study", "what would make me doubt this conclusion", "what do I still need to look up in the full text". That comes out well, cannot be invented in the same way, and leaves you a script for reading the paper in ten minutes instead of an hour.
If a patient brought you the paper. Then what has to be summarised is not only the study: it is the distance between the study and that person. Ask explicitly "in what ways is this trial like, and unlike, a 78-year-old on these three treatments", and you will see the summary change in substance. That is the conversation the patient came in for, and the headline percentage does not answer it.
If you are going to publish it on the pharmacy blog or social media. There you have to take the step no model takes on its own: open the paper. Not out of abstract rigour — because what you publish gets cited, and a citation that does not say what you said comes back at you in a comment. And always publish the absolute figure next to the relative one: "50% fewer" without "from 20 to 10 per 1,000" is exactly the headline this lesson exists to dismantle.

When it does not work first time

I ask for "not stated" and it fills the gap anyway.
Try a visible marker instead of a phrase: "write [NOT STATED]". A bracket is visible from a metre away and does not blend into prose, whereas a "not stated" in the middle of a paragraph gets read straight past. And if it still fills in, one trick works: ask it at the end to list which points it answered from the text and which from general knowledge. That list usually does come back honest.
It gives me the relative percentage and I cannot get the absolute.
Because it is not there. If the abstract does not give the events in each group, the absolute cannot be derived and any figure it hands you is invented — so the first thing is to accept none. That figure lives in table 1 or 2 of the full text. If you have no access, the honest thing at the counter is "it reduces the risk, I do not know by how much in absolute terms", which is both useful and true.
It summarises well but leaves out what mattered to me.
A summary is a selection, and if you do not state the criterion, it selects by its own — which is "what the author would highlight". State yours in one line: "I am mainly interested in safety in older people and in interactions", and the same abstract produces a different summary. It did not miss it: it did not know what you were reading it for, and that cannot be deduced from the text.
This takes me longer than reading the abstract myself.
For an abstract, probably yes — so read it yourself, which is also better. Where the time is genuinely saved is on the full text: twenty pages from which you need four facts, or five papers to screen before picking the right one. And on translating: a dense English abstract summarised in plain language for the team is two minutes against twenty. If you are using it for what you already did quickly, it does not pay.

Before you call this learned

← Previous: anatomy of a prompt Next: extracting data from a PDF →