AI School · Level 1 · Lesson 8

Images and voice: the multimodal risks

Léelo en español →

The previous two lessons were about text: what you write, and what happens to what you write. A photo or a voice note is something else entirely, and the reason is simple: a piece of text only contains what you decided to type, but a photo contains everything that was in front of the camera at the moment you took it, whether you meant it to or not.

The whole thing in one sentence. An image or an audio clip identifies by what is around what you are asking about, not just by what you are asking. You can write a perfectly anonymous question and still upload a photo that is not anonymous at all, because nobody asked you to crop the background before taking the shot.

Four ways an image identifies without you noticing

What slips inWhy it goes unnoticed
A name on a label or prescription, at the edge of the frame. You frame the shot thinking about the pack or the active ingredient you want to ask about, not about what is left at the edge of the frame — and a dispensing label usually carries a full name.
The file's own metadata (location, date, exact time). The photo carries that information attached even though it is not visible on screen; many tools read it anyway when processing the file.
A recognisable background from your own counter. A screen with a patient's record open behind it, a pigeonhole with a visible name, a calendar with an appointment — none of that was the point of the photo, and all of it is still in the frame.
A recognisable voice in an audio note. If you record a patient's voice describing a symptom so as not to lose any detail, that voice is itself identifying data, even if you never say a name anywhere in the recording.
Reflections in glass or a metal surface at the counter. A shop window, a display case or even a metal scale can bounce back the reflection of a screen or of somebody standing behind you, outside what looked like the main frame — and it is the hardest of the four to spot on sight, because it does not show up by looking at the centre of the photo but at its reflected edges.
Why re-reading the text lesson does not fix this. The "strip name, ID number, date of birth" rule works on what you write, because you decide every word. In an image you do not decide every pixel: you decide the framing, and everything inside it travels along with it, whether you looked at it before uploading or not. That is why this needs a separate habit — the same rule applied from memory is not enough. It is a different habit because the physical gesture is different too: writing invites you to re-read before sending, while pointing and firing a camera is a gesture done and forgotten in the same second, without that built-in review step that text carries simply by being slower.

A full example, and what is wrong with it

A pharmacist wants to ask about a possible interaction for a drug newly launched on the market that is not yet in the site's tool's dataset, and takes a photo of the pack so as not to mistype the name.

What happens

The photo comes out exactly right for what she wanted: the brand name is clearly readable on the box. The framing, done quickly over the counter between one customer and the next, also includes, in the background, the computer screen with the pharmacy management software open — with a full name visible in the active window — and, in the corner, a printed electronic prescription waiting to be dispensed, with the CIP number legible.

And now, what is wrong with taking that photo as fine:

  1. The question was perfectly legitimate; the photo was not. Asking about the brand name of a drug carries no privacy problem at all. The problem is not what was asked, it is what else ended up in the frame without anybody looking at it before pressing send.
  2. Nobody reviews the background of a photo the way they review a sentence. Written text gets re-read before sending, out of habit. A photo gets taken and uploaded in almost the same gesture, without that review step — and it is exactly that step that needs adding as a habit.
  3. The prescription's CIP number is as strong an identifier as an ID card number. It is exactly the kind of data the privacy lesson forbids typing by hand, and here it got in without a single letter being typed, purely by being inside the frame.
  4. The fix is not "stop taking photos", it is "look at the frame before uploading it". Three seconds looking at what is around what you want to show — not just the centre of the frame — would have been enough to crop the photo or retake it without the screen and the prescription in the background, with no need to give up the convenience of photographing the pack instead of typing it out by hand.

A three-step routine, so you never have to think it through from scratch

Memorising "be careful with photos" is not much use in the middle of a busy counter, because it is too vague to apply in a hurry. What actually works is a short routine, always the same, so it becomes a reflex instead of a decision you have to make from zero every time.

  1. Before you shoot: isolate the object. Move closer or further away until the frame only contains what you want to ask about, with no screens, no other people's paperwork, nothing behind it besides a neutral surface.
  2. After you shoot and before you upload: scan the whole frame, not just the centre. A two-second look at the four corners of the photo is enough to catch most problematic backgrounds.
  3. After you use it: delete it from the phone's camera roll. Do not let the photo pile up in the phone's gallery once the query is resolved — it is the easiest trace to forget about and the longest-lived one, because nobody scrolls back through their camera roll looking for photos from three months ago.

With these three steps turned into habit, the risk in this lesson stops depending on remembering it in the moment: it becomes part of the very gesture of taking the photo, the same way looking both ways before crossing a street is already a reflex.

And if your pharmacy is not like that

If you work at a narrow counter with little room to frame things well. This is exactly where something is most likely to slip into the background, because there is no distance to separate the object of the photo from the rest of the environment. The cheap fix is a fixed neutral background — a folder, a cleared patch of surface — to always rest whatever you are photographing on, instead of improvising the framing each time over the counter as it happens to be.
If you use voice notes so as not to lose detail of what a patient tells you. Swap the habit for summarising in your own words on the spot, with no name, instead of recording the original voice. Some nuance is lost, that is true, but you gain not having to manage an audio file with a recognisable voice afterwards that you have to remember to delete.
If you often photograph prescriptions or reports on paper to ask something. Here the risk is not in the background, it is in the document itself: cover the personal-data section with your hand or a piece of paper before taking the shot, instead of photographing the whole document and trusting yourself to crop it later in a hurry.
If you have security cameras recording the counter. This has no direct bearing on what you upload to an AI, but it is worth keeping in mind when choosing where to take a photo to ask about something: an angle chosen to avoid appearing in your own security camera is usually also an angle that avoids the problematic background of the other photo.
If you use a personal phone to take these photos. Remember that beyond the risk of what shows in it, that photo also stays saved in your phone's camera roll — with its own location metadata — until somebody deletes it by hand. Delete it from the camera roll right after using it, not only from the conversation with the AI.
If you dictate notes to your phone's voice assistant to transcribe later. The risk here is not only the voice of whoever is dictating: if while dictating you say a patient's name out loud because it feels natural when speaking, that name becomes text just as if you had typed it, with the added catch that speaking makes it easier for a detail to slip out that you would never have written by hand. Review the transcript before using it, the same way you would review your own written text.

When it does not work first time

I do not have time to review the framing every time, we are busy at the counter.
Then automate the step instead of relying on remembering: a piece of card or a folder always within reach to rest whatever you photograph on solves the background problem at a stroke, without having to think about it each time or lose any real time.
Is location metadata as serious as the rest?
On its own, it does not usually identify a patient — it identifies your pharmacy, which is already public. The problem appears when it is combined with another piece of data from the same photo: your pharmacy's location plus a label with a name in the same image are two pieces of data that together say far more than either one alone.
I accidentally recorded a voice note with the patient's name said out loud.
Delete it before uploading it anywhere, and if you have already uploaded it, delete it from the conversation just as you would a badly written piece of text — there is no difference in how to treat a voice-based piece of data once the mistake is spotted, except that here the mistake is easier to make without noticing.
Can I just ask the AI to ignore the background of a photo?
You can ask, but do not trust that it truly complies: a model that analyses images processes the whole frame before deciding what is relevant to your question, so the information in the background has already been "seen" — and potentially retained — even if it never gets mentioned in the answer.
The photo resolution is so low that no name is readable — does it not matter then?
It helps, but it is no guarantee: a model built to read text in images often extracts more from a blurry photo than you would expect, especially if the text is large or well contrasted. Low resolution reduces the risk, it does not remove it — the habit of checking the frame is still needed.
What if I need to keep the photo as a record, not just to ask a one-off question?
Then it is a different case from this lesson: a photo kept as part of a record needs the same treatment as any document you archive, with its own retention and access rules — the "crop the background before asking an AI" rule is built for a one-off, disposable query, not for keeping anything long term.

Before moving on

← Revisit: automation bias
Next: prompt injection, the basics →
← Back to the AI School