Skip to main content
Sections BusinessInsuranceAI Tools
MyBusinessFeed

Friday, 4 September 2026

AI Tools

The Word That Nearly Got Signed Off: What Goes Wrong When AI Writes Clinical Notes

The patient had moderate aortic stenosis. The summary said severe. One word, and everything downstream of that note would have been built on it.

risk-doctor-computer

Dr Anwar signed off forty-one summaries on a Tuesday afternoon. The scribe had been running for six weeks by then and it was good — genuinely good, the notes were tidier than his own. On the thirty-third he stopped, went back two lines, and read a word he had nearly approved.

The patient had moderate aortic stenosis. The summary said severe.

One word. Everything downstream of that note — the urgency, the referral, the conversation with the family — would have been built on it. He changed it, signed the rest more slowly, and spent the evening wondering how many Tuesdays he had already had.

A clinician working at a computer between appointments
The error rate is low. The volume is not.

That substitution is a documented case, and it is the clearest illustration of why the risks here are not the ones people expect.

The error rate is low. That is the problem.

Transcription and summarisation tools in clinical settings report error rates of roughly 1 to 3%. Older transcription models fabricated text in around 1.4% of transcriptions — in one case inventing a medication that does not exist, described as “hyperactivated antibiotics”.

Put in any other setting, 98% accuracy would be excellent. In clinical documentation it means something different, because the denominator is enormous and the consequences are not evenly distributed. A busy clinic generating a few hundred notes a week at a 2% error rate is producing several flawed records every week, indefinitely.

And the errors are not random noise. They cluster into four recognisable types:

  • Hallucination — content that was never said appears in the note
  • Critical omission — something that was said does not appear
  • Misattribution — a statement is assigned to the wrong person
  • Contextual misinterpretation — the words are right and the meaning is wrong

Dr Anwar’s near-miss was the fourth kind, and it is the hardest to catch, because nothing about the sentence looks wrong.

Omission is worse than invention

Most attention goes to fabrication, because a made-up medication is a dramatic failure and easy to explain.

Clinical notes waiting to be checked
A missing line does not look like anything at all.

But an invention is visible. You read the note, you see a drug nobody prescribed, you delete it. An omission looks like a perfectly good note that happens to be shorter.

There is a documented example worth sitting with: an AI-generated summary left out symptoms relating to premenstrual dysphoric disorder, even though medication adjustments in the record were addressing that diagnosis. The note was not wrong in what it said. It was wrong in what it left out, and the gap pointed in a direction — raising the possibility that summarisation bias becomes clinical bias, one absent paragraph at a time.

Checking a note for errors is a different task from checking it for absences, and only the first one is natural. Reviewing your own notes teaches you to read what is there.

Automation bias: the risk is you, not the model

This is the part that deserves the most attention and gets the least.

Automation bias is the well-documented tendency to over-rely on automated output and to discount information that contradicts it. It appears wherever people supervise automation, and it has been observed specifically in AI-assisted medical decision-making.

The consultation record now sits between clinician and patient
The better the tool gets, the less carefully it is checked.

The mechanism is uncomfortable because it is rational. A tool that is right 98% of the time trains you, correctly, to expect it to be right. After six weeks of good notes, reviewing becomes skimming. Six weeks is roughly how long Dr Anwar’s scribe had been running.

Which produces the central paradox of these tools: the better they perform, the less carefully they get checked, and the more damage the remaining errors do. The safety of the system does not improve smoothly with accuracy. It has a dip in the middle, exactly where “reliable enough to trust” arrives before “reliable enough to not need checking”.

There is a related finding worth knowing. When leading models were tested on physician-validated clinical vignettes that contained a single incorrect detail, hallucination rates ran between 50% and 82%. Feed a model one wrong fact and it will build confidently on it. In a documentation workflow that means an early misheard word can propagate through an entire summary — and automation bias makes the confident output harder, not easier, to question.

What actually reduces the risk

Not “review the output”. Everyone already believes they do that. Reviewing degrades over weeks, silently, in people who are certain it has not.

Check against the source, not for plausibility

The only reliable check is comparing the summary to what was actually said. Reading it to see whether it sounds right tests coherence, and coherence is the thing these tools are best at.

Read for what is missing

Deliberately, as a separate pass. Before opening the summary, note the two or three things from the consultation that must appear. Then check they are there.

Watch the modifiers

Severe, moderate, mild. Acute, chronic. Denies, reports, is unsure. Improving, stable, worsening. These carry most of the clinical weight in the fewest characters, and they are where a substitution does the most damage while looking least like an error.

Reviewing a generated summary before signing it off
Severity words carry the most clinical weight in the fewest characters.

Audit on a schedule, not on suspicion

Pull a random sample of notes every month and check them properly against the recordings. Track what you find. This is the only way to know whether your review quality is holding or quietly eroding, and it is the number that should decide whether you widen the deployment.

Assume your attention decays

Rotate who reviews. Cap how many consecutive summaries one person signs off. Treat sustained review as a fatigue problem, which it is.

What this does not mean

It does not mean the tools should not be used. The time they release is real and measured, and a clinician with more attention available is safer than one drowning in typing.

Errors that survive the consultation travel with the record
An error that survives sign-off does not stay in the clinic.

It means the risk sits somewhere other than where most deployments look for it. Teams test accuracy before rollout and then stop measuring. The thing that changes after rollout is not the model. It is the reviewing.

Dr Anwar still uses the scribe. He signs off in batches of ten now, with a break between, and he reads the severity words twice.

If you are at the stage of deploying one of these, the governance requirements are set out in our guide to adopting AI scribes in UK healthcare, which covers where the regulatory line falls and what has to be in place first. For the wider picture of what the NHS is funding, see where the NHS is actually spending on AI.

Frequently asked questions

How accurate are AI clinical documentation tools?

Reported error rates sit at roughly 1 to 3%, with older transcription models fabricating content in around 1.4% of transcriptions. In clinical documentation even a low percentage matters, because the volume is high and the consequences are unevenly distributed.

What kinds of errors do AI scribes make?

Four recognisable types: hallucination, critical omission, misattribution, and contextual misinterpretation. One documented case substituted “severe” for “moderate” in describing aortic stenosis.

What is automation bias?

The tendency to over-rely on automated output and discount contradictory information. It is documented in AI-assisted medical decision-making, and it means review quality degrades as the tool proves reliable.

Why are omissions more dangerous than fabrications?

A fabrication is visible — you read something that should not be there. An omission looks like a shorter, perfectly normal note. One documented example left out symptoms of premenstrual dysphoric disorder despite the record addressing that diagnosis.

Can a small error propagate?

Yes. When leading models were tested on clinical vignettes containing a single incorrect detail, hallucination rates ran between 50% and 82% — a model will build confidently on a wrong premise.

How should a clinician review AI-generated notes?

Check against what was said rather than for plausibility, run a separate pass for omissions, pay particular attention to severity and certainty words, audit a random sample monthly, and limit how many consecutive summaries one person signs off.

Drawn from published clinical research on AI scribes and automation bias, including work in npj Digital Medicine and related 2026 reviews. General information for context, not clinical guidance — follow your organisation’s own protocols.

Related stories

Scroll to Top