← Back to all posts
News

27 AI Scribes Write NHS Records. The Patients Do the Proofreading.

September 1, 2026 · 04:11 UTC · News
27 AI Scribes Write NHS Records. The Patients Do the Proofreading.

TL;DR

Healthwatch England, the statutory patient watchdog, told the Guardian on August 31 that AI scribes transcribing GP and hospital consultations are getting drug names and diagnoses wrong, and that in the documented cases it was the patient, not the clinician, who caught the error. One scribe dropped the word "null" from "null demyelination," turning a clear nerve-damage test into a positive one. 27 different AI scribes are in use across England's health service, the MHRA has so far left them outside the medical-device regime, and Healthwatch's own polling says nearly 90% of recent patients had no idea a scribe was in the room. If you build anything that turns speech into records, this is your product category's first real safety audit, and it did not go great.


One dropped word, one reversed diagnosis

The flagship case is brutal in its smallness. A woman, herself an NHS health professional, got a summary saying she had demyelination, the nerve damage behind conditions like multiple sclerosis. The actual test result read "null demyelination." The scribe deleted the one word carrying the entire meaning, and she described the aftermath as "a very traumatising experience."

where the meaning fell out of the pipeline test result:null demyelination AI scribe summarydrops one word patient record:demyelination
One deleted negation reversed a diagnosis, and the patient found it, not the doctor.

It was not a one-off. Healthwatch's casework includes a scribe that confused a prescribed drug with a different one of a similar name, and an AI-generated summary letter that omitted a consultant's instruction to get a migraine medication repeat prescription from the GP. London GP Dr Shier Ziser Dawood, who calls the tools "a double-edged sword," recounted a scribe charting instructions to continue Prozac when the antidepressant was never prescribed or discussed. In every headline case, the last line of defense was the one person in the room without medical training.

Healthwatch's warning is exactly the failure mode you would predict: these inaccuracies "may persist in their records if the patient doesn't catch them."

27 scribes, zero device certificates

The scale is the story. GPs and hospital doctors across England are using 27 different AI scribe products, also called ambient voice technology: record the consultation, transcribe it, then have a generative model produce structured notes, letters, and referrals that flow into the patient record and out to other providers. The government's 10-year health plan expects them to "liberate staff from their current burden of bureaucracy and administration." So far they have also liberated at least one test result from the word "null."

Here is the regulatory gap. The MHRA has not classified AI scribes as medical devices, which keeps them outside the regime that would test safety and effectiveness before deployment. NHS England's ambient scribing guidance, updated to version 3 on July 29, puts the burden on buyers and suppliers instead: organizations must complete a DCB0160 clinical safety case, suppliers must provide a DCB0129 one, and every AI-generated output is supposed to be reviewed and approved by a human before it enters clinical use. Whether a product counts as a medical device hangs on the manufacturer's self-declared "intended purpose," which is a definition you can draft your way around.

The overconfidence loop

The mandated human review exists. The problem is what the humans believe about the machine. A survey of 1,003 UK GPs cited in the Guardian's reporting found more than half believed their AI-written records were more accurate than their own. Once the reviewer trusts the tool more than themselves, review-and-approve decays into approve.

The error pattern is not random, either. Dr Charlotte Blease, an AI-in-healthcare researcher at Uppsala University, notes the risk climbs with multi-person consultations, complex medical histories, and patients who do not speak English as a first language. And drug names are the worst possible vocabulary for a speech-plus-language-model pipeline: think of a spell-checker that silently corrects unfamiliar words into familiar ones, except the unfamiliar words are the medications, and the model's statistical pull is always toward the more common, similar-sounding token. The rarest words in the transcript are simultaneously the most clinically important and the most likely to get rounded off.

What patients think of the robot in the room

Healthwatch ran the numbers on consent months before the error casework landed. Its YouGov poll of 4,039 UK adults, fielded April 16-27, found comfort collapses as the consultation gets more sensitive.

comfortable with an AI scribe listening, by topic (% of UK adults) routine check48% sexual health29% mental health28% domestic abuse23%
Fewer than half are comfortable even for a routine check; for sensitive topics it drops to roughly a quarter. (YouGov for Healthwatch England, April 2026)

The consent picture is worse than the comfort picture. Of 2,795 respondents who had an appointment in the previous 12 months, nearly 90% did not know whether AI scribing was used. 81% want to be informed and asked before a scribe is switched on, and 69% say they would be more comfortable if clinicians explicitly committed to verifying the AI's output. Overall sentiment is a dead heat, 38% supportive against 37% opposed, but the intensity is lopsided: 21% strongly opposed versus 11% strongly supportive.

patients vs the scribe (YouGov for Healthwatch England, April 2026) unaware it ran~90% want consent81% want checks69%
Nine in ten recent patients could not say whether an AI wrote part of their record.

Rachel Power, chief executive of the Patients Association, put the stakes plainly: "Trust and confidence in this technology depend on good communication and genuine partnership with patients."

Why this lands on builders

Ambient scribing is one of the few AI verticals with real, paying, at-scale deployment, and this is what its first public safety accounting looks like. If you are building anything speech-to-record shaped, medical or not, the lessons transfer directly.

  • Eval the negations. "Null," "no," "denies," "ruled out": dropped negation is the highest-damage, lowest-visibility failure in summarization. If your eval suite does not specifically score negation preservation and rare-entity fidelity, you are shipping the demyelination bug.
  • Review UX is the product. A wall of plausible prose invites rubber-stamping. Surfacing low-confidence spans, drug names, and negations for explicit confirmation is what makes "human review" mean something.
  • Beware users who trust you too much. When over half of professional reviewers think the model writes better notes than they do, your safety story cannot be "a human checks it."

Key Takeaways

  • Healthwatch England documented AI scribes reversing a diagnosis by dropping "null," confusing similar drug names, omitting a prescription instruction, and charting a never-discussed Prozac prescription; patients caught the errors.
  • 27 different AI scribes are in use across England's health service, feeding GP and hospital records at national scale.
  • The MHRA has so far left AI scribes outside the medical-device regime, so there is no England-wide pre-deployment test of safety or effectiveness; NHS England guidance leans on DCB safety cases and mandatory human review instead.
  • A survey of 1,003 UK GPs found over half trust the AI's notes more than their own, which quietly hollows out that mandatory human review.
  • Healthwatch's YouGov polling: ~90% of recent patients unaware scribes were used, 81% want consent first, and comfort falls from 48% for routine checks to 23% for domestic abuse consultations.
  • For builders: score negation preservation and rare-entity accuracy explicitly, and design review flows that fight rubber-stamping instead of inviting it.

Sources: The Guardian (syndicated on AOL), Healthwatch England / YouGov survey, NHS England ambient scribing guidance, The Next Web

AIhealthcareNHSAI scribesspeech to texthallucinationsregulationtrends
CONSOLE
$