← All posts / Policy

NHS Watchdog Warns Doctors' AI Scribes Get Drug Names and Diagnoses Wrong

Healthwatch England says NHS patients are catching dangerous AI transcription errors — wrong drugs, wrong diagnoses — that doctors miss, as 27 scribe tools spread with no England-wide oversight.

NHS Watchdog Warns Doctors' AI Scribes Get Drug Names and Diagnoses Wrong

AI scribes — the ambient voice tools that listen to a GP consultation and draft the clinical note — can put patients at risk by getting the names of drugs and illnesses wrong. That is the blunt warning from Healthwatch England, the statutory NHS patient champion, in an exclusive reported by The Guardian on 31 August 2026. And in a pattern that should make anyone deploying these systems uncomfortable, several of the most serious errors were caught not by the doctors using the tools, but by the patients themselves.

What the watchdog found

The cases Healthwatch highlighted are exactly the failure modes critics of ambient AI have feared:

  • A false diagnosis of nerve damage. One woman was left badly shaken when an AI scribe’s summary of her conversation wrongly stated she had demyelination — serious nerve damage that can lead to multiple sclerosis. The error only surfaced because the patient happened to be an NHS health professional and queried the tool’s record of her MRI scan result. The hospital eventually corrected the entry to what it should have been: “null demyelination.” “This was eventually corrected but was a very traumatising experience to be given an incorrect diagnosis because of AI and then be told it’s a typo,” she said.
  • Drug-name confusion. In another case, an AI scribe confused the drug the GP had actually prescribed with a different drug of a similar name — again spotted by the patient, not the doctor.
  • A dangerous omission. An AI-generated summary letter failed to record that a hospital consultant had told the patient to seek a repeat prescription from their GP for their migraines — a gap that could have left them unable to get their medication.

Healthwatch says it has heard “multiple stories from patients who have noticed these errors when a health professional hasn’t,” and warns that undetected inaccuracies “may persist in their records” — with consequences for future care decisions built on top of them.

Why this matters right now

The timing is what makes the warning significant. The UK government’s 10-year health plan for the NHS in England explicitly expects AI scribes to “liberate staff from their current burden of bureaucracy and administration, freeing up time to care and to focus on the patient.” Ambient scribing is central to the plan’s promised “big shift” from an analogue health service to a digital one. GPs and hospital doctors in England are already using 27 different AI scribe tools, and ministers have been warned that the NHS and individual medics could face lawsuits over mistakes made by AI.

Yet the oversight regime has not kept pace. Healthwatch called it “worrying” that the Medicines and Healthcare products Regulatory Agency (MHRA) has decided not to classify AI scribes as medical devices — which means there is no England-wide mechanism to assure that the tools are safe and effective before or during deployment. A Healthwatch spokesperson put it plainly: “Healthcare has never been error-free. But our findings show the urgent need for clarity over how patients can report and get corrected any mistakes made by AI scribing tools or the professionals that use them.”

Rachel Power, chief executive of the Patients Association, added: “Trust and confidence in this technology depend on good communication and genuine partnership with patients and right now both are missing.”

Hallucinations, accents, and the time-saving paradox

The evidence base behind these concerns has been building. Dr Shier Ziser Dawood, a London GP, warned in the British Journal of General Practice that AI scribes may prove “a double-edged sword” for family doctors. She recounted an incident in which an AI scribe’s note claimed she had told a patient to “continue their Prozac” — a drug she had neither prescribed nor discussed. That is a classic hallucination: the model confidently documenting something that never happened in the consultation.

Her more subversive point is about the economics. If doctors must carefully review every transcript to catch such errors, the tools are not yet saving time — yet NHS planners are already budgeting the gains. Dawood warned that family doctors may be expected to see two more patients every day on the assumption that scribes free them from note-taking, even though UK GP consultation times are already among the shortest in the world.

Dr Charlotte Blease, an expert on AI in healthcare at Uppsala University, found that GPs who use ambient voice technology believe errors creep in most often when more than one person is present in the consultation, when the patient has a complex medical history, or when English is not the patient’s first language. Her survey of 1,003 UK GPs found that more than half believed their ambient AI records were more accurate than notes they wrote themselves — a striking vote of confidence that cuts both ways. “The fact is, doctors can and do make mistakes without AI. And it is certainly possible the error rate is worse,” Blease noted.

The accent problem is real too: patients in Rotherham complained to their local Healthwatch that an AI receptionist used by some GP practices simply could not understand their strong Yorkshire accents.

The accountability gap

Pull the threads together and a consistent picture emerges. Public attitudes are ambivalent — Healthwatch’s July 2026 survey of 2,795 recent patients found nearly 90% were unaware AI scribing was being used at their appointments, and support and opposition split roughly evenly. Deployment is racing ahead across 27 tools. Detection of errors currently depends heavily on patients — the least-informed party in the room — noticing discrepancies in documents they were often never told were machine-generated. And the regulator has opted out of treating these systems as medical devices, leaving no national assurance layer at all.

“Healthcare has never been error-free” is the standard defense, and it is true. But traditional clinical documentation errors at least had an author who could be asked what they meant. An AI scribe’s hallucinated instruction has no memory of the consultation at all — only a probabilistic reconstruction of it. If the NHS’s digital-first vision is to survive contact with real patients, Healthwatch’s asks are the minimum: patients must know when AI is writing their record, must have a clear route to report mistakes, and must have the power to get corrections made before those errors propagate into every downstream decision about their care.

None of this argues against ambient AI in the clinic. It argues for deploying it like safety-critical software rather than office productivity tooling — with adverse-event reporting, error-rate transparency, and a regulatory classification that matches the stakes. Until then, the system’s last line of defense is a patient sharp enough to ask, “Excuse me — did the machine just give me MS?”