Medical AI Has a Proof Problem: Clinicians Push Back on Deployment Beyond Diagnostics
The FT's Big Read argues medical AI is being deployed at enormous scale on remarkably thin evidence that it improves patient outcomes — and the clinicians being asked to trust it are pushing back.
The most consequential AI story of September 20, 2026 is not a model release. It is a warning from the people expected to use the technology: medical AI is being deployed at enormous scale on remarkably thin evidence that it actually improves patient outcomes.
That is the thesis of “Medical AI has a proof problem,” Sarah Neville’s Big Read in the Financial Times published September 19-20, 2026, which has become the top-shared AI story of the weekend among health professionals. The piece captures a tension the industry has been circling for months: adoption is racing far ahead of the evidence base, and clinicians — the ones whose signatures ultimately put these tools into patient workflows — are increasingly refusing to treat deployment as a substitute for proof.
What the FT piece says
Neville, the FT’s global health editor, reports that clinicians are pushing back on medical AI’s expansion beyond its proven territory — imaging and diagnostic assistance — into broader clinical workflows: triage, documentation, treatment suggestions, and decision support. The core complaint is not that the tools are useless. It is that hospitals are rolling them out with limited peer-reviewed performance data behind the specific deployments.
Three concerns dominate the clinician response documented in the piece:
- Deskilling. Recent surveys cited by the FT flag deskilling fears among roughly 74 percent of clinicians — the worry that clinicians who lean on AI recommendations gradually lose the independent judgment that makes them a safety net rather than a rubber stamp.
- Hallucinations. For generative tools built on large language models, studies of clinical decision support estimate hallucination rates of 8 to 20 percent — a range that is tolerable in a search engine and alarming in a treatment recommendation.
- Unclear governance. Institutional AI policies remain in flux across hospital systems, with no settled answer to who is accountable when a deployed model drifts, degrades, or simply fails in production.
The piece lands at a moment of maximum irony for the field. The same month, Alibaba’s DAMO Academy published Damo Radar in Science — a vision-language model that beat 23 of 26 expert radiologists across 146 abdominal CT findings on roughly 40,000 real-world exams. The frontier of what medical AI can do keeps advancing. The FT’s argument is about a different question: what deployed medical AI has actually been shown to do, for which patients, in which workflows.
“Shockingly poor evidence”
The line from the FT piece that clinicians keep quoting: “there is really shockingly poor evidence that AI actually makes an impact on patient outcomes.”
That quote resonated immediately with Amy Abernethy, the former FDA deputy commissioner who now co-leads Highlander Health. In a widely shared LinkedIn response to the article, Abernethy — who was quoted in the FT piece itself — wrote: “We are deploying medical AI at enormous scale on remarkably thin evidence that it improves patient outcomes. I was quoted in this piece, but I have been thinking about this topic ever since my Duke days.”
Her diagnosis is structural, not motivational: “The obstacle is not a lack of will to evaluate these tools. But traditional clinical trial approaches don’t scale to meet this moment.” Running a randomized controlled trial for every AI tool, in every workflow, at every update cadence, is simply not feasible when models are revised weekly. Her proposed alternative is to evaluate tools using real-world data already generated by the healthcare delivery system itself — an approach her organization is building toward through clinician-led “learning labs” inside health systems.
She is blunt about why this is slow: “When you ask a health system to contribute records to an evaluation, you are asking its general counsel to accept risk with little corresponding protection. The result is a field that demands proof alongside a data infrastructure that makes proof slow and expensive to produce.”
The numbers behind the unease
The clinician skepticism documented by the FT is not free-floating anxiety — it maps onto a measurable evidence gap:
- The FDA has authorized more than 1,450 AI-enabled medical devices as of end-2025, clearing 295 in 2025 alone. Yet fewer than 2 percent of cleared devices were supported by randomized clinical trials.
- Two-thirds of clinicians now use AI in their work, according to the American Medical Association.
- A 2025 analysis of FDA-cleared AI devices found most 510(k) premarket summaries lack details on study design, sample sizes, and demographic representation.
And the oversight picture is getting weaker, not stronger. In January 2026, the Department of Health and Human Services released a proposed rule, HTI-5, that would remove the source-attribute transparency requirements and intervention risk-management provisions that the agency itself finalized in late 2023 — the closest thing U.S. clinical AI had to a model-card mandate. As legal analyst Michael Craige argued in The Regulatory Review, the rollback “does not relieve a burden; it shifts the burden onto the buyer,” leaving hospital procurement and compliance teams to construct bespoke diligence for every algorithm in every workflow, at exactly the moment the fastest-growing category — generative AI assistants — falls largely outside FDA jurisdiction altogether.
The counterweight: where proof does exist
The FT piece is careful to note that the proof problem is not uniform across the field. Where rigorous real-world evaluation has been done, results can be striking.
The strongest example comes from Kenya. A landmark study by Penda Health and OpenAI, embedded in sixteen Nairobi clinics covering some 20,000 patient visits, found that an AI “safety net” tool called AI Consult — an off-the-shelf GPT-4-based assistant reviewing clinicians’ work — reduced diagnostic errors by 16 percent and treatment errors by 13 percent, with a 32 percent decrease in history-taking omissions and zero new safety events introduced. The study, published in Nature partner journal Medicine in 2026, is precisely the kind of evidence clinicians say they want: prospective, real-world, outcome-measured, and workflow-embedded.
Notably, that study evaluated AI as a checker of human work rather than a replacement for it — a deployment pattern that sidesteps the deskilling trap by keeping the clinician firmly in the loop while AI audits the encounter.
Why this matters now
The timing of the FT piece matters. It arrives amid a week of conflicting signals on AI safety in general — the same weekend as reports that U.S. frontier labs are coordinating safety standards while a class action accuses them of coordinating a slowdown, and days after leaked forecasts showed OpenAI expecting nearly $280 billion in cash burn through 2030. Medical AI sits at the intersection of both currents: enormous commercial pressure to deploy, and a growing institutional insistence that deployment without evidence is a liability, not a product.
The practical stakes are straightforward. If the evidence gap persists while adoption accelerates, the likely endpoint is a high-profile failure — a hallucinated treatment plan, a drifted model quietly degrading care, a deskilled clinician missing what the AI missed — that triggers a regulatory reckoning far harsher than the transparency rules the industry is currently dismantling.
If the field instead builds the real-world evaluation infrastructure Abernethy describes — learning labs, curated longitudinal datasets, decision-to-outcome audit trails — the Kenya results suggest the tools can genuinely earn the trust clinicians are being asked to extend.
The proof problem, in other words, is solvable. But as this weekend’s most-shared AI story makes clear, solving it requires the industry to spend on evaluation what it currently spends on deployment. Right now, it isn’t close.
Sources
- [1] https://www.ft.com/content/34319b00-f874-4119-aa28-8376d81e7190
- [2] https://www.linkedin.com/posts/amyabernethy_medical-ai-has-a-proof-problem-activity-7506813136679776257-ntKL
- [3] https://www.theregreview.org/2026/08/06/craige-the-governance-gap-in-clinical-ai/
- [4] https://www.nature.com/articles/s44360-026-00082-5
- [5] https://time.com/7304457/ai-prevents-medical-errors-clinics/