← All posts / Research

1,357 AI Medical Devices Cleared by the FDA — Only 3 Tested on Patient Outcomes

A PLOS Digital Health review finds that of 1,357 FDA-authorized AI medical devices, just 3 were evaluated on outcomes patients actually care about — survival, hospitalization, quality of life.

1,357 AI Medical Devices Cleared by the FDA — Only 3 Tested on Patient Outcomes

Artificial intelligence has quietly become part of the furniture of American medicine. It flags suspected strokes on CT scans, estimates cardiovascular risk, prioritizes mammograms for radiologists, and helps surgeons plan procedures. But a sweeping review published this week in PLOS Digital Health raises a deeply uncomfortable question: how many of these systems have ever been shown to make patients healthier?

The answer, according to researchers led by Rawan Abulibdeh of the University of Toronto, is almost none. Of the 1,357 AI-enabled medical devices authorized by the U.S. Food and Drug Administration through December 5, 2025, only three were evaluated using patient-centered outcomes — measures like survival, stroke, hospitalization, or quality of life. Everything else rests on thinner evidence: technical benchmarks, retrospective datasets, or demonstrations of similarity to tools already on the market.

What the Study Found

The analysis, titled “1,357 AI medical devices cleared, 3 actually tested on patient outcomes” (PLOS Digital Health, 5(8): e0001597, DOI: 10.1371/journal.pdig.0001597), systematically examined how FDA-cleared AI devices were evaluated in human patients before entering clinical care. The numbers are stark:

  • 1,357 AI/ML-enabled devices authorized by the FDA through December 5, 2025
  • 34 devices (about 2.5%) were linked to registered clinical trials at all
  • 12 devices had trial results publicly available
  • 12 devices had peer-reviewed manuscripts published
  • 3 devices — less than 0.25% — were studied against outcomes directly meaningful to patients

The devices in question are not niche experiments. They span surgical planning support, cardiovascular risk estimation, mammography and other medical imaging analysis, and a widening range of diagnostic and clinical decision-support tools. Most rely on machine-learning models that detect statistical patterns in large datasets and use them to classify images, generate risk scores, or recommend clinical actions.

The core problem, the authors argue, is that technical performance is not the same as improved health. An algorithm can detect an abnormality accurately without proving that its use leads to earlier treatment, fewer complications, longer survival, or better quality of life. Accuracy on a curated dataset is a laboratory property; benefit at the bedside is a clinical one — and the gap between them is where patients live.

Why the Evidence Gap Exists

Much of the explanation lies in how medical devices reach the U.S. market. Many AI tools enter through a pathway in which manufacturers demonstrate “substantial equivalence” to an already-authorized device — the logic being that a new product sufficiently similar in safety and intended use to a predicate device can be cleared without fresh clinical trials.

That framework made sense for traditional devices where similarity was physical and intuitive. For AI systems, it is far shakier. A newer algorithm may produce results faster, or match expert interpretations on selected cases, but that does not show it reduces diagnostic errors in routine practice — let alone that it improves outcomes across the diverse populations encountered in real health systems. Each generation of “equivalent” software can drift further from the original evidence base, a game of regulatory telephone played at machine speed.

The review also documented who is missing from the evidence that does exist. Studies were concentrated in highly resourced healthcare environments, while pregnant women, adults over 75, and people who do not speak English were frequently excluded from evaluations. These omissions matter because medical algorithms can behave differently when applied to patients whose age, language, disease severity, imaging equipment, or treatment access differs from the development dataset. A system that performs well in a specialized academic hospital may be markedly less reliable in a rural clinic or an under-resourced health system.

The Stakes

The consequences extend beyond statistical uncertainty. If an AI system produces more false negatives for a particular group, clinicians may miss serious disease in exactly the patients already facing worse outcomes. If it generates excessive false positives, patients undergo unnecessary tests, procedures, anxiety, and expense. These effects compound when an algorithm is embedded directly into electronic health records and clinical workflows, where its recommendations can carry an aura of authority the underlying evidence doesn’t support.

There is also an international dimension. Many countries — especially those with limited regulatory resources — look to decisions made by agencies in wealthier nations when deciding which medical technologies to adopt. A device cleared in the United States but never tested across diverse populations or care settings may end up deployed in low- and middle-income countries where disease patterns, infrastructure, staffing, and follow-up care all differ. Patients elsewhere can become inadvertent participants in a large, uncontrolled experiment.

The authors also point to structural incentives: developers face commercial pressure to reach market quickly, while outcome-focused trials are expensive, slow, and logistically complex. Proving that a device changes hospitalization rates requires larger and longer studies than proving it detects a feature on an image. In that environment, technical validation becomes the dominant benchmark — even though the ultimate purpose of medical technology is to help people live longer, healthier lives.

A Proposed Fix

Abulibdeh and colleagues propose redesigning the evaluation framework around a three-phase approach: moving beyond a single premarket demonstration of similarity, requiring evidence that systems work across diverse patient subgroups and healthcare settings, and connecting pre-approval validation with continued evaluation after deployment. The shift they describe is philosophical as much as procedural — from asking whether a machine can reproduce a label, to asking whether its use changes decisions, improves care, avoids harm, and delivers benefits fairly.

Notably, the study does not claim that AI medical devices are ineffective, or that algorithms have no place in healthcare. Its point is narrower and more damning: we largely don’t know what these tools do to patients, because almost nobody has checked. As AI becomes increasingly visible in hospitals and clinics, regulatory clearance risks being mistaken for proof of benefit. The researchers insist these are fundamentally different claims — a device can be legally cleared because it resembles an existing product while still lacking any evidence that it helps people live longer or better.

For regulators weighing how to govern medical AI, for hospitals deciding what to purchase, and for patients encountering algorithmic recommendations in their care, the study lands as a bracing reality check. The question it leaves hanging is urgent: before intelligent systems become routine parts of medicine, what standard of evidence should be required to show they actually make care better?