← All posts / Research

One Model, Three Countries, a Million Patients: Google's ARDA Retinal AI Publishes Its Scaling Lessons

A Nature Medicine Comment from Google and clinical partners in India, Thailand and Australia distils what a decade of deploying diabetic-retinopathy AI taught about scaling clinical models beyond the pilot stage.

One Model, Three Countries, a Million Patients: Google's ARDA Retinal AI Publishes Its Scaling Lessons

One Model, Three Countries, a Million Patients: Google’s ARDA Retinal AI Publishes Its Scaling Lessons

Most clinical AI never leaves the hospital it was validated in. On September 23, 2026, Google published the receipts for one that did: a Comment in Nature Medicine documenting how a single deep-learning system for diabetic-retinopathy screening has now been used on more than one million patients across three radically different health systems — Aravind Eye Care System in Madurai, India; Rajavithi Hospital in Bangkok, Thailand; and Lions Outback Vision in Western Australia.

The paper, led by Richa Tiwari, Rajroshan Sawhney, and Kasumi Widner of Google with clinicians at each partner institution, is not a trial write-up. It is something rarer: a deployment retrospective. The authors deliberately frame it around “cross-cutting insights that may inform the expansion of healthcare artificial intelligence globally” — the unglamorous, between-the-algorithm-and-the-patient work that determines whether medical AI survives contact with a real clinic.

The disease and the tool

Diabetic retinopathy is a leading cause of preventable blindness. Nearly half of people with diabetes will develop some form of it, and the global patient population — at least 537 million adults, with roughly 227 million in the Asia-Pacific region alone — dwarfs the number of ophthalmologists available to screen them. In India and Thailand, the specialist shortage is acute enough that many patients are diagnosed only after their vision has already begun to fail.

The tool at the centre of the paper is ARDA — Automated Retinal Disease Assessment — a deep-learning model that reads retinal photographs and flags referable diabetic retinopathy and diabetic macular edema. Google began the research programme nearly a decade ago with Aravind and Rajavithi as foundational partners, published early validation work in JAMA as far back as 2016, and has been running real clinical deployments since 2018-2019. An earlier post-marketing report covered the first 600,000 patients screened in India; the new Comment pushes the documented total past one million across all three anchor deployments.

Critically, Google no longer runs these programmes alone. The model is licensed to health-tech partners — Forus Health and AuroLab in India, Perceptra in Thailand — who handle local regulatory approvals, hardware integration, and clinic operations. The stated ambition is a combined 6 million AI-supported screenings at no cost to patients across India and Thailand over the next decade, with Thailand’s Ministry of Public Health Department of Medical Services working the model into the national screening innovation programme.

What a million patients actually taught

The paper’s value is in its refusal to pretend scaling is a software problem. Three themes dominate the lessons the authors draw.

First, the bottleneck is workflow, not accuracy. A model that performs well in a lab changes nothing until it is embedded in a clinic’s patient flow: who takes the image, who counsels the patient, what happens in the ninety seconds after the model returns a result, and who handles the referral if the answer is “referable.” Each of the three anchor sites solved this differently — Aravind inside high-volume specialty eye hospitals, Rajavithi inside a public general hospital tied to a national screening programme, Lions Outback Vision through mobile and remote outreach across vast rural distances in Western Australia. The same algorithm succeeded in all three because the surrounding process was redesigned each time.

Second, deployment surfaces data-integrity debt that pilots hide. As clinical-operations analysts reading the Comment noted, local validation metrics look clean while aggregate performance conceals variance driven by inconsistent image labelling, site-specific acquisition protocols, and patient demographic mismatches. A camera firmware difference between sites is invisible to the model and decisive for its output. The million-patient dataset makes that variance measurable in a way no single-site pilot ever could.

Third, monitoring is a permanent obligation, not a launch task. The earlier India post-marketing report was itself an argument for publishing algorithm performance after deployment — “consistent with recommendations by regulatory bodies” — and the new Comment extends that stance: performance must be watched continuously across the full demographic and operational range of the deployed system, not certified once at study start. This aligns with the direction regulators are already moving, from the FDA’s total-product-lifecycle thinking for AI-enabled devices to the WHO’s six regulatory domains for health AI.

Why the Comment format matters — and what it omits

It is worth being precise about what the paper is not. It publishes no per-country accuracy figures, no camera-model breakdowns, and no false-referral rates. As AI Weekly’s editorial note put it, “the Comment format leaves per-country failure modes and accuracy figures off the page.” The framing is deliberately workflow-and-deployment lessons rather than clinical performance numbers — those live in the earlier peer-reviewed evaluations.

That restraint cuts both ways. It means the million-patient figure cannot be independently decomposed into sensitivity and specificity by country from this paper alone. But it also means the authors are publishing something vendors rarely volunteer: the operational lessons that determine whether a clinically excellent model dies in a pilot or reaches a million people.

The broader significance

For the clinical-AI field, the ARDA retrospective lands at a moment of sober reckoning. Much of the literature consists of single-centre validation studies whose real-world performance collapses on contact with new EHR schemas, imaging protocols, and patient populations. A public, decade-long, three-country deployment record — including the failures and course corrections — is a genuine reference point for hospital purchasers, regulators, and the growing number of companies now shipping clinical foundation models.

It also sketches a template for how Big Tech exits clinical AI operations responsibly: Google contributes the model and the research partnership, local licensees own regulatory approval and delivery, and a public health ministry integrates the tool into an existing national programme. Whether that template generalises to less well-characterised conditions than retinopathy — where the ground truth is a specialist’s judgement rather than a photograph — remains the open question.

One million screenings is a large number and a small one at once: a rounding error against 537 million people with diabetes, and an existence proof that clinical AI can be governed, monitored, and scaled across borders without collapsing. The lesson the paper actually teaches is that the hard part was never the model.

Sources