← All posts / Research

One Model, 146 Diseases: Alibaba's DAMO RADAR Reads Abdominal CT Scans Better Than 23 of 26 Radiologists — and It's Fully Open-Sourced

Published in Science on September 17, DAMO RADAR is a vision-language model trained on 400,000+ CT exams that matches expert radiologists across 146 abdominal findings — and Alibaba has released the weights, code, and training framework for anyone to build on.

One Model, 146 Diseases: Alibaba's DAMO RADAR Reads Abdominal CT Scans Better Than 23 of 26 Radiologists — and It's Fully Open-Sourced

For a decade, medical imaging AI has lived in a strange niche: thousands of narrow tools, each licensed to spot one thing — a lung nodule here, a breast lesion there — while the radiologist reading the scan still had to synthesize everything else. On September 17, that paradigm took a direct hit. Science published a paper from Alibaba’s DAMO Academy and the First Affiliated Hospital of Zhejiang University School of Medicine describing DAMO RADAR (Rapid Abdominal Diagnosis with AI and Radiology), a single vision-language model that reads contrast-enhanced abdominal CT scans across 18 anatomical structures and 146 clinical findings — and the team has open-sourced the model, the code, and the entire training framework.

In a head-to-head reader study, RADAR’s average performance exceeded that of 23 of 26 radiologists drawn from 14 hospitals, falling short of only three senior physicians. When paired with the model, human readers’ diagnostic sensitivity rose by roughly 10% while their average reading time dropped by more than 30%. A junior doctor working with RADAR outperformed a senior doctor working alone.

Why the abdomen was the final boss

If you want to understand why this paper landed in Science rather than a specialist venue, look at the anatomy. The abdomen is widely considered the hardest region in conventional diagnostic imaging. The digestive, urinary, and reproductive systems are packed together; the intestines coil through the entire cavity; and nearly every organ is soft tissue of similar density. Spotting a lesion often comes down to a fine sensitivity to density differences that takes radiologists years to develop.

It is also a graveyard for the supervised-learning approach that produced most clinical AI to date. Zhang Jianpeng, a senior algorithm expert at DAMO Academy, noted that a radiologist typically needs two to three years to master a single disease type, while a real CT scan frequently contains multiple organs and multiple lesions at once. Hand-labeling a dataset broad enough to cover that reality is economically impossible — which is why prior tools stayed narrow, and why “AI nodule detectors” for lungs shipped years ago while the abdomen remained a no-go zone.

How RADAR learns without manual labels

RADAR sidesteps the annotation bottleneck entirely. Instead of doctors drawing boxes slice by slice, the model learns directly from what hospitals already produce in enormous volumes: CT images and the diagnostic reports that accompany them. The training corpus exceeded 400,000 contrast-enhanced abdominal CT examinations, converted into roughly 15 million anatomy-wise image-text pairs, according to the paper’s abstract.

Two methodological innovations carry the weight:

  • Fine-grained organ-level alignment. The model first localizes individual organs — liver, pancreas, gallbladder, kidneys — and then aligns each organ unit with the corresponding passage in the report. Rather than matching a whole scan to a whole report, it learns which sentence describes which structure.
  • Adaptive contrastive modeling. In standard contrastive learning, every other sample is pushed away. In medicine that’s wrong: two patients with healthy livers are clinically similar. RADAR adjusts the distance between samples according to medical knowledge, so the geometry of its embedding space reflects clinical relationships rather than arbitrary identity.

The combination is what lets one model generalize across 146 findings instead of overfitting to a handful of heavily annotated diseases.

The numbers that convinced the reviewers

The validation stack in the paper is unusually deep, spanning internal cohorts, external hospitals, emergency cases, and pathology-anchored cancer benchmarks:

EvaluationCasesResult
Internal consecutive real-world cohort~39,000mean AUC 0.913 across 146 findings
External hospitals (8 sites, varied regions/scanners)24,000+AUC 0.895
Emergency cases (never in the training objective)~27,000AUC 0.904
Liver, pancreatic, gastric, colorectal cancers vs. pathology—AUC 0.891–0.984

An AUC of 1.0 represents perfect discrimination; 0.5 is a coin flip. Holding above 0.89 on out-of-distribution scanners and emergency populations — a setting the model was never explicitly trained for — is the difference between a demo and something a hospital can actually deploy. And the scaling curves in the paper show performance still rising with more data, with no sign of saturation.

Zhang Ling, another senior algorithm expert on the team, told Chinese media that Science rarely publishes medical imaging AI because the field has long been treated as an engineering problem. RADAR, he argued, is the first demonstration that generalist medical imaging AI is a realistic technical path — and that it belongs in a science journal, not just an engineering one.

The open-source release is the real story

A Science paper would be notable on its own. The fact that everything shipped publicly at the same time is what changes the practical picture. The GitHub repository (alibaba-damo-academy/damo-radar, live since July with final updates on September 18) carries the weights, inference code, and the training framework, with documentation also mirrored on Zenodo and model weights on Hugging Face. Any hospital, university, or startup can now reproduce, fine-tune, and extend the system without negotiating a license with Alibaba.

That stands in sharp contrast to the dominant Western pattern, where strong medical AI models tend to live inside FDA-cleared commercial products or locked-down APIs. It also continues a deliberate Chinese strategy of open-weight releases — from DeepSeek to Qwen — that has been steadily converting global developer mindshare into geopolitical soft power. A radiology department in Brazil, Egypt, or Vietnam can download expert-level abdominal CT triage tonight. Whether that is a public-health triumph or a regulatory headache depends on jurisdictions that have not written rules for it yet.

The honest caveats

The clinicians involved are not uniformly celebratory, and their concerns are worth taking seriously. Xiao Wenbo, director of radiology at Zhejiang University’s First Affiliated Hospital, said she initially could not believe the validation data, and that many doctors in her department pushed for immediate deployment. But her principal worry is not accuracy — a human still signs off on every report. It is de-skilling: the risk that junior doctors who grow up with RADAR never build the imaging reasoning that currently takes years to develop. Her proposed mitigation is to run RADAR as a one-on-one teaching aid instead — letting trainees self-check their draft reports and instantly see which lesion they missed.

There are also boundaries the paper itself acknowledges. RADAR covers contrast-enhanced CT of the abdomen; the team says the organ-level alignment approach should transfer to MRI, PET, and ultrasound “given suitable data,” but that transfer is future work, not a demonstrated result. The reader study, while multi-hospital, is still a study — prospective clinical trials, regulatory clearance, and integration into messy hospital workflows remain between this paper and routine patient care. And a generalist model that flags 146 findings will inevitably generate its own flood of incidental findings that someone has to triage.

What it means

For a decade the industry has argued about whether “generalist” medical AI — one model covering the full breadth of a specialty — was even possible, or whether medicine’s error tolerance made it a fantasy. The RADAR team’s framing is pointed: medicine is the domain with the most extreme demand for precision and the lowest tolerance for error, so if the generalist paradigm works here, it is likely to carry over to other high-stakes settings.

Three things to watch next: how many hospitals actually deploy the open framework (the repo had already gathered hundreds of stars within hours of the Science publication); how well the approach transfers to other modalities; and whether Western regulators respond to open-weight clinical AI with acceptance, restriction, or silence. What is no longer debatable is the direction of travel. The one-model-per-disease era of medical AI just got its most credible alternative — and anyone can download it.