← All posts / Research

One Model, 146 Diseases: Alibaba's DAMO RADAR Reads Abdominal CT Better Than 23 of 26 Radiologists

Alibaba DAMO Academy and Zhejiang University's open-source RADAR, published in Science, identifies 146 conditions across 18 abdominal organs from contrast CT scans — AUC 0.913 on ~39,000 internal exams and 0.895 across 8 external hospitals — beating 23 of 26 radiologists in a head-to-head reader study.

One Model, 146 Diseases: Alibaba's DAMO RADAR Reads Abdominal CT Better Than 23 of 26 Radiologists

For a decade, medical imaging AI has followed one template: pick a disease, get radiologists to label thousands of scans, train a classifier, repeat. This week in Science, Alibaba’s DAMO Academy and the First Affiliated Hospital of Zhejiang University published a deliberate break from that template. Their model — RADAR, short for Rapid Abdominal Diagnosis with AI and Radiology — is a single generalist vision-language system that reads contrast-enhanced abdominal CT and flags 146 distinct diseases across 18 anatomical structures, from liver and pancreatic cancer to fatty liver disease and acute appendicitis. And in a reader study against 26 practicing radiologists from 14 hospitals, it outperformed 23 of them.

The model, its code, and its framework are fully open-sourced on GitHub, Hugging Face, and Zenodo — an unusually aggressive release for a system the researchers describe as the first expert-level generalist medical imaging AI.

Why the abdomen, and why generalism

The abdomen is widely regarded as the hardest playground in radiology. The liver, gallbladder, pancreas, spleen, kidneys and coiled intestines are packed together, nearly all of them soft tissue with similar density; a single study contains hundreds of slices, and the clinically decisive finding may be one tiny lesion on one organ. Xiao Wenbo, director of radiology at the First Affiliated Hospital of Zhejiang University, notes that junior doctors dread being assigned to the abdominal imaging rotation — it is the domain where reading skill depends almost entirely on a radiologist’s trained sensitivity to subtle density differences.

DAMO’s team arrived at generalism out of frustration with the alternative. The academy had already published five Nature Medicine papers on single-disease detectors — early pancreatic cancer, gastric cancer, colorectal cancer — but senior algorithm expert Zhang Jianpeng put the math bluntly: at two to three years to conquer one disease, and tens of thousands of known human diseases, the dedicated-model approach can never finish. Worse, real patients do not arrive with one disease at a time, and a purpose-built model is blind to everything outside its target.

How RADAR learns — without hand labels

RADAR’s most consequential design choice is that it requires no manual annotation. It was trained on more than 400,000 contrast-enhanced abdominal CT examinations paired with the clinical reports that hospitals already produce, yielding roughly 15 million anatomy-aware image–text pairs. Instead of learning from pixel labels, it learns from the correspondence between images and the sentences radiologists wrote about them — closer to a medical student shadowing ward rounds than to a classifier drilled on a spreadsheet.

Two technical insights make this work at acceptable fidelity:

  • Organ-level fine-grained alignment. Naively aligning a whole CT volume with a whole report fails — the abnormal signal from one small lesion drowns in hundreds of normal slices, and the model learns only crude statistics. RADAR first segments the volume into anatomical units (liver, pancreas, gallbladder, kidney, and so on), then aligns each unit with the corresponding sentences in the report. It mirrors how a radiologist actually reads: organ by organ, not abdomen-as-blob.
  • Adaptive contrastive modeling. Standard contrastive learning pushes every patient’s features apart by default. That is wrong in medicine: two healthy livers should sit close together, and two livers with the same tumor should too. RADAR dynamically adjusts inter-sample distances according to this clinical logic rather than patient identity.

Notably, the paper’s scaling curve is still rising with more data — no saturation yet — which is the strongest argument that 146 diseases is a waypoint rather than a ceiling.

The numbers

The validation is layered, and the external cohorts are where the claims earn their keep:

CohortScaleResult
Internal consecutive real-world exams~39,000average AUC 0.913 across 146 findings
External validation, 8 hospitals>24,000 scansAUC 0.895, single-center AUCs from 0.874 upward
Emergency-department cases (never seen in training)27,000AUC 0.904
Four cancers vs. pathological gold standard (liver, pancreas, stomach, colorectal)biopsy-confirmed ground truthAUC 0.891–0.984

The external set matters because it spans different regions, different scanners, and different imaging protocols. The emergency result matters even more: trauma and acute-care scans are rushed, contrast timing is suboptimal, patients move — and RADAR, whose training distribution explicitly excluded emergency cases, still held a 0.904 AUC. That is a genuine generalization signal, not a curve fit to one hospital’s scanner fleet.

The cancer benchmark is the sharpest test of all. For hepatocellular carcinoma, pancreatic cancer, gastric cancer and colorectal cancer, the team stopped grading against radiologist reports and used pathology biopsy results as ground truth — the actual gold standard of oncology — and the model still landed between 0.891 and 0.984 AUC.

Human versus machine — and human plus machine

The reader study pitted RADAR against 26 radiologists from 14 hospitals, all independently diagnosing the same case batches. RADAR beat 23, falling only slightly short of three senior subspecialists. But the more operationally interesting finding is the collaboration experiment: when doctors read with RADAR’s assistance, overall sensitivity rose by roughly 10 percentage points and average reading time dropped by more than 30% — and the combination of a junior doctor plus AI exceeded the sensitivity of senior radiologists working alone.

Contextualize that against the workflow math: even a senior radiologist needs about 20 minutes to write a full abdominal CT report, tracing hundreds of images. Departments like Xiao’s process thousands of scans daily. A tool that compresses that time by a third while catching 10% more true positives is not a party trick — it is a capacity multiplier for the exact bottleneck in diagnostic medicine.

Fully open, with caveats worth stating

The release is unusually complete: training and inference code (Apache 2.0) on GitHub and Zenodo, checkpoints and supporting data on Hugging Face, preprocessing pipelines for external datasets like MERLIN, plus a demo inference case. The research community can fine-tune, audit, and extend it — which is precisely what a claim as large as “expert-level generalist radiology AI” needs. The model weights carry a CC BY-NC-SA 4.0 license, so commercial clinical deployment requires separate arrangements, and none of the validation yet substitutes for prospective regulatory trials. An AUC measured on retrospective Chinese cohorts does not automatically transfer to a US hospital’s scanner mix, prevalence base rates, or patient demographics — the emergency-cohort robustness is encouraging, but it is still within one health system’s universe.

The trajectory, though, is the story. Medical AI spent a decade proving it could match specialists at one narrow task at a time. RADAR is evidence that the generalist route — vision-language learning over routinely collected clinical data, no manual labels required — can cover the full breadth of the hardest imaging domain at expert level, and that its capability has not yet hit the ceiling. Zhang Jianpeng’s complaint that single-disease AI could never finish the job now has a concrete alternative: a model that already reads 146 diseases, whose scaling curve is still climbing, and whose weights anyone can download today.

For hospitals short of senior radiologists — which is most hospitals on Earth — that is not an incremental improvement. It is a different category of artifact.