← All posts / Research

146 Conditions, One Pass, Open Weights: Alibaba's DAMO RADAR Reads Abdominal CT Better Than 23 of 26 Radiologists

Published in Science and open-sourced under CC BY-NC-SA, Alibaba DAMO's RADAR vision-language model was trained on 424,911 CT exams without manual annotation, hit a mean AUC of 0.913 across 146 abdominal findings, and runs on a single 16GB GPU.

146 Conditions, One Pass, Open Weights: Alibaba's DAMO RADAR Reads Abdominal CT Better Than 23 of 26 Radiologists

For a decade, medical imaging AI has lived by one rule: one model, one disease. Hundreds of FDA-cleared tools can spot a lung nodule or flag a stroke, but each is a narrow specialist. This month, a paper in Science (vol. 393, eaec6129, DOI 10.1126/science.aec6129, appeared September 18, 2026) made the case that the generalist era has arrived in radiology. The model is called RADAR — Rapid Abdominal Diagnosis with AI and Radiology — built by Alibaba’s DAMO Academy with the Hupan Laboratory and the First Affiliated Hospital of Zhejiang University School of Medicine. And in a field where flagship medical AI usually stays locked behind hospital contracts and per-scan API fees, the team did something unusual: they released the weights.

What RADAR actually does

Contrast-enhanced abdominal CT is one of the hardest generalization targets in medical imaging. A single study covers some 18 organs, and the space of things that can go wrong runs to hundreds of findings — tumors, cysts, calcifications, inflammations — many of which are visually subtle and depend heavily on anatomical context. Reading one well is precisely why abdominal radiologists train for years.

RADAR, a vision-language model, was trained on 424,911 contrast-enhanced abdominal CT examinations, paired with the radiology reports doctors had already written for them — 1.5 million image–text pairs, plus more than 15 million anatomy-specific pairs. The crucial engineering choice: instead of treating a CT volume as a blob of voxels, RADAR decomposes it into individual anatomical structures and links each structure to the corresponding language in its report, using large-scale contrastive learning to bind visual patterns to clinical descriptions. Because the reports themselves act as the labels, no manual annotation was needed at scale.

At inference time the model takes two inputs: the CT scan and an organ segmentation mask — a precomputed outline telling the model where the liver, pancreas, stomach, and other organs sit in this particular patient. That mask narrows the task from “find anything unusual in this 3D volume” to “score these organ regions against these known conditions.” The output is a CSV row with a confidence score for each of 146 organ–condition pairs, from liver lesions to aortic calcification. In effect, it is a triage layer — what the DAMO team calls a second pair of eyes — rather than an autonomous reader.

The numbers

Across roughly 40,000 real-world examinations, RADAR achieved a mean AUC of 0.913 across all 146 findings, versus 0.776 for the best competing vision-language model — a decisive gap. Three further results stand out:

  • Emergencies it was never trained for. On more than 27,000 emergency CT cases, RADAR scored a mean AUC of 0.904, despite no emergency-specific training data. Emergency reads are where radiologist fatigue and time pressure bite hardest, so generalization into that setting matters more than any benchmark.
  • Generalization across hospitals. Tested on cohorts from eight external centers, the model held a mean AUC of 0.895 — evidence that it learned findings, not the quirks of one institution’s scanners and protocols. Performance stayed strong across organs, diseases, and uncommon findings.
  • The reader study. In head-to-head reading, RADAR outperformed 23 of 26 expert radiologists across the full finding set. More clinically interesting: when radiologists used RADAR as an assistant, their sensitivity improved by roughly 10% — the shape of deployment that actually matters, a human plus a model beating both alone.

Why open weights change the story

Cancer-screening AI has historically lived behind enterprise licensing and per-scan API pricing, which means the hospitals that need a second reader most — district hospitals, rural clinics, facilities in low-resource settings — are the last to get it. RADAR flips the distribution model. The code, checkpoints, and supporting files are on GitHub (alibaba-damo-academy/damo-radar), Hugging Face, and ModelScope.

Independent testing found the hardware footprint modest by frontier-model standards: under 1GB of VRAM to initialize, settling around 16GB during inference on a single consumer-grade GPU, with tunable KV-cache and context-length parameters to shrink it further. The setup is a standard open-source ML workflow — clone the repo, run the download script, load the checkpoint alongside a BERT-based text encoder used for the vision-language alignment, feed in a scan plus organ mask, and read out the scored CSV. The researchers also ship training, fine-tuning, and preprocessing code, so the release is a research framework, not just a frozen artifact.

The caveats — and they are real

The GitHub README is blunt: RADAR is intended for research purposes only, and prospective clinical studies are still required before clinical deployment. The license is CC BY-NC-SA 4.0 — non-commercial, which means hospitals cannot simply productize it without a separate agreement. And the output format demands expertise to interpret: a 0.70 confidence on aortic calcification means something only to a trained clinician with the patient in front of them. Reading the spreadsheet without that background invites misunderstanding.

The release also lands mid-debate about evidence standards. Just days ago, clinicians quoted in the Financial Times warned that medical AI is being deployed with thin peer-reviewed evidence beyond imaging’s traditional beachhead. RADAR sits on the right side of that argument — this is a Science-paper release with weights attached, not a product launch with a press release — but the gap between “beats 23 of 26 radiologists in a reader study” and “safe to deploy in your hospital’s workflow” is exactly the gap the paper’s own authors say still needs to be crossed.

The bigger picture

Radiology AI has been stuck in a specialist trap: each narrow model needs its own labeled dataset, its own deployment, and its own integration. RADAR’s bet is that vision-language pretraining on existing clinical reports — data every hospital already generates — is the escape hatch. A generalist that scores 146 conditions in one pass, generalizes across eight external centers, holds up in emergency settings it never saw in training, and runs on hardware a district hospital can buy, redraws the boundary of what “deployable” means.

The name to watch is the pattern, not just the model: frontier-grade medical AI arriving as open research artifacts rather than locked APIs. If the generalist-plus-open-weights combination holds up in prospective trials, the second opinion that today costs a per-scan API fee could become something a clinic downloads once and runs locally — and the bottleneck shifts from access to clinical validation. That is a much better problem to have.

Sources are listed in the frontmatter of this post.