St. Jude's AdaptiveFlow Screens 69 Billion Molecules for 1,000x Less: AI Drug Discovery Gets a Cloud-Native Rewrite
St. Jude's open-source AdaptiveFlow platform steers 69-billion-molecule virtual screens with active learning, scales to 5.6 million cloud CPUs, and cuts screening costs up to 1,000-fold while discovering validated nanomolar FSP1 and PARP-1 inhibitors.
Finding a new drug starts with an brutal numbers game: chemists estimate there are something like 10^60 drug-like molecules, and no lab on Earth can synthesize and test more than a sliver of them. Virtual screening is how the field has fought back — computationally “docking” candidate molecules into a target protein’s binding pocket and ranking them by predicted affinity. But even in silicon, the math turns hostile. Exhaustively screening Enamine’s REAL Space, the largest commercially accessible library of ready-to-synthesize compounds at 69 billion molecules, would cost far more than most academic labs — or even most companies — could ever justify spending on a single target.
On September 1, 2026, St. Jude Children’s Research Hospital announced AdaptiveFlow, an open-source platform that it says “redefines large-scale cloud computing for drug discovery.” The claims are specific and large: near-linear scaling up to 5.6 million CPUs on AWS, compatibility with more than 1,500 docking protocols, and up to a 1,000-fold reduction in computational cost for ultra-large virtual screens (ULVSs) — achieved not by brute force, but by making the search adaptive.
The problem with brute force
Virtual screening has always scaled with money. The previous generation of platforms, including the same group’s VirtualFlow 2.0, proved that billion-scale screens were technically feasible on cloud infrastructure. The catch was economics: cost grows roughly linearly with library size, so going from 1 billion to 69 billion compounds means a 69-fold budget increase for a marginal gain in hit diversity. As a result, ULVSs remained the province of well-funded collaborations, and the vast middle of chemical space stayed unexplored.
AdaptiveFlow’s answer is a method the authors call Adaptive Target-Guided Virtual Screening (ATG-VS). Instead of docking all 69 billion molecules, the platform first organizes chemical space into a multi-dimensional grid of molecular properties, then screens a small set of representative molecules from each “tranche” — roughly 12 million representatives for the entire REAL Space, a 5,750-fold reduction in the initial workload. An active-learning loop then uses the prescreen results as labels: molecules classified as high-confidence binders earn full docking, while unpromising regions of chemical space are deprioritized before any expensive computation is spent on them. The platform applies a threshold-based classification where the 75th percentile of docking scores serves as the decision boundary, concentrating resources where the evidence says the hits probably are.
The headline result: applied to the full 69-billion-compound REAL Space, ATG-VS cuts screening costs by up to 1,000-fold compared to exhaustive search — while, per the paper, “preserving strong enrichment of top hits.”
Three modules, 5.6 million CPUs, and GPUs too
AdaptiveFlow is engineered as three core components. The AdaptiveFlow Ligand Preparation (AFLP) module handles preprocessing of massive libraries — the unglamorous but decisive work of converting 69 billion catalog entries into a ready-to-dock 3D format, which the team completed for the free, screening-ready REAL Space release now available through the AWS Registry of Open Data. The AdaptiveFlow for Virtual Screening (AFVS) engine is the execution layer, supporting more than 1,500 docking protocol combinations, mixing classical physics-based docking with deep-learning-based docking protocols. And AdaptiveFlow Unity (AFU) ties the pieces into a unified, modular workflow that runs on both CPU clusters and GPU architectures.
That last point matters more than it might seem. Existing platforms were largely CPU-only, and the field’s compute center of gravity is shifting to GPUs. AdaptiveFlow integrates deep-learning docking and GPU acceleration for a further 10- to 100-fold throughput increase, and demonstrated near-linear scaling on up to 5.6 million vCPUs in the AWS Cloud — which the authors describe as a new benchmark for parallelization in cloud-based drug discovery.
It found real molecules, and proved them with crystal structures
Benchmarks alone don’t validate a screening platform; wet-lab results do. The team deployed AdaptiveFlow against two disease-relevant targets: ferroptosis suppressor protein 1 (FSP1), an emerging cancer target that protects cells from iron-mediated cell death, and poly(ADP-ribose) polymerase 1 (PARP-1), the target of an established class of oncology drugs. The screens surfaced nanomolar inhibitors of both.
Critically, the hits were not left as in-silico predictions. Leveraging newly solved crystal structures of FSP1 in complex with NAD+, FAD, and coenzyme Q1, the researchers experimentally validated the discovered inhibitors and determined co-crystal structures of FSP1 bound to the small molecules — revealing binding mechanisms that were previously unknown. For a field where “the hits didn’t survive the assay” is the default punchline, structural validation of computationally discovered compounds is the strongest evidence a platform can offer.
Why an open-source bet matters
The senior research story here spans institutions: the work was led from St. Jude’s Department of Structural Biology by Christoph Gorgulla — who originally developed VirtualFlow at Harvard Medical School with Haribabu Arthanari’s group — together with collaborators at Stanford, Dana-Farber, Amazon Web Services, Enamine, and TU Berlin. The platform is released as open source, and the prepared 69-billion-compound library is freely accessible on AWS.
That choice is the announcement’s real significance. A 1,000-fold cost reduction doesn’t just save money for groups that could already afford ULVSs; it moves the technique across a threshold where an academic lab, a pediatric research hospital, or a startup can realistically run a screen of this scale on a grant. Diseases with small patient populations — the defining constraint of pediatric oncology, St. Jude’s home territory — are exactly where commercial drug development is weakest and where cheaper early-stage discovery has the most leverage.
There are honest limits. Virtual screening still produces candidates, not drugs: hits need medicinal chemistry, ADME optimization, and years of clinical validation. Active learning can also encode biases from its prescreen labels, and the 1,000-fold figure is an upper bound tied to the REAL Space library and specific targets tested. But as a demonstration that the economics of ultra-large virtual screening can be bent by machine learning rather than budget, AdaptiveFlow is one of the more consequential infrastructure papers of the year — and, starting today, anyone can download it and try.
Sources
- [1] https://www.stjude.org/media-resources/news-releases/2026-medicine-science-news/ai-informed-adaptiveflow-redefines-large-scale-cloud-computing-for-drug-discovery.html
- [2] https://pmc.ncbi.nlm.nih.gov/articles/PMC13041786/
- [3] https://pubmed.ncbi.nlm.nih.gov/41929058/
- [4] https://www.biorxiv.org/content/10.1101/2023.04.25.537981v2.full-text
- [5] https://registry.opendata.aws/vf-libraries/
- [6] https://www.stjude.org/people/g/christoph-gorgulla.html