Deep Origin's DODock Cracks Virtual Screening: 30% Hit Rate on CD73, 100x the AI Benchmark
A physics-plus-ML docking engine held 80% pose accuracy on novel targets where AlphaFold 3-class models fall below 25% — and turned a 0.3% CD73 screen into 30.6%.
Virtual screening has been a staple of computational drug discovery for four decades, and for most of that time it has been a disappointment. Legacy pipelines are credited as the primary hit-finding strategy in only about 1% of drug discovery campaigns, largely because false positives that look great in silico routinely evaporate when they hit a wet lab. On August 5, 2026, South San Francisco–based Deep Origin published a bioRxiv preprint that claims a substantive way out of this trap — and the wet-lab numbers to back it up.
What Deep Origin Announced
The company introduced two intertwined models: DODock, which predicts how a small molecule binds to a protein, and DOScore, which ranks binding affinity across ultra-large chemical spaces. The headline result came from a prospective screen against CD73, a nucleotidase enzyme that modulates immune suppression and tumor metabolism — a well-known but hard target in cancer immunotherapy. Screening across a library of roughly 80 billion virtual compounds, Deep Origin’s pipeline delivered a 30.6% hit rate: 56 active compounds out of 183 tested, 54 of them active below 100 µM, with a representative non-nucleotide hit reaching 570 nM cellular potency.
The comparison point is stark. A prior large-scale AI screen against CD73 — the 318-target AtomNet campaign — confirmed a single weak inhibitor at 176 µM among 335 compounds, a hit rate of roughly 0.3%. Deep Origin’s result is a ~100-fold improvement in screening efficiency on the same target class.
The “Memorization Trap” — and How to Escape It
The preprint’s most intellectually interesting claim concerns why existing models fail. Deep Origin argues that standard time-based benchmarks leak look-alike proteins and twin chemical skeletons into training sets, letting AI models score high simply by memorizing historical ligand–protein pairs. When confronted with genuinely novel biology and chemistry, that memorization is worthless.
The evidence they marshal is a distinction the field hasn’t clearly drawn before: leading AI co-folding models (AlphaFold 3, Boltz-1, Chai-1, Protenix) are actually good at reconstructing protein pockets — their accuracy collapses specifically during ligand placement and docking mechanics. On the independent Runs N’ Poses benchmark, these models achieve 75–88% pose accuracy on familiar targets but drop below 25% on unfamiliar ones and novel chemical space. DiffDock falls to 2% on the OpenBind benchmark of novel-target biology, where co-folding models manage only 4–28%.
DODock’s numbers on the same tests: 89% on familiar complexes, over 50% on the most novel ones in Runs N’ Poses, and 80% on OpenBind. Under strict zero-leakage splits — Deep Origin removes any protein sharing ≥30% sequence identity and any ligand exceeding 0.4 Tanimoto similarity from training — DOScore maintains more than 10-fold early hit enrichment (EF@1%) across the majority of DUD-E and DEKOOS 2.0 targets.
Physics Grounding Is the Difference
How does it work? Instead of relying purely on patterns learned from data, DODock generates broad 3D pose proposals with a diffusion model, then refines them through an 80-parameter physical energy engine called DOFast that calculates fundamental molecular forces — not memorized associations. A final AI model ranks the refined poses by evaluating atomic contact points across the binding interface. On the PoseBusters benchmark of physical validity, DODock holds 84% → 80% accuracy from most-similar to least-similar targets, while Glide SP falls from 80% to 47%, AutoDock Vina from 64% to 50%, and DiffDock from 23% to 2%.
DOScore tackles the chronic data-scarcity problem in affinity modeling more cleverly: Deep Origin used DODock itself to generate high-quality 3D structures for hundreds of thousands of assay-labeled compounds, creating its own physically grounded training substrate for ranking true binders above decoys.
A Blind Test That’s Hard to Argue With
Perhaps the most persuasive single data point in the paper is a fully blind prospective test. DODock predicted the binding pose for laroprovstat (AZD0780), AstraZeneca’s Phase 3 oral PCSK9 inhibitor candidate, at an atypical C-terminal domain site — months before any crystal structure existed. Deep Origin then solved the 2.07 Å crystal structure experimentally, confirming the blind prediction to 1.2 Å heavy-atom RMSD. Predicting an unexpected binding mode, on a molecule you’ve never seen, at an unusual site, and being right to roughly an atom’s width — that is the kind of validation that benchmarks can’t fake.
The four prospective campaigns also followed a deliberate difficulty gradient: IRAK4 (kinase) yielded a 15.0% hit rate including a 158 nM inhibitor; Factor XIa (protease) 4.3%; and IL-17A — a protein–protein interaction target where small molecules historically struggle — 3.1%, with hits that allosterically disrupt a distal PPI site. Across all four, validated hits showed high scaffold novelty (ECFP4 Tanimoto similarity of 0.22–0.27 versus known binders), meaning the screens found genuinely new chemical matter rather than remixes of existing drugs.
Why It Matters
The economics are the real story. Roughly 90% of drugs fail in clinical development, and around 74% of new-molecular-entity cost accrues after preclinical testing — where the discovery funnel has already collapsed by a factor of 10,000×. If docking tools can hold their accuracy on novel targets, the number of molecules worth taking to the bench rises dramatically, and the cost of finding starting points for “undruggable” targets falls.
There are caveats worth stating plainly. A 30% hit rate on one target family does not guarantee generalization to all biology; the preprint has not yet passed peer review; and Deep Origin notes its internal production models already exceed the published metrics — a reminder that this is also a commercial positioning exercise for its partnership business (40 slots, 2 filled, lead out-licensing asset GPR75). But the company has done something unusually scientific for the AI-drug-discovery space: it published the whole method — every algorithm, hyperparameter, and data split across 80+ pages of supplementary documentation — and explicitly invites the community to audit and replicate the results on molecules it has never seen. Co-founder and CEO Michael Antonov (co-founder of Oculus) frames it as testing “on cases built to make it fail,” and CSO Garegin Papoian, formerly Monroe Martin Professor of biochemistry at the University of Maryland, co-led the research with Head of AI Garik Petrosyan.
For a field that has learned to live with a low accuracy ceiling for 40 years, a physically grounded model that keeps its promises on novel targets — and proves it in wet labs, blind tests, and open methodology — is exactly the kind of result that deserves attention.
Sources
- [1] https://deeporigin.com/news/virtual-screening-architecture-100-fold-hit-rate
- [2] https://www.globenewswire.com/news-release/2026/08/05/3339125/0/en/deep-origin-introduces-novel-virtual-screening-architecture-achieves-100-fold-hit-rate-jump-on-challenging-target.html
- [3] https://www.hpcwire.com/aiwire/2026/08/19/deep-origin-claims-breakthrough-in-ai-drug-discovery-platform/
- [4] https://www.synbiobeta.com/read/deep-origin-unveils-breakthrough-in-virtual-screening-with-dodock