The Language of the Cell: Biohub's $1.8 Billion Virtual Biology Bet Pulls In Google, Meta, and Washington
Meta, Google DeepMind, Isomorphic Labs, the DOE, and NIH pool $1.8 billion into Biohub's Virtual Biology Initiative to build the open datasets needed to train AI models that predict how living cells behave.
One of the strangest coalitions in modern science assembled this week: Mark Zuckerberg’s nonprofit research organization Biohub announced that Google DeepMind, Meta, drug-discovery startup Isomorphic Labs, the U.S. Department of Energy, and the National Institutes of Health are together pouring $1.8 billion into a single goal — generating the biological data needed to train AI systems that can predict how living cells behave.
The project, called the Virtual Biology Initiative, is aiming at one of AI’s hardest open problems: building models capable of simulating a cell closely enough that researchers can test ideas digitally before committing years of lab work and millions of dollars to physical experiments. Biohub describes the combined commitment as the largest coordinated effort ever to generate AI-ready biological data.
Where the $1.8 billion comes from
The money arrives from several directions at once:
- Biohub itself committed $500 million to the initiative back in April, funded by Zuckerberg and Priscilla Chan.
- Google DeepMind, Isomorphic Labs, and Meta are contributing a combined $300 million.
- The U.S. Department of Energy plans to invest more than $500 million over five years in laboratory measurement, modeling, computation, and cell research.
- The National Institutes of Health is bringing datasets and repositories developed through more than $500 million in previous federal funding.
That stacks up to roughly $1.8 billion across new funding, existing datasets, computing resources, and measurement technologies. NVIDIA is also providing computing infrastructure and technical support, though it is not among the named funders.
Why AI biology is data-starved
The bet reflects an uncomfortable truth in computational biology. AI has already transformed protein structure prediction — AlphaFold’s successors and Meta’s ESM models can fold proteins with remarkable accuracy — but predicting an entire living cell is a categorically bigger problem.
Today’s datasets contain information on hundreds of millions of cells. Biohub researchers believe genuinely useful predictive models may require data covering billions, and eventually trillions, of cells and cellular states. The initiative will collect data showing how cells respond to changes in their environment, genetic modifications, drugs, and other interventions, using techniques including spatial transcriptomics, advanced microscopy, cryo-electron tomography, and large-scale cellular screening.
“We need to capture the language of biology; we need to capture the language of the cell. And that doesn’t exist today,” Biohub Head of Science Alex Rives told Reuters.
The vision turns biological research into something closer to simulation. A scientist could ask a model how a cell might respond to a drug candidate or a gene edit, then use the answer to decide which experiments are actually worth running at the bench. “An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally,” Rives said. “The insights that come from this could unlock a far greater understanding of disease and open up completely new paths for cures.”
Timeline and open-data terms
Biohub expects the first dataset from the project in roughly a year, and says accurate predictive models could emerge within five years — an aggressive schedule for a field where “soon” usually means “in a decade.”
The data governance model is a pragmatic compromise. Datasets are expected to become public, but the private companies funding some of the work will receive temporary early-access windows before broader release. Government-funded data, by contrast, will not carry those restrictions. For AI labs locked in a race for proprietary training data, a few months of head start on billions of cellular measurements is a meaningful carrot — and for the research community, eventual open release means the data ultimately compounds for everyone.
An unusually broad coalition
The partner list reads like a who’s-who of biology and AI. Beyond the funders, research organizations involved include the Broad Institute, the Allen Institute, the Human Cell Atlas, the Human Protein Atlas, the Gladstone Institutes, and the Wellcome Sanger Institute.
The initiative also lands amid a broader wave of AI investment in biology. Anthropic has built out wet-lab capabilities, and the OpenAI Foundation has launched a grant program worth more than $125 million aimed at biological and medical datasets. In other words, Biohub’s project is chasing a bottleneck that several frontier labs have independently identified: AI cannot learn biological behavior that scientists have never measured.
The real bottleneck isn’t compute
That framing is the most interesting part of the announcement. For most of the past three years, the limiting reagent in AI has been compute — GPUs, power, and capital. Biology inverts that equation. No amount of parameter scaling teaches a model how a hepatocyte responds to a novel compound if no one has ever measured that response at sufficient resolution and scale.
If the Virtual Biology Initiative succeeds, the next major AI breakthrough may come less from building a larger model and more from giving existing models a vastly richer picture of how life actually works. It is a wager that the “world models” that matter next won’t be trained on video of the physical world, but on painstakingly measured interiors of cells — and that the organization that funds the measurement, not just the training run, gets to shape what follows.
There are risks, of course. Coordinated mega-science projects have a mixed track record, five-year timelines in biology have a habit of slipping, and $1.8 billion is small change against the true cost of mapping trillions of cellular states. But as a signal of where the frontier labs believe the next wall sits — data, not compute — this week’s announcement is about as clear as they come.
Sources
- [1] https://www.reuters.com/business/healthcare-pharmaceuticals/us-government-google-join-zuckerberg-backed-biohub-18-billion-push-ai-biology-2026-10-07/
- [2] https://techstartups.com/2026/10/07/zuckerbergs-biohub-lands-google-and-u-s-government-backing-for-1-8-billion-ai-biology-push/
- [3] https://biohub.org/news/virtual-biology-initiative/
- [4] https://qz.com/google-meta-us-government-biohub-virtual-biology-initiative-100726