← All posts / Industry

Data as the New Drug: OpenAI Foundation Commits $125M+ to Public Data for Health, With $40M to UNC for Cancer Vaccines

The OpenAI Foundation's second science program funds open biological datasets — a $40M grant to UNC Lineberger will directly measure tumor antigens and T-cell responses to make personalized cancer vaccines less of a guessing game.

Data as the New Drug: OpenAI Foundation Commits $125M+ to Public Data for Health, With $40M to UNC for Cancer Vaccines

On the morning of September 15, 2026, the OpenAI Foundation announced its second science program, Public Data for Health, committing more than $125 million in initial grants to create and preserve open-access scientific datasets. The headline grant — $40 million to UNC Lineberger Comprehensive Cancer Center in Chapel Hill — aims to solve a deceptively mundane problem that has quietly stalled one of oncology’s most promising frontiers: personalized cancer vaccines are designed computationally from data nobody has actually measured.

The announcement

The program, unveiled in a post authored by Abhishaike Mahajan and Jacob Trefethen of the Foundation’s Life Sciences and Curing Diseases division, takes a different shape from the Foundation’s first science program, AI for Alzheimer’s, launched in April 2026. Where that effort targets a single devastating disease, Public Data for Health is deliberately horizontal: it funds the creation of datasets that researchers and AI tools worldwide can build on, across many diseases at once.

The Foundation’s framing is blunt. Scientific data, it argues, is the foundational input to discovery. In mathematics, AI systems have recently begun contributing new knowledge without collecting new data. Biology is the opposite: models can increasingly analyze biological information at scale and recover hidden structure even from incomplete evidence, but they cannot conjure observations that were never made. And some datasets of enormous public value may never exist at all — because no single company or academic lab has the incentive or the budget to create them.

That market-failure argument is the program’s thesis: philanthropy should fund the data layer of the life sciences the way it once funded reference genomes and protein structure archives.

Three grants, three theories of data

The initial tranche spans molecular, epidemiological, and regulatory layers of drug development, each embodying one of the Foundation’s three “starting hypotheses” about what makes data valuable — connected data that follows biology across multiple steps, scarce data that saves what cannot be recreated, and direct data that measures what actually matters.

OpenADMET (UCSF and collaborators) will build open datasets, benchmarks, and blinded competitions to test whether AI can predict how small molecules are absorbed and distributed through the body. The stakes are spelled out in the announcement: roughly 90% of drug candidates fail in clinical trials, often because absorption and distribution are hard to predict. The team cites AlphaFold2’s reliance on Protein Data Bank data during CASP competitions as the precedent — high-quality public data paired with public prediction challenges is what unlocked protein structure prediction. OpenADMET wants to repeat that playbook for drug pharmacokinetics, measuring key molecular properties for tens of thousands of compounds, linking them to transporter-protein structures, and validating a subset against human blood-brain barrier models.

CTD Commons, led by Josh Morrison of 1Day Sooner, will test whether Common Technical Documents from failed or shelved drug programs — the full regulatory dossiers spanning animal toxicology, manufacturing details, and FDA correspondence — can be acquired and published before companies shut down and the records disappear. Only a sliver of that work ever appears in published papers. If it works, early-stage developers get a map of what regulators have required of similar products, and machine learning systems get a corpus from which to learn patterns in why drugs fail.

The University of North Carolina, the largest grant at $40 million, will establish the Initiative for Generative Immunotherapy under immunologist Dr. Benjamin Vincent and computational biologist Dr. Alex Rubinsteyn.

The UNC project: measuring the red flags

Personalized neoantigen vaccines are among the few medicines designed computationally for each individual patient. The workflow: sequence a patient’s tumor, predict which tumor-specific proteins (neoantigens) appear on its surface, and build a vaccine that trains the immune system to attack cells waving those particular red flags.

The catch is that tumor sequencing is a proxy. Directly measuring which targets actually appear on the cell surface — and how strongly that patient’s T cells respond to each one — is difficult and expensive, so almost nobody does it at scale. Vaccines end up built on predictions stacked on predictions, which is why, as Rubinsteyn put it, “many fail in the development process because they are too arbitrary.”

The UNC team will attack this directly. They plan to analyze hundreds of de-identified tumor tissue and immune cell samples drawn from three biobanks, directly measuring tumor cells’ surface proteins and patients’ T-cell responses rather than inferring them from sequencing. The result will be one of the first real training and evaluation datasets in the field — data that AI models can use to learn which tumor antigens are genuinely good targets, not just predicted ones. In parallel, clinical trials will compare multiple vaccine formulations for their ability to raise tumor-specific immune responses in triple-negative breast cancer, one of the hardest-to-treat subtypes.

“Personalized cancer vaccines are finally starting to show signs of clinical efficacy, but still have gaps which might take decades to fill under the traditional model of therapeutic development,” Rubinsteyn said. “High-quality data can help us close those gaps faster.”

All of it will be released as de-identified public data. Jacob Trefethen, the Foundation’s Head of Life Sciences and Curing Diseases, was explicit about the point: UNC’s work “will generate data that researchers don’t have at the scale they need today and make those findings broadly available so scientists around the world can build on them.”

Why this matters

The timing is notable. Just two weeks ago, the field’s first Phase 3 win for a personalized mRNA cancer vaccine — built on a neoantigen-picking algorithm — demonstrated that computational vaccine design can clear the highest bar in clinical medicine. But one successful trial design doesn’t generalize, because the underlying target-selection models are trained on thin, indirect data. The UNC dataset is an attempt to fix the field’s data problem at the root rather than tuning models on top of it.

There is also a structural signal here. The OpenAI Foundation — the non-profit parent that governs the for-profit OpenAI Group PBC — is spending its scientific credibility on open public goods rather than exclusive partnerships. Grantees are encouraged to publish analyses as preprints and share data continuously rather than only at project end. If the Foundation’s bet is right, the highest-leverage thing an AI-era philanthropist can buy isn’t compute or models — it’s the observations the models have never been able to learn from.

For patients, the path from dataset to drug is long. But if the Initiative for Generative Immunotherapy delivers what it promises, the next generation of vaccine candidates could enter human dosing backed by a far stronger set of targets — and the design of personalized cancer vaccines, in Vincent’s words, becomes less about guesswork and more about measurement.