Nine Billion Variants, One Petabyte: DeepMind's AlphaGenome Atlas Precomputes Every Possible DNA Letter Change
Google DeepMind has released AlphaGenome Atlas, a 1-petabyte dataset predicting the molecular effects of all ~9 billion possible single-letter DNA variants in the human genome — 30x larger than the AlphaFold Database — free for academic use.
Nine billion. That is the number of single-letter changes — single-nucleotide variants, in the language of genetics — that are possible across the roughly three billion base pairs of the human genome. Testing even a tiny fraction of them in a laboratory is practically impossible. On September 8, 2026, Google DeepMind removed the need to try: the company released AlphaGenome Atlas, a platform containing precomputed predictions for the molecular effects of every single one of those variants, packaged as a 1-petabyte dataset that any academic researcher can query through a free web portal.
The release is the latest escalation in DeepMind’s strategy of turning frontier AI models into open scientific infrastructure — the same playbook that made the AlphaFold Database one of the most-used resources in the history of biology. AlphaGenome Atlas is more than 30 times larger than the AlphaFold Database, and DeepMind is explicitly framing it as the genomic counterpart to the protein-structure resource that earned a share of the 2024 Nobel Prize in Chemistry.
From model to map
AlphaGenome, the underlying AI model, was introduced in June 2025 as a unified DNA sequence model capable of predicting how genetic variants affect thousands of molecular processes — gene expression, RNA splicing, chromatin accessibility and more — across hundreds of human and mouse cell types. The model proved powerful for analyzing individual variants, but it left a practical gap: most biologists do not want to run inference on a frontier AI system. They want an atlas they can browse.
By precomputing AlphaGenome’s predictions across the entire genome, DeepMind has created exactly that. Where an atlas links together features of the land like altitude and location, AlphaGenome Atlas charts the molecular effects of DNA variants across all of the genome’s terrain — including the 98% that does not code for proteins, where most disease-associated variants actually reside and which conventional tools have always struggled to interpret.
Alongside the raw predictions, the team released the AlphaGenome Variant Impact (AVI) score, a single number that condenses the strengths of AlphaGenome and AlphaMissense (DeepMind’s model for protein-altering variants) into a unified impact ranking. The AVI score works for both coding and non-coding regions, and DeepMind reports best-in-class performance across variant pathogenicity and rare disease benchmarks. Each AVI score is further decomposed into feature attributions — which molecular process, such as RNA splicing or chromatin accessibility, is predicted to be most disrupted — so researchers get a mechanism, not just a number.
The Atlas also ships with a compendium of more than 2,500 recurrent DNA sequence motifs — the “words” of the genome — and their locations, letting scientists connect variants directly to the functional sequences they disrupt.
Real results, already validated
The most persuasive part of the release is not the dataset’s size but the validation work DeepMind’s external collaborators have already done with it.
At the Broad Institute, Laura Covill and Anne O’Donnell-Luria, working with the GREGoR Consortium on unsolved rare disease cases, used the AVI score to prioritize variants that previous research had overlooked. The team identified a variant affecting DNM1, a gene strongly linked to epileptic encephalopathy. Crucially, AlphaGenome’s predictions explained exactly how the variant caused harm: it created an incorrect splice site, a mistake in the cell’s genetic instructions that produces an abnormally extended protein. Experimental screens validated the prediction and turned up nearby variants with similar effects.
At the University of Exeter, Medical Research Council fellow Gareth Hawkes applied the Atlas to whole-genome data from more than 54,000 UK Biobank participants. By grouping rare variants according to their predicted molecular effects, he uncovered 22% more non-coding genetic associations than standard approaches — signals that would otherwise have been lost in the statistical noise of millions of harmless changes. The work pinpointed regulatory variants driving the abundance of circulating proteins including PLA2G7 (linked to aging) and EGLN1 (a vital cellular oxygen sensor), and a BMI-focused analysis identified 19 genetic regions for targeted follow-up.
At the Stowers Institute for Medical Research, Julia Zeitlinger and Melanie Weilert used the motif compendium to categorize which transcription factors merely affect DNA accessibility versus which ones can actually turn genes on and off — a basic-science question that previously required years of dedicated experiments.
Why it matters
The economics of genetic interpretation change when predictions are precomputed. A clinician facing a sick child with an unexplained phenotype must today sift through thousands of candidate variants; a ranked, mechanistically annotated shortlist can compress months of analysis into an afternoon. Population geneticists studying common traits face a different version of the same problem — finding the few functional variants hidden in a background of millions — and the Exeter results suggest the Atlas measurably improves that signal-to-noise ratio.
There are caveats worth stating plainly. The predictions are model outputs, not experimental measurements, and they inherit the limitations of AlphaGenome itself. DeepMind describes the Atlas as “a baseline rather than an endpoint” — as the underlying model improves, the maps will need to be regenerated. Access is free for non-commercial use through the web portal, with commercial licensing coming to Google Cloud, a two-tier arrangement that mirrors AlphaFold’s and will likely renew debates about how much of the life-science commons should sit inside a single company’s infrastructure.
Still, the trajectory is unmistakable. AlphaFold compressed fifty years of expected protein-structure determination into a searchable database; AlphaGenome Atlas now attempts the same move for the regulatory genome, the portion of our DNA that has resisted systematic interpretation since it was first sequenced. And DeepMind is already positioning the Atlas as a component in larger systems: the resource is accessible through the AlphaGenome API and as a skill in Google Antigravity, the company’s agentic development platform, pointing toward a future where an AI agent can chain variant ranking, mechanism attribution and literature review into a single scientific workflow.
For the researchers who spent decades mapping the genome letter by letter, the release of a precomputed atlas of every possible change to it is a fitting next chapter — the reference genome told us what we are made of; AlphaGenome Atlas begins to tell us what each letter of it does.
Sources
- [1] https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/
- [2] https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphagenome-atlas/
- [3] https://spectrum.ieee.org/alphagenome-atlas
- [4] https://aiweekly.co/alerts/google-deepmind-ships-alphagenome-atlas-with-predictions-for-all-9b-single