Google DeepMind introduced AlphaGenome Atlas on 8 September 2026, describing it as a searchable database of predicted molecular effects for approximately 9 billion possible single-nucleotide changes in the human genome. The company says the pre-calculated dataset is about 1 petabyte in size and can be queried through a web portal without coding.

The Atlas is built around AlphaGenome, an artificial-intelligence model that predicts how DNA changes may affect molecular processes. It also provides an AlphaGenome Variant Impact (AVI) score that combines predicted effects from both coding and non-coding regions to help researchers prioritise variants for further investigation.

The resource is presented as a research and variant-prioritisation tool, not as a clinical diagnostic or treatment product. The announcement does not provide independent validation, accuracy metrics or regulatory approval.

Contents

What AlphaGenome Atlas contains

A single-nucleotide variant is a change involving one DNA base at a particular position. For each relevant position, different single-letter substitutions can produce different potential molecular consequences.

Google DeepMind says it used AlphaGenome to pre-calculate predicted effects for approximately 9 billion such changes and organised the results into the Atlas. Pre-calculation changes how researchers access the predictions: instead of generating a model output separately for every candidate variant, users can query a prepared catalogue.

The stated scale is central to the resource's purpose. The resulting dataset is approximately 1 petabyte, according to Google. The supplied announcement does not clarify whether the 9 billion figure covers every possible substitution across a particular reference-genome scope or uses another defined set of positions.

The Atlas includes predictions for coding and non-coding DNA. Coding sequences contribute directly to protein production. Non-coding sequences do not encode a protein themselves but can influence processes such as gene regulation and RNA processing. The source says humans have about 3 billion DNA base pairs, while roughly 2% codes for proteins, leaving a much larger non-coding portion whose effects are less completely understood.

The AVI score is intended to combine predicted effects from these two broad regions into one prioritisation measure. That does not make it a measurement of biological impact. It remains a model-derived hypothesis about what a variant might do.

How the Atlas is meant to be used

The Atlas is positioned as infrastructure for researchers studying many candidate variants. A researcher could use the AVI score to rank variants and decide which ones deserve laboratory, statistical or clinical follow-up.

This may be particularly relevant to rare-disease research, where a patient's genome can contain many variants and the disease-causing change may not lie inside a protein-coding sequence. A non-coding variant can affect gene regulation or RNA processing even when it does not alter a protein's amino-acid sequence.

The same approach can be applied to population-scale genetic analysis. Rather than treating variants only as locations in the genome, researchers can group them according to their predicted molecular effects and then examine whether those groups are associated with traits.

Google describes the Atlas as available through a website portal intended for users without coding skills. The announcement does not specify access conditions, usage limits, licensing terms, downloadable data or whether an application programming interface is available.

Reported research applications

Google says researchers at the Broad Institute used the AVI score in an unsolved rare-disease case involving the DNM1 gene. In that example, AlphaGenome predicted that a variant created an incorrect splice site.

A splice site is a sequence involved in RNA processing. Cells use splice sites to remove introns and join exons before producing mature RNA. An incorrectly used splice site can change the resulting RNA and potentially the protein produced.

Google characterises the DNM1 prediction as supporting evidence that helped solve the case. It does not say that the Atlas alone established the diagnosis, and the supplied announcement does not state whether the prediction was experimentally validated or how much it contributed compared with other evidence.

The company also says the Atlas was applied to data from more than 54,000 UK Biobank participants in an analysis of complex traits. According to Google, grouping variants by predicted molecular effects uncovered 22% more non-coding genetic associations. When researchers focused on the top 1% of variants predicted to have the greatest impact, the analysis identified 19 genetic regions linked to body mass index.

These findings describe statistical associations, not proof that the identified variants or regions cause changes in body mass index. The announcement does not provide the analysis's exact statistical methods, covariate adjustments, association definitions or replication results.

What to watch next

Independent evaluations will need to test whether AlphaGenome's predictions agree with experimentally measured regulatory effects across different biological contexts. Such work could also show how the AVI score should be calibrated and how much it improves variant prioritisation compared with existing genomic annotations and prediction methods.

Follow-up studies will be important for determining whether variants prioritised by the Atlas produce reproducible biological or clinical findings. For the UK Biobank application, replication and causal analyses would help distinguish useful biological signals from statistical associations.

The practical value of the Atlas will also depend on access. Google says the portal is available globally and does not require coding, but the supplied announcement does not detail data-download options, API access, licensing or usage limits. India-specific pricing, availability and regulatory information have not been provided.

Sources