Google DeepMind’s AlphaGenome Atlas Maps the Impact of 9 Billion DNA Changes
Google DeepMind has introduced a new artificial intelligence (AI) tool that predicts how changes to individual DNA letters could affect human tissues and cellular processes.
Announced on September 8, the AlphaGenome Atlas contains data on approximately 9 billion possible changes to the human genetic code. The database estimates how each change could influence different tissues and biological processes, and provides a simple score to help researchers assess the potential impact of genetic variants.
What is the AlphaGenome Atlas?
The human genome contains around 3 billion DNA base pairs. Understanding how this enormous sequence is packaged, distributed and read to create human life remains one of the central challenges of genetics research.
Last year, DeepMind launched AlphaGenome, an AI model designed to predict how changes in DNA sequences could affect proteins, cells and other biological processes. The AlphaGenome Atlas builds on that work by making the model’s predictions available in a searchable database.
DeepMind hopes the Atlas will make AlphaGenome more accessible to researchers who may not have the bioinformatics expertise or computing resources needed to run the model themselves.
Why non-coding DNA matters
About 98% of human DNA is non-coding, meaning it does not directly provide instructions for making proteins. For many years, this portion of the genome was sometimes described as “junk DNA.” Geneticists now understand that non-coding regions can contain regulatory instructions that control when and where protein-coding genes are activated.
These regulatory instructions are distributed throughout the genome rather than arranged in a simple sequence. As a result, a change in one DNA letter can potentially affect regulatory activity across a much larger region.
Using AlphaGenome, researchers can investigate how a single DNA-letter change may affect up to one million nearby DNA letters. This gives the model a broader view of each gene’s regulatory network than previous approaches.
Making genetic variant research easier
Before the AlphaGenome Atlas was released, researchers generally needed bioinformatics experience to access AlphaGenome through its application programming interface. Calculating the predicted effect of every genetic change could also place a significant burden on academic computing resources.
The Atlas addresses both challenges by precomputing the effects of every possible DNA base change. The resulting dataset totals about 1 petabyte, or 1 million gigabytes, and is available free of charge through an accessible web portal.
“You don’t have to be an expert in these methods to go in there and look things up in your browser.”
Tuuli Lappalainen, Professor of Genomics at KTH Royal Institute of Technology and senior associate faculty member at the New York Genome Center
Greg Findlay, a group leader at the Francis Crick Institute in London who was not involved with the AlphaGenome Atlas, described it as a useful resource. However, other researchers have emphasized that the technology cannot answer every major question in genetics.
What is the AlphaGenome Variant Impact score?
The AlphaGenome Atlas also includes an AlphaGenome Variant Impact score. This metric predicts how much biological impact a particular genetic change could have.
In a preprint paper, DeepMind researchers showed that the score could distinguish between disease-associated and harmless variants in clinical datasets.
Lappalainen said the score’s simplicity could make some variants difficult to interpret. Nevertheless, it may be useful for researchers who need to quickly examine large lists of genetic variants and identify those that warrant further study.
Why AlphaGenome is not a perfect predictor
Although the AlphaGenome Atlas could improve the study of genetic mutations, it is not a complete or infallible predictive tool.
A September 11 preprint study from a research team led by Katie Pollard, director of the Gladstone Institute for Data Science and Biotechnology and a professor at the University of California, San Francisco, found that AlphaGenome was effective at identifying potentially causative mutations but consistently underestimated their impact.
Pollard’s team also found that the model could not always connect changes in regulatory elements with the genes they control, particularly when those genes were located far away in the genome.
“My perhaps naive wish is that people take these predictions with a grain of salt.”
Tuuli Lappalainen
AI predictions still need laboratory testing
Despite these limitations, Lappalainen said tools such as the AlphaGenome Atlas are part of a broader shift toward collaborative, data-driven genomics. She also cautioned that researchers must place equal emphasis on laboratory, or “wet lab,” experiments to test whether predicted genetic effects are accurate.
Laboratory testing is slower and more labor-intensive than searching a database, but it is essential. Experimental results provide the biological data that AI models need to improve their predictions.
“Data in biology is still very limited and we need to generate that data,” Lappalainen said.
Source: www.livescience.com


