The molecular effects of mutating every DNA base pair in the human genome one by one have been mapped in the AlphaGenome Atlas. The deep learning algorithm from Google modelled changing each of the three billion letters of DNA one at a time and scored each change according to its predicted impact.
This includes what was once labelled junk DNA – sequences that do not code for proteins but are now recognised as crucial in turning genes on or off or regulating their activity. Less than 2% of the human genome codes for proteins, with the non-coding ‘dark genome’ making up the rest. The consequences of variation in this dark genome are less well understood. Now, a geneticist can look up how changing a DNA base anywhere in the genome influences a gene.
‘Just as a topographical map charts height at different geographical locations, the Atlas charts the impact of mutations at different locations in the genome,’ says Žiga Avec, a computational biologist at Google DeepMind, which created the Atlas using its AlphaGenome AI model. ‘It’s a collection of maps because there is a separate map for every cell type and every regulatory process.’
What do people mean when they talk about AI in science?
Artificial intelligence (AI) is an umbrella term often incorrectly used to encompass a variety of connected but simpler processes.
AI is the ability of machines and computer programmes to perform tasks that typically only humans could do, such as reasoning, responding to feedback and decision making.
Generative AI is a newer variant of AI that analyses and detects patterns in training datasets to generate original text, images and videos in response to requests from users. ChatGPT, Microsoft Copilot, Google Gemini and more recently X’s Grok are all examples of chatbots that use generative AI.
Neural networks are an interconnected array of artificial neurons, akin to biological brains, that identify, analyse and learn from statistical patterns in data.
Machine learning is a subset of AI that allows machines to learn from datasets and make predictions based on new data, without programmers explicitly asking it to do so. Machine learning models improve their performance as they receive more data.
Deep learning is an enhanced type of machine learning that uses neural networks with many layers to analyse complex data from very large datasets. Applications of deep learning include speech recognition, image generation and translation.
Large language models or LLMs are a type of deep learning trained on large amounts of data to understand and generate language. LLMs learn patterns in text by predicting the next word in the sequence and these models are now able to write prose, analyse text from the internet and hold dialogues with users.
Testing all single-nucleotide variants in a lab is virtually impossible, but the Google DeepMind team used AlphaGenome to make predictions on the impact of genetic changes. ‘They took datasets from all sorts of experiments in different tissues and then tried to predict the effect of a variant using those across the entire genome,’ says Caroline Wright, a geneticist at the University of Exeter, UK, who contributed to the recent Atlas preprint.
AlphaGenome, which was unveiled in January, is able to predict which genes are expressed in different tissues, where they get spliced, or which DNA letters are accessible, physically close to one another or bound by proteins. The Atlas applied insights from this AI tool throughout the human genome for each of the three alternatives across three billion bases, as well as millions of additions and deletions observed in biobanks. It then generated a variant score to rank mutations by impact. ‘It gives you a ranked list for what the biggest predicted impacts are, but you can also pick out a cell type that you are interested in,’ says Carl de Boer, a genomics researcher at the University of British Columbia, Canada.
Shedding new light on a rare disease
The Atlas revealed rare non-coding variants that drive circulating protein levels. It also used its AlphaGenome Variant Impact to implicate a variant in a gene (DNM1) in a rare disease: epileptic encephalopathy. It predicted that this variant caused an error that leads to an abnormally extended protein – subsequent experiments supported the prediction.
‘If you find a gene or region associated with whatever disease you are interested in, you can then dig into it in more detail,’ says Wright. ‘It could suggest that increasing or decreasing a protein could impact a disease, which could be useful for finding new drug targets.’
‘The genome is littered with these switches called enhancers that turn on genes, off genes, and you can have variants there that influence diseases,’ says Jorge Ferrer at the Centre for Genomic Regulation (CRG) in Barcelona, Spain, who analyses genomes from patients with diabetes. ‘The most common diseases are largely influenced by variants that act not in protein coding genes, but on these switches and enhancers. They might have a tiny effect, but if you aggregate the effect of many variants from different parts of the genome, it can strongly influence disease susceptibility.’ These include heart disease, lipid disorders, diabetes, obesity and Alzheimer’s disease. The Atlas could help pinpoint the most important variants.
Variants don’t exist in isolation, however, and can be influenced by other variants or masked by redundancy. And there are differences between our own cells and in our life histories. ‘We’re all a bit different in terms of what our cells are doing, what we had for lunch, the last virus we got. These influence how cells read DNA but are not included in these predictions,’ says de Boer. ‘How much this matters is not clear.’
The developers of the Atlas stress that it would not be suitable for diagnosing patients. ‘It is somewhat limited in the cell types we cover and there are some [regulatory] mechanisms that we might not cover,’ says Jung Cheng at Google DeepMind. And it makes predictions, he adds, not determinations. ‘I would not be hanging my hat on predictions at this point,’ agrees de Boer. ‘It requires some follow-up experimentation but it could still narrow down possibilities quite a bit.’
Still, it should help researchers navigate dark DNA. ‘The non-coding part of the genome has been a tough nut to crack,’ says de Boer. ‘It’s a complicated system and we need a lot of data and very powerful models. This looks state-of-art. It works surprisingly well, though it does need lots of improvement.’






No comments yet