More than 3 billion “letters” make up the human genome. Science can already read nearly all of them, but reading does not necessarily mean understanding.
Less than 2% of the genome directly codes for proteins, while most of it consists of noncoding regions, including regulatory mechanisms that determine when genes are activated, how strongly they are expressed and in which cells. Many of these regions have already been studied, but our understanding of their roles remains significantly limited. Yet it is precisely there that explanations may lie for diseases, differences between individuals and even why a particular treatment works for one person but not another.
An ambitious project led by ARC, the innovation arm of Sheba Medical Center, together with the Icahn School of Medicine at Mount Sinai in New York and NVIDIA, is entering this vast territory.
Under a three-year collaboration involving a total investment of tens of millions of dollars, the three partners are developing a Genomic Foundation Model, or gFM, designed to use artificial intelligence to identify patterns in DNA sequences and link them to biological mechanisms, diseases and responses to treatment.
The concept is somewhat similar to the way large language models learn from text. Instead of words and sentences, the model is exposed to enormous quantities of DNA sequences and attempts to learn the “rules” hidden within them: which sequences appear together, which variations may disrupt regulatory mechanisms and how a genetic difference relates to the activity of a gene.
But in the human body, that language is particularly complex. A small change in one region can affect a gene located far away, and the same sequence may behave differently in a brain cell, the liver or the immune system.
That is where the partnership offers an advantage. Sheba and Mount Sinai bring genetic databases, medical information and clinical expertise, while NVIDIA provides computing power, AI infrastructure and expertise in building large-scale models.
The aim is not merely to teach a computer to identify interesting sequences but to connect its predictions with information about real people: who developed a disease, how the disease manifested itself and how the patient responded to treatment.
The project will initially focus on mental health and neuropsychiatric disorders, where genetics is particularly complex. In diseases such as schizophrenia, for example, there is generally no single mutation that explains the condition. Instead, risk is associated with a combination of many variations across the genome, each contributing a small amount.
A model capable of examining thousands of regions simultaneously and identifying relationships among them could help researchers determine which variations truly matter and which biological pathways warrant further investigation.
The project is also not starting from scratch. The field of genomic AI has developed rapidly in recent years, with models including Evo 2 from the Arc Institute and NVIDIA and AlphaGenome from Google DeepMind. The innovation here, therefore, is not simply the use of AI to read DNA, but the effort to directly connect an advanced genomic model with large medical databases and questions emerging from clinical practice.
Sheba and NVIDIA have also begun laying the groundwork. During 2026, researchers involved in the partnership published initial studies aimed at improving how genomic models learn from DNA sequences and how such models can be compared more accurately. These remain early steps toward the broader objective, but they mark a transition from an ambitious announcement to actual research and development.
And the goal is particularly ambitious.
The project is not attempting to produce, within three years, a “dictionary” explaining every letter in the 98% of the genome that does not code for proteins. Rather, it aims to find patterns and connections within that enormous landscape that existing tools struggle to identify.
If AI can successfully connect genetic variations with regulatory mechanisms, biological pathways and diseases, it could give researchers access to regions of the genome whose medical significance remains far from fully understood.
Since the completion of the Human Genome Project, our ability to determine what is written in DNA has improved dramatically. Sheba, Mount Sinai and NVIDIA are now trying to tackle the next, far more complicated challenge: understanding what it all means.
If their model succeeds in connecting genetic variations, biological mechanisms and disease, it could transform vast portions of the genome from information we know how to read into information that can be used to detect diseases earlier, identify new targets for treatment and tailor precision medicine more closely to each individual.




