AI Maps More Than 2,800 Viral Protein Complexes to Help Prepare for Future Pandemics
When COVID-19 emerged, scientists had a crucial advantage: decades of prior coronavirus research had revealed enough about the virus’s key proteins to support vaccine design in record time. The next pandemic may not offer the same head start.
To help close that knowledge gap, NVIDIA has joined a coalition of global research institutions, including Google DeepMind and the European Bioinformatics Institute of the European Molecular Biology Laboratory (EMBL-EBI), to publish predicted 3D structures of more than 2,800 viral protein complexes.
The dataset is freely available to scientists worldwide through the AlphaFold Database. Its structures were predicted using AlphaFold2, Google DeepMind’s artificial intelligence model for predicting how proteins fold into three-dimensional shapes, alongside the NVIDIA BioNeMo Inference Runtime.
This approach enabled the team to extend protein-structure inference across thousands of viral proteomes and predict the complexes—groups of interacting proteins—encoded by each virus.
“Our goal with the AlphaFold database has always been to democratize access to basic biology at scale,” said Richa Patel, life sciences partnerships manager at Google DeepMind. “This collaboration, which brings thousands of viral complexes into the database, will give scientists around the world the insights they need to prepare for future epidemics.”
NVIDIA also contributed the BioNeMo structure prediction pipeline, a GPU-accelerated workflow that can help researchers predict the 3D structures of their own targets from protein sequences.
Preparing for the next pandemic must start now. According to an analysis by the Center for Global Development, there is approximately a 50% probability that, by 2050, the world will face a pandemic as severe as COVID-19.
“When the next pandemic happens, something may suddenly happen and we will lack the knowledge we had about coronavirus,” said Joe Grove, professor of molecular virology at the Medical Research Council/University of Glasgow Centre for Virus Research and a collaborator on the project. “What we’re trying to do is build up some of the knowledge up front.”
Approximately 30% of the protein interactions added to the database are completely new to science. They exhibit interaction geometries that have never previously been documented in the Protein Data Bank, a major repository of experimentally determined protein structures. These findings offer new biological insights for researchers to investigate and build on.
“This database is a hypothesis-generation engine,” said Chris Dallago, applied research science team leader for digital biology at NVIDIA. “By enabling biologists and the AI community to study protein interactions as complexes rather than just single molecules, we can advance the field as a whole.”
Why Viral Protein Complexes Matter
Most proteins do not function alone. Instead, they form molecular complexes that carry out sophisticated biological functions. These structures are often the targets that vaccines and drugs must interact with to disrupt viral activity.
For example, understanding the 3D structure of the COVID-19 virus’s spike protein helped support vaccine design. For thousands of other viruses, comparable structural knowledge does not yet exist. This new dataset begins to fill that gap.
Traditional methods for determining protein structures, such as crystallizing a protein and exposing it to X-rays, can take years and cost thousands of dollars per structure. Using AlphaFold2 and NVIDIA BioNeMo on NVIDIA GPUs, researchers can predict structures in minutes and run predictions in bulk. Scientists can then evaluate promising predictions through experimental methods.
For this project, the team systematically examined protein structures from virus families known to infect humans, ranging from common cold viruses to emerging threats such as mpox.
Global Collaboration and Open Access
The collaboration brings together the Epidemic Preparedness Innovations Coalition, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow.
The dataset’s release coincides with the United Nations General Assembly convened by the World Economic Forum on pandemic prevention, preparedness and response in New York City this week. It contributes to the AlphaFold Database, which currently contains more than 260 million predicted protein and protein-complex structures, covering nearly every cataloged protein known to science.
“Opening this data is critical to understanding virus diagnostics and developing treatments and vaccines,” said Joe McEntire, interim director of EMBL-EBI. “This dataset also covers less-studied viruses, lowering the barrier for scientists directly facing outbreaks in resource-strapped settings.”
Predictions in open datasets are labeled by confidence. The structures provide information about what viral complexes may look like and how individual proteins interact within a viral proteome.
“When I started my Ph.D., the proteins I was studying had no structure. It was like working in the dark and having to guess what was going on,” Grove said. “This dataset is a powerful tool for all current Ph.D. researchers, providing high-quality structural data to accelerate basic science.”
Overall, the release expands the structural information available to scientists working in digital biology, virology and disease research. By making predicted viral protein complexes openly available, the collaboration aims to give researchers a stronger foundation for studying emerging pathogens and preparing for future outbreaks.
Explore the viral protein complex datasets in the AlphaFold Database Pandemic Preparedness Portal. Use the BioNeMo structure prediction pipeline to predict protein targets and learn more about NVIDIA BioNeMo.
Source: blogs.nvidia.com


