Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

NVIDIA, Google DeepMind and Partners Release Open Dataset of Viral Protein Complexes for 2,800+ Viruses

NVIDIA has joined Google DeepMind, EMBL-EBI and other research organizations to release predicted 3D structures for the protein complexes of more than 2,800 viruses through the AlphaFold Database, free for any scientist to use. The structures were generated with AlphaFold2 optimized by NVIDIA's BioNeMo Inference Runtime, and about 30% of the protein interactions are entirely new to science.

Published
英伟达联合谷歌 DeepMind 等机构,开源 2800 余种病毒的蛋白复合物结构
Image source: blogs.nvidia.com

NVIDIA has joined a coalition of global research organizations, including Google DeepMind and the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), to release predicted 3D structures for the protein complexes of more than 2,800 viruses. The data is openly available to any scientist, anywhere, through the AlphaFold Database, with the aim of stockpiling structural biology knowledge before the next pandemic rather than during it.

The structures were inferred using AlphaFold2, Google DeepMind's model for predicting how proteins fold into 3D shapes, with optimization from the NVIDIA BioNeMo Inference Runtime. That combination allowed the team to scale inference to thousands of viral proteomes, predicting the complexes, or groups of interacting proteins, encoded within each virus.

NVIDIA is also openly releasing the BioNeMo Structure Prediction Pipeline, the GPU-accelerated workflow used to generate the dataset. Researchers can apply the same sequence-to-structure process to their own targets instead of rebuilding the inference chain from scratch.

About 30% of the protein interactions being added to the database are completely new to science, showing interaction shapes that have never been documented in the Protein Data Bank, the main repository of experimentally determined protein structures. That translates into new material for the biological community to explore and build on.

Structure matters because most proteins do not work alone: they come together in complexes of multiple molecules to perform sophisticated functions, and those structures are often what a vaccine or drug must target to disrupt viral function. Understanding the 3D structure of the COVID-19 spike protein proved foundational to vaccine design; for thousands of other viruses, no such structural knowledge exists today, and the dataset begins to fill that gap.

Traditional methods for determining protein structures, crystallizing proteins and shooting X-rays at them, can take years and cost thousands of dollars per structure. AlphaFold2 optimized to run on NVIDIA GPUs predicts a structure in minutes and can be run in bulk, after which scientists can verify high-confidence predictions through experimental methods. For this project the team systematically worked through viral families known to infect humans, from common-cold viruses to emerging threats such as Mpox.

The collaboration spans the Coalition for Epidemic Preparedness Innovations, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow. The release coincides with a United Nations General Assembly meeting convened by the World Economic Forum on pandemic prevention, preparedness and response taking place this week in New York City, and the AlphaFold Database now holds more than 260 million protein and protein complex predictions covering nearly every cataloged protein known to science.

Those involved framed the timing around urgency. Chris Dallago, applied research science team lead in digital biology at NVIDIA, calls the database "an engine for hypothesis generation," while Joe Grove, professor of molecular virology at the University of Glasgow, says the intent is to stockpile knowledge ahead of time. An analysis by the Center for Global Development estimates a roughly 50% chance of the world facing a pandemic as severe as COVID-19 by 2050, which is why open structural data may matter as much for preparedness as for research speed.

Why it matters

Releasing both the dataset and the generation pipeline turns a years-long, thousands-of-dollars-per-structure effort into bulk prediction, lowering the barrier for research teams in low-resource settings. For pandemic preparedness it builds a searchable structural baseline before an outbreak rather than after, which could shorten the path from identifying a pathogen to choosing vaccine and drug targets.

NVIDIAGoogle DeepMindOpen SourceLife Sciences
Back to realtime news

Nearby Updates

All