Next upAI x Bio Pitch Contest
News

NVIDIA joins coalition releasing an open viral-protein dataset

NVIDIA, Google DeepMind, EMBL-EBI and other research organizations released predicted protein-complex structures from 2,812 viral proteomes through the AlphaFold Database, alongside an open GPU-accelerated workflow.

D
Sep 26, 2026 · 2 min read

NVIDIA joined Google DeepMind, EMBL-EBI and other research organizations in releasing predicted protein-complex structures spanning 2,812 viral proteomes through the AlphaFold Database. The collaborators also opened the GPU-accelerated workflow behind the collection, allowing researchers to run it on their own protein targets.

According to the database, the collaboration covered 23 virus families relevant to human health. It produced 5,279 high-confidence heterodimers, which pair different protein chains, and 2,749 high-confidence homodimers, which pair matching chains. The resource is available under the CC BY 4.0 license for academic and commercial use, subject to the database attribution guidance and dataset-specific metadata.

These structures are computational predictions, not experimentally determined structures. The NVIDIA announcement says each prediction is labeled by confidence, while the database describes the additions as high-confidence predictions generated with AlphaFold2 and AlphaFold-Multimer. Confidence scores estimate how reliable a model may be; they do not establish how a virus behaves.

EMBL-EBI said the models cannot show how genetic variation affects a virus, establish host-pathogen interactions or determine whether changes make a virus more transmissible or deadly. Laboratory investigation is still required. “A protein complex structure alone doesn’t tell us what happens when a virus mutates,” Joe Grove, a professor of molecular virology at the University of Glasgow, said in the project announcement. The opened evidence does not demonstrate that the release has improved outbreak response, vaccine development, treatment development or pandemic prevention.

NVIDIA also published the BioNeMo Structure Prediction Pipeline under the Apache License 2.0. For context on another public NVIDIA developer workflow, DataPhoenix has covered CUDA-Q Logical for fault-tolerant quantum development. The open-source BioNeMo pipeline builds multiple-sequence alignments, predicts structures for individual proteins and multi-chain complexes, reports per-residue and aligned-error confidence measures, scores complex interfaces, checks for steric clashes, and exports ModelCIF and BinaryCIF files. Its documentation says researchers supply their own sequences, model weights and reference data, then run the containerized workflow on a Slurm cluster equipped with NVIDIA GPUs.

More news