Independent education resourceInformation here does not replace care from a qualified health professional.
Peptide Therapy GuideClear peptide education

Educational guide

AlphaFold Database expands with millions of predicted protein complexes

A new collaboration between EMBL's European Bioinformatics Institute (EMBL-EBI), Google DeepMind, NVIDIA, and Seoul National University has made millions of AI-predicted protein complex structures openly available through the AlphaFold Database. To maximise gl

Written by Peptide Therapy Guide Editorial Team
For education only

This guide cannot diagnose a condition or recommend a personal treatment plan. Discuss medical questions with a qualified professional.

A new collaboration between EMBL's European Bioinformatics Institute (EMBL-EBI), Google DeepMind, NVIDIA, and Seoul National University has made millions of AI-predicted protein complex structures openly available through the AlphaFold Database. To maximise global health impact, the dataset prioritizes proteins important for understanding human health and disease. This is the largest dataset of protein complex predictions currently available.

Proteins are the building blocks of life. They interact to create protein complexes which fulfil biological functions. By visualizing protein interactions, scientists can uncover the molecular mechanisms that drive cell behavior, identify what goes wrong when someone gets sick, and develop new drugs and therapies. Predicting the structure of protein complexes is extremely challenging because, in nature, proteins change shape and interact in many different ways.

Science thrives on collaboration. By making this foundational protein complex dataset openly available to the world, we're inviting researchers to test, refine, and build on it to drive the next wave of biological discoveries." Jo McEntyre, Interim Director of EMBL-EBI

Protein complexes for global health impact

The latest AlphaFold Database update spans millions of homodimers – protein complexes formed of two identical proteins. It focuses on 20 of the most studied species, including humans, as well as the World Health Organization's bacterial priority pathogens list. This approach aims to bring significant and immediate value for global health challenges.

"By expanding the AlphaFold Database to include protein complexes, we are addressing a critical need expressed by the scientific community," said Anna Koivuniemi, Head of the Google DeepMind Impact Accelerator. "We hope that by lowering the barrier to these complex predictions, we can empower researchers everywhere to pursue the next wave of discoveries that could ultimately improve human health on a global scale."

Scientific expertise meets technical innovation

The collaboration builds on Google DeepMind's AI system AlphaFold, which, since 2021, accurately predicted the structure of millions of proteins. To democratize access to AlphaFold predictions, Google DeepMind and EMBL-EBI developed the AlphaFold Database, an open resource that anyone can access. The database has over 3.4 million users from 190 countries.

Through ongoing dialogue with the scientific community, a clear need emerged to expand the AlphaFold database to include protein complexes. In response to this need, EMBL-EBI, Google DeepMind, NVIDIA, and Seoul National University teamed up, contributing specialist expertise and resources, to calculate and integrate millions of protein complexes into the AlphaFold Database.

The collaboration brought together deep biological expertise and technical innovations. NVIDIA and the Steinegger Lab at the Seoul National University developed the methodology, based on Google DeepMind's AI system AlphaFold, including accelerations to multiple sequence alignment calculations and deep learning inference. NVIDIA provided cutting-edge AI infrastructure and scaled out inference pipelines to overcome limitations that historically made this scale of calculations challenging. EMBL-EBI enabled the collaboration by bringing the other parties together and contributing expertise in scientific and biodata management, as well as analysis. As a champion of open science, EMBL-EBI, together with Google DeepMind, integrated the new dataset into the AlphaFold Database.

"NVIDIA's ambition is to consistently contribute orders-of-magnitude accelerations for fundamental digital biology workloads, enabling what was not possible before," said Anthony Costa, NVIDIA Director of Digital Biology. "This release is a great example of how AI infrastructure and software can uniquely enable new scales of biological understanding."

"By making predicted protein complexes accessible at an unprecedented scale, we are illuminating an unseen landscape of molecular interactions across the tree of life," explained Martin Steinegger, Associate Professor at Seoul National University.

Open science at scale

It takes a blend of AI-scale infrastructure and deep technical knowledge in accelerating complex workflows to generate AI predictions for protein complexes at this scale. The collaboration is centrally hosting data that would otherwise require around 17 million hours of GPU (graphics processing unit) computing to recreate.

By making these calculations once and adding the information into the AlphaFold Database, this collaboration aims to help democratize access to protein complex predictions. It enables scientists everywhere to investigate how proteins interact in the vast protein universe, and accelerate discoveries that could lead to new medicines, new products, and a deeper understanding of life itself.

This is the first step in an ambition to add a wide range of protein complex structure predictions to the AlphaFold Database. The partnership has already calculated predictions for 30 million complexes. Of these, 1.7 million high-confidence homodimer predictions have been added to the AlphaFold Database. Another 18 million are lower-confidence homodimers, which are available as a list and for bulk download. The rest are heterodimers, currently being analysed and assessed. More protein complex predictions will be calculated and high-confidence predictions will be added to the AlphaFold Database in the coming months. The work is described in more detail in a preprint.

"The human genome has just over 20,000 different proteins. Despite this relatively small genome, human beings display incredibly complex pathways, processes and regulation. Much of this complexity arises from the intermolecular interactions between proteins, and with small molecule ligands and DNA. Adding predicted protein-protein homodimeric interactions to the AlphaFold Database is a first step towards a comprehensive description of the human interactome, the basis by which human biology will be described and understood. This has relevance for the design of new therapeutics, understanding host-pathogen interactions, and more. Making these structures accessible to all, allows every researcher around the world to build on these data, moving one step closer to predicting the biology of life," said Dame Janet Thornton, Director Emeritus of EMBL-EBI.

Connected reading

Helpful context for this guide

Source-derived material selected through this article’s indexed topics.

Related questions

01What are 'dancing molecules'?

Stupp and his team posited that "dancing molecules" might encourage the stubborn tissue to regenerate. Previously invented in Stupp's laboratory, dancing molecules are assemblies that form synthetic nanofibers comprising tens to hundreds of thousands of molecules with potent signals for cells. By tuning their collective motions through their chemical structure, Stupp discovered the moving molecules could rapidly find and properly engage with cellular receptors, which also are in constant motion and extremely crowded on cell membranes. Once inside the body, the nanofibers mimic the extracellular matrix of the surrounding tissue. By matching the matrix's structure, mimicking the motion of biological molecules and incorporating bioactive signals for the receptors, the synthetic materials are able to communicate with cells. "Cellular receptors constantly move around," Stupp said. "By making our molecules move, 'dance' or even leap temporarily out of these structures, known as supramolecular polymers, they are able to connect more effectively with receptors."

Source: www.news-medical.net ↗
02Could you discuss any ongoing or future research projects that you are particularly excited about in the field of axon biology and sncRNAs?

Certainly. We're keenly exploring extracellular vesicles as mediums for cells to communicate. These cell made vesicles often carry microRNAs and other non-coding RNAs, presenting a unique avenue to understand how neurons modulate their environment, which is especially interesting in neurological conditions. We are exploring extracellular vesicles as these tools with which cells can communicate and transfer gene expression patterns. And we're looking at, for example, how early life brain tumours such as medulloblastoma can impact neuron development and activity and how this can affect later life pain processing and neurological conditions. This has been possible via funding from the Medical Research Foundation, which supported a big collaboration between the labs of Gareth Hathway, Beth Coyle, Vicky James, Anna Grabowska and myself in Nottingham.

Source: www.news-medical.net ↗
03What knowledge has your research provided, specifically regarding Alzheimer's disease?

The main protein or peptide in Alzheimer's disease patients is something called beta-amyloid. It has a 42 amino acid chain, and accumulates in the plaque between nerve cells in Alzheimer's disease patients. Beta-amyloid has been found to epimerize. Certain amino acids epimerize in the beta-amyloid at specific locations, and it turns out those amino acids are mainly aspartic acid and serine. There was no good way to analyze for such aberrations in beta-amyloid until now, and we have found a number of techniques that not only detect it, but can point out exactly where, within the beta-amyloid chain, the aberration is occurring and if there are more aberrations in Alzheimer's patients than in normal people. Image Credit:Shutterstock/Tavarius

Source: www.news-medical.net ↗
04What’s your vision for BenevolentBio?

I want us to disrupt the drug discovery and development process and to look at each place on the drug discovery and development pipeline, so that we can be much better at getting the right target, much quicker at getting the right compound and much more confident that those compounds have the right characteristics which mean they will be safe and well tolerated. Then we can go to the right patient population with the right dose, so we would have a much leaner, more successful process and be able to demonstrate the value of our AI technology.

Source: www.news-medical.net ↗
05How does the zwitterionic form of an amino acid relate to the pI?

When the pH is exactly at the pKa value, a buffer arises in which the deprotonated and protonated amino acids exist in equilibrium. For example, when the pH = 2.34 (pKa of glycine), the solution comprises of 50% neutral molecules in which the carboxyl is deprotonated, and 50% positive molecules where the carboxyl is protonated. This pH produces the carboxyl buffer zone. If the pH s increased to that of the pKa of the amino group (9.60), another buffer is produced where there is an equilibration between the protonated neutral zwitterion and the deprotonated negative amino acid. The isoelectric point can, therefore, be approximated by averaging the two pKa values. More complex amino acids have more than two pKa values due to the presence of additional pKa values for their side chains.

Source: www.news-medical.net ↗
P

About the author

Peptide Therapy Guide Editorial Team

Editorial team for Peptide Therapy Guide.

View all articles →