Product

Showing posts with label protein. Show all posts
Showing posts with label protein. Show all posts

Sunday, 5 June 2022

Flexible and protected – new findings on SARS-CoV-2 protein shed light on virus’s ability to infect cells

 In the fight against the coronavirus, SARS-CoV-2 researchers from multiple research institutions in Germany have combined their resources to study the spike protein on the surface of the virus. With its spikes, the virus binds to human cells and infects them. The study gave surprising insights into the spike protein, including an unexpected freedom of movement and a protective coat to hide it from antibodies. The results are published in Science.

This is a simulation of four spike proteins (red, orange, blue and grey) of the SARS-CoV-2 virus. The proteins and lipids are shown in surface representation. The protective glycan chains are shown in green. Credit: Sören von Bülow, Mateusz Sikora, Gerhard Hummer/MPI of Biophysics

At the start of a COVID-19 infection, the coronavirus SARS-CoV-2 docks onto human cells using the spike-like proteins on its surface. The spike protein is at the centre of vaccine development because it triggers an immune response in humans. A group of German scientists, including members of EMBL in Heidelberg, the Max Planck Institute of Biophysics, the Paul-Ehrlich-Institut, and Goethe University Frankfurt have focused on the surface structure of the virus to gain insights they can use for the development of vaccines and of effective therapeutics to treat infected patients.

The team combined cryo-electron tomography, subtomogram averaging, and molecular dynamics simulations to analyse the molecular structure of the spike protein in its natural environment, on intact virions, and with near-atomic resolution. Using EMBL’s state-of-the-art cryo-electron microscopy imaging facility, 266 cryotomograms of about 1000 different viruses were generated, each carrying an average of 40 spikes on its surface. Subtomogram averaging and image processing, combined with molecular dynamics simulations, finally provided the important and novel structural information on these spikes.

The results were surprising: the data showed that the globular portion of the spike protein, which contains the receptor-binding region and the machinery required for fusion with the target cell, is connected to a flexible stalk. “The upper spherical part of the spike has a structure that is well reproduced by recombinant proteins used for vaccine development,” explains Martin Beck, EMBL group leader and a director of the Max Planck Institute (MPI) of Biophysics. “However, our findings about the stalk, which fixes the globular part of the spike protein to the virus surface, were new.”

“The stalk was expected to be quite rigid,” adds Gerhard Hummer, from the MPI of Biophysics and the Institute of Biophysics at Goethe University Frankfurt. “But in our computer models and in the actual images, we discovered that the stalks are extremely flexible.” By combining molecular dynamics simulations and cryo-electron tomography, the team identified the three joints – hip, knee and ankle – that give the stalk its flexibility.

“Like a balloon on a string, the spikes appear to move on the surface of the virus and thus are able to search for the receptor for docking to the target cell,” explains Jacomine Krijnse Locker, group leader at the Paul-Ehrlich-Institut. To prevent infection, these spikes are targeted by antibodies. However, the images and models also showed that the entire spike protein, including the stalk, is covered with chains of glycans – sugar-like molecules. These chains provide a kind of protective coat that hides the spikes from neutralising antibodies: another important finding on the way to effective vaccines and medicines.

Solving the protein structure puzzle

 Proteins are beautiful molecular structures and understanding what they look like has been a goal for scientists for more than half a century. After years of arduous work and frustratingly slow progress, a game-changing artificial intelligence method is poised to disrupt the field.

We call proteins the building blocks of life because they make up all living things, from the smallest virus or bacterium to plants, animals, and humans. But, in reality, proteins don’t look anything like blocks. They are beautifully complex structures and every single one of them is unique. Their shape, also called a structure, is linked to their function, which means their shape determines what they do. For example, haemoglobin transports oxygen around the body, while insulin maintains the delicate balance of sugar within the blood.

Simple question, complex answer

Studying protein structure means you’re faced with a very simple question that requires a very complex answer. A protein is a string of small organic molecules called amino acids, connected in a chain, a bit like beads on a string. This chain of amino acids spontaneously folds up to create a unique and beautiful structure. The simple question is: what does the structure look like?

This problem has been around now for at least 50 years and, after many failed attempts, I came to believe that the only way to make progress was to gather more data to make better predictions. I was proven wrong.

Figuring out the structure of just one protein can take years of experimental work, using expensive equipment and incredibly complex methodology. One method is X-ray crystallography, which blasts crystalline molecules with an X-ray beam. This beam diffracts into many directions and, by measuring the angles and intensities, crystallographers can produce a 3D picture of the density of electrons within the crystal. This reveals the structure of complex biological molecules, including proteins. One of the difficulties of the method is obtaining the crystals, and sadly this method simply hasn’t worked for some proteins.

Experimental meets computational

Luckily, there is an incredibly active and tenacious community of scientists who have dedicated their lives to predicting protein structures or how the chain folds from their amino acid sequences. All newly determined structures are stored in the Protein Data Bank (and its European node, PDBe) and are freely available for anyone in the world to look at.

In the mid 1990s, the need to coordinate efforts and assess progress became clearer than ever, so the community embarked on a worldwide experiment, called Critical Assessment of protein Structure Prediction (CASP). Every two years the organisers launch the challenge of predicting the structure of several proteins. The objective is to test and independently assess new computational methods for structure prediction. These methods use computers, not lab experiments, to predict protein structure. The methods, now increasingly powered by artificial intelligence (AI), had been improving over the past few years, but a solution still seemed a long way off.

That is until this week, when – during the latest CASP conference – the assessors announced that one team, DeepMind’s AlphaFold, had put forward an AI system that achieved unparalleled levels of accuracy. This approach built on our extensive knowledge of protein structures obtained in the lab over the past 60 years. But this was the first time a computational model was deemed to be competitive with experimental methods. And something that would have taken years of experimental work can now be deduced within just days using a new type of neural network.

Why does it matter?

There are millions of proteins that make up the living world, but we only know the structures of a tiny number of them. In fact, we only have experimental  structures (or even partial structures) for 10% of the 20 000 proteins that make up the human body. A powerful AI model could unveil the structures of the other 90%. This is important not just because it improves our understanding of human biology, health, and disease, but also because in the longer term it would offer avenues of research, for example designing new drugs.

Most existing drugs are designed using 3D structures, but they currently target only about a quarter of human proteins. AlphaFold could help unlock more proteins as potential drug targets and open up new approaches to therapies. Furthermore, easily predicting the structure of viruses can help us understand their biology and the diseases they cause. Finally, there may be significant opportunities to understand and treat neglected tropical diseases, where research is currently under-resourced.

The potential goes beyond human health. Understanding plant and animal proteins (as well as their genomes) could help us improve crop yields or breeding procedures. This would hold significant potential for feeding a growing population.

Finally, at a more scientific level, being able to predict structure from sequence is the first real step towards protein design: building proteins that fulfil a specific function. From protein therapeutics to biofuels or enzymes that eat plastic, the possibilities are endless. 

A fine time for protein science

Understanding proteins is a bit like putting together a large 3D jigsaw puzzle in a dark room. You know what some of the pieces look like and you can sometimes match a few together in clusters, but it’s incredibly arduous and a complete solution would rarely be found. A fast and accessible method for determining the whole structure in the computer solves the puzzle automatically.

As a lover of everything protein, the most exciting thing for me is that this breakthrough is not an end, but a whole new beginning, bringing with it electrifying opportunities and follow-on questions. The structures allow us to understand better how the proteins function and, in turn, this could enable us to fine-tune this function for the benefit of people and the planet. Just like the Human Genome Project facilitated the birth of new scientific disciplines, such as genomics, solving the protein structure question could bring about new and exciting fields of research. One thing is for sure, it’s a fine time to be a protein scientist!

Friday, 3 June 2022

Deep learning models help predict protein function

 Deep learning models can improve protein annotations and has helped expand the Pfam database

Our protein family database – Pfam – is used by a diverse range of researchers across the globe. Open access to the protein family data stored in Pfam has helped experimental biologists understand protein function, aided structural biologists’ insights into protein structure, given computational biologists rapid access to protein sequence information, and let evolutionary biologists trace the origins of proteins.

Pfam gives researchers access to vital protein annotations, structures, and multiple sequence alignments. It is a resource widely used to classify protein sequences into phylogenies and identify domains – functional regions – to provide insights into protein function.

With help from new deep learning models, Pfam has increased the protein sequence annotation and function data available within the database by unprecedented amounts. Research published in the journal Nature Biotechnology demonstrates how deep learning methods developed by Google Research could be trained using data from Pfam to accurately annotate many previously undescribed protein domains, shedding light on potential protein function. This new data added to Pfam has expanded the database to such an extent, it would have taken several years to achieve the same result manually.

Deep learning and protein function

“Initially I was rather sceptical about using deep learning to reproduce the protein families within Pfam. Then I started collaborating more closely with Lucy Colwell and her team at Google Research and my scepticism quickly changed to excitement for the potential of these methods to improve our ability to classify sequences into domains and families,” said Alex Bateman, Senior Team Leader of Protein Sequence Resources at EMBL-EBI. “These models exceed my expectations. They’re not just copying the data already in Pfam, they’re able to learn from the data and find new information that is yet to be discovered. What this gives us is the ability to expand the Pfam collection and potentially that of other resources using these same deep learning methods.”

By combining deep learning models with existing methods to add new data into Pfam, the researchers were able to expand the database by almost 10%. This exceeds all expansion efforts made to the database over the last decade. The deep learning methods were also able to predict the function for 360 human proteins that had no previous annotation data available in Pfam.

Expanding Pfam

Using additional protein family predictions generated from the Google Research team’s neural networks – a series of algorithms that looks for underlying structure in the sequences of protein domains and families – created a supplement to Pfam called Pfam-N, where N stands for network. Pfam-N adds a further 6.8 million protein sequences to the Pfam database.

“We’re also now building on these established deep learning methods to expand the information in the database even further,” said Bateman. “We’re changing the way the existing deep learning model works so that we can call multiple protein domains at once. This new update to the database should be ready very soon.”

“My personal view is that there’s still a lot of scope to improve the deep learning models we’re currently using,” Bateman added. “We’re in the early days of this and I’m very hopeful for what it will mean for the future classification of protein families. This may even be something that will get solved in the next five years.”

Find out more

Find out more about Pfam’s collaboration with Google Research and get a detailed introduction to Pfam-N in this Xfam blog post.

Funding

This work is funded by the Wellcome Trust as part of a Biomedical Resources grant awarded to the Pfam database.


Source article(s)

Connect broadband

Why do governments, corporations, and experts promote eggs, meat, and other animal foods?

  Your question combines nutrition, public policy, ethics, religion, psychology, and AI. It's useful to separate evidence-based facts ...