Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Computational method to reduce the search space for directed protein evolution.

We introduce a computational method to optimize the in vitro evolution of proteins. Simulating evolution with a simple model that statistically describes the fitness landscape, we find that beneficial mutations tend to occur at amino acid positions that are tolerant to substitutions, in the limit of small libraries and low mutation rates. We transform this observation into a design strategy by applying mean-field theory to a structure-based computational model to calculate each residue's structural tolerance. Thermostabilizing and activity-increasing mutations accumulated during the experimental directed evolution of subtilisin E and T4 lysozyme are strongly directed to sites identified by using this computational approach. This method can be used to predict positions where mutations are likely to lead to improvement of specific protein properties.

Bacteriophage T4↗

A general empirical model of protein evolution derived from multiple protein families using a maximum-likelihood approach.

Phylogenetic inference from amino acid sequence data uses mainly empirical models of amino acid replacement and is therefore dependent on those models. Two of the more widely used models, the Dayhoff and JTT models, are estimated using similar methods that can utilize large numbers of sequences from many unrelated protein families but are somewhat unsatisfactory because they rely on assumptions that may lead to systematic error and discard a large amount of the information within the sequences. The alternative method of maximum-likelihood estimation may utilize the information in the sequence data more efficiently and suffers from no systematic error, but it has previously been applicable to relatively few sequences related by a single phylogenetic tree. Here, we combine the best attributes of these two methods using an approximate maximum-likelihood method. We implemented this approach to estimate a new model of amino acid replacement from a database of globular protein sequences comprising 3,905 amino acid sequences split into 182 protein families. While the new model has an overall structure similar to those of other commonly used models, there are significant differences. The new model outperforms the Dayhoff and JTT models with respect to maximum-likelihood values for a large majority of the protein families in our database. This suggests that it provides a better overall fit to the evolutionary process in globular proteins and may lead to more accurate phylogenetic tree estimates. Potentially, this matrix, and the methods used to generate it, may also be useful in other areas of research, such as biological sequence database searching, sequence alignment, and protein structure prediction, for which an accurate description of amino acid replacement is required.

Algorithms↗

Domain deletions and substitutions in the modular protein evolution.

The main mechanisms shaping the modular evolution of proteins are gene duplication, fusion and fission, recombination and loss of fragments. While a large body of research has focused on duplications and fusions, we concentrated, in this study, on how domains are lost. We investigated motif databases and introduced a measure of protein similarity that is based on domain arrangements. Proteins are represented as strings of domains and comparison was based on the classic dynamic alignment scheme. We found that domain losses and duplications were more frequent at the ends of proteins. We showed that losses can be explained by the introduction of start and stop codons which render the terminal domains nonfunctional, such that further shortening, until the whole domain is lost, is not evolutionarily selected against. We demonstrated that domains which also occur as single-domain proteins are less likely to be lost at the N terminus and in the middle, than at the C terminus. We conclude that fission/fusion events with single-domain proteins occur mostly at the C terminus. We found that domain substitutions are rare, in particular in the middle of proteins. We also showed that many cases of substitutions or losses result from erroneous annotations, but we were also able to find courses of evolutionary events where domains vanish over time. This is explained by a case study on the bacterial formate dehydrogenases.

Amino Acid Motifs↗

Selective constraints, amino acid composition, and the rate of protein evolution.

What are the major forces governing protein evolution? A common view is that proteins with strong structural and functional requirements evolve more slowly than proteins with weak constraints, because a stringent negative selection pressure limits the number of substitutions. In contrast, Graur claimed that the substitution rate of a protein is mainly determined by its amino acid composition and the changeabilities of amino acids. In this paper, however, we found that the relative changeabilities of amino acids in mammalian proteins are different for transmembranal and nontransmembranal segments, which have very distinct structural requirements. This indicates that the changeability of a given residue is influenced by the structural and functional context. We also reexamined the relationship between substitution rate and amino acid composition. Indeed, the two kinds of segments exhibit contrasting amino acid compositions: transmembranal regions are made up mainly of hydrophobic residues (a total frequency of approximately 60%) and are very poor in polar amino acids (<5%), whereas nontransmembranal segments have frequencies of 30% and 22%, respectively. Interestingly, we found that within a given integral membrane protein, nontransmembranal segments accumulate, on average, twice as many substitutions as transmembranal regions. However, regression analyses showed that the variability in amino acid frequencies among proteins cannot explain more than 30% of the variability in substitution rate for the transmembranal and nontransmembranal data sets. Furthermore, transmembranal and nontransmembranal segments evolving at the same rate in different proteins have different compositions, and the compositions of slowly evolving and rapidly evolving segments of the same type are similar. From these observations, we conclude that the rate of protein evolution is only weakly affected by amino acid composition but is mostly determined by the strength of functional requirements or selective constraints.

Amino Acid Substitution↗

The quest for the universals of protein evolution.

The sequences of proteins change at dramatically different rates. Unraveling the association between this variability and cellular and ecological processes will enable a better understanding of protein function and evolution. Although weak associations of protein evolutionary rates with a plethora of variables have been reported, none withstands comparison with the effect of expression levels. A recent report by Drummond et al. suggests that the dominant role of expression levels in slowing the rate of protein evolution stems from selection for translation robustness.

Animals↗

Converging on a general model of protein evolution.

The availability of high-throughput genomic databases that establish protein dispensability, expression and interaction networks enables rigorous tests of competing models of protein evolution. Recent research utilizing these new data sets shows that protein evolution is more complex than was previously thought. Several variables, including protein dispensability, expression, functional density, and genetic modularity, appear to have independent effects on the evolutionary rate of proteins, suggesting that proteomes have evolved via an assembly of selectional regimes. These results indicate that a general model of protein evolution will emerge as more functional genomic data from a diversity of organisms accumulate.

Evolution, Molecular↗

Function driven protein evolution. A possible proto-protein for the RNA-binding proteins.

We introduce a hypothesis that present day proteins evolved from "proto-proteins," small 15-20 residue peptides with some elements of secondary structure and primitive function. Increasingly stable and functional proteins arose by adding structural elements to produce the small domains or protein modules that we would recognize today. From this point of view, the surprising similarities between small structural fragments of large proteins, that are usually taken as examples of convergent, function-driven evolution, are interpreted in exactly the opposite way--as traces of common evolutionary origin. As an example, a hypothetical evolutionary tree for two families of RNA binding proteins, the OB fold, a family of all beta proteins, and RBD fold, an alpha/beta protein family is presented. We argue that both protein families could have evolved from the same RNA-binding proto-protein, which had a form of beta-loop-beta RNA binding motif.

Amino Acid Sequence↗

Constant relative rate of protein evolution and detection of functional diversification among bacterial, archaeal and eukaryotic proteins.

BACKGROUND: Detection of changes in a protein's evolutionary rate may reveal cases of change in that protein's function. We developed and implemented a simple relative rates test in an attempt to assess the rate constancy of protein evolution and to detect cases of functional diversification between orthologous proteins. The test was performed on clusters of orthologous protein sequences from complete bacterial genomes (Chlamydia trachomatis, C. muridarum and Chlamydophila pneumoniae), complete archaeal genomes (Pyrococcus horikoshii, P. abyssi and P. furiosus) and partially sequenced mammalian genomes (human, mouse and rat). RESULTS: Amino-acid sequence evolution rates are significantly correlated on different branches of phylogenetic trees representing the great majority of analyzed orthologous protein sets from all three domains of life. However, approximately 1% of the proteins from each group of species deviates from this pattern and instead shows variation that is consistent with an acceleration of the rate of amino-acid substitution, which may be due to functional diversification. Most of the putative functionally diversified proteins from all three species groups are predicted to function at the periphery of the cells and mediate their interaction with the environment. CONCLUSIONS: Relative rates of protein evolution are remarkably constant for the three species groups analyzed here. Deviations from this rate constancy are probably due to changes in selective constraints associated with diversification between orthologs. Functional diversification between orthologs is thought to be a relatively rare event. However, the resolution afforded by the test designed specifically for genomic-scale datasets allowed us to identify numerous cases of possible functional diversification between orthologous proteins.

Animals↗

Domain rearrangements in protein evolution.

Most eukaryotic proteins are multi-domain proteins that are created from fusions of genes, deletions and internal repetitions. An investigation of such evolutionary events requires a method to find the domain architecture from which each protein originates. Therefore, we defined a novel measure, domain distance, which is calculated as the number of domains that differ between two domain architectures. Using this measure the evolutionary events that distinguish a protein from its closest ancestor have been studied and it was found that indels are more common than internal repetition and that the exchange of a domain is rare. Indels and repetitions are common at both the N and C-terminals while they are rare between domains. The evolution of the majority of multi-domain proteins can be explained by the stepwise insertions of single domains, with the exception of repeats that sometimes are duplicated several domains in tandem. We show that domain distances agree with sequence similarity and semantic similarity based on gene ontology annotations. In addition, we demonstrate the use of the domain distance measure to build evolutionary trees. Finally, the evolution of multi-domain proteins is exemplified by a closer study of the evolution of two protein families, non-receptor tyrosine kinases and RhoGEFs.

Databases, Protein↗

Protein evolution is faster outside the cell.

Some proteins are highly conserved across all species, whereas others diverge significantly even between closely related species. Attempts have been made to correlate the rate of protein evolution to amino acid composition, protein dispensability, and the number of protein-protein interactions, but in all cases, conflicting studies have shown that the theories are hard to confirm experimentally. The only correlation that is undisputed so far is that highly/broadly expressed proteins seem to evolve at a lower rate. Consequently, it has been suggested that correlations between evolution rate and factors like protein dispensability or the number of protein-protein interactions could be just secondary effects due to differences in expression. The purpose of this study was to analyze mammalian proteins/genes with known subcellular location for variations in evolution rates. We show that proteins that are exported (extracellular proteins) evolve faster than proteins that reside inside the cell (intracellular proteins). We find weak, but significant, correlations between evolution rates and expression levels, percentage of tissues in which the proteins are expressed (expression broadness), and the number of protein interaction partners. More important, we show that the observed difference in evolution rate between extra- and intracellular proteins is largely independent of expression levels, expression broadness, and the number of protein-protein interactions. We also find that the difference is not caused by an overrepresentation of immunological proteins or disulfide bridge-containing proteins among the extracellular data set. We conclude that the subcellular location of a mammalian protein has a larger effect on its evolution rate than any of the other factors studied in this paper, including expression levels/patterns. We observe a difference in evolution rates between extracellular and intracellular proteins for a yeast data set as well and again show that it is completely independent of expression levels.

Amino Acid Sequence↗

A structure-centric view of protein evolution, design, and adaptation.

Proteins, by virtue of their central role in most biological processes, represent one of the key subjects of the study of molecular evolution. Inherent in the indispensability of proteins for living cells is the fact that a given protein can adopt a specific three-dimensional shape that is specified solely by the protein's sequence of amino acids. Over the past several decades, structural biologists have demonstrated that the array of structures that proteins may adopt is quite astounding, and this has lead to a strong interest in understanding how protein structures change and evolve over time. In this review we consider a large body of recent work that attempts to illuminate this structure-centric picture of protein evolution. Much of this work has focused on the question of how completely new protein structures (i.e., new folds or topologies) are discovered by protein sequences as they evolve. Pursuant to this question of structural innovation has been a desire to describe and understand the observation that certain types of protein structures are far more abundant than others and how this uneven distribution of proteins implicates on the process through which new shapes are discovered. We consider a number of theoretical models that have been successful at explaining this heterogeneity in protein populations and discuss the increasing amount of evidence that indicates that the process of structural evolution involves the divergence of protein sequences and structures from one another. We also consider the topic of protein designability, which concerns itself with understanding how a protein's structure influences the number of sequences that can fold successfully into that structure. Understanding and quantifying the relationship between the physical feature of a structure and its designability has been a long-standing goal of the study of protein structure and evolution, and we discuss a number of recent advances that have yielded a promising answer to this question. Finally, we review the relatively new field of protein structural phylogeny, an area of study in which information about the distribution of protein structures among different organisms is used to reconstruct the evolutionary relationships between them. Taken together, the work that we review presents an increasingly coherent picture of how these unique polymers have evolved over the course of life on Earth.

Adaptation, Biological↗

Generality of the structurally constrained protein evolution model: assessment on representatives of the four main fold classes.

The Structurally Constrained Protein Evolution (SCPE) model simulates protein evolution by introducing random mutations into the evolving sequences and selecting them against too much structural perturbation. Given a single protein structure, the SCPE model can be used to obtain a whole set of site-dependent amino acid substitution matrices. The set of SCPE substitution matrices for a given protein family can be seen as an independent-sites model of evolution for that family. Thus, these matrices can be compared with other substitution-matrix-based models of evolution. So far, SCPE has been tested only on left-handed parallel beta helix (LbetaH) proteins. Here, we address the question of generality by assessing the SCPE model on representatives of the four main classes of folds: alpha, beta, alpha+beta, and alpha/beta. We compare with other models using the likelihood ratio test with parametric bootstrapping. We show that SCPE performs better than the popular JTT model for all cases considered. Furthermore, by considering the relative contributions of mutation and selection, we found that the key to the success of the SCPE model is the selection step.

Algorithms↗

Functional genomic analysis of the rates of protein evolution.

The evolutionary rates of proteins vary over several orders of magnitude. Recent work suggests that analysis of large data sets of evolutionary rates in conjunction with the results from high-throughput functional genomic experiments can identify the factors that cause proteins to evolve at such dramatically different rates. To this end, we estimated the evolutionary rates of >3,000 proteins in four species of the yeast genus Saccharomyces and investigated their relationship with levels of expression and protein dispensability. Each protein's dispensability was estimated by the growth rate of mutants deficient for the protein. Our analyses of these improved evolutionary and functional genomic data sets yield three main results. First, dispensability and expression have independent, significant effects on the rate of protein evolution. Second, measurements of expression levels in the laboratory can be used to filter data sets of dispensability estimates, removing variates that are unlikely to reflect real biological effects. Third, structural equation models show that although we may reasonably infer that dispensability and expression have significant effects on protein evolutionary rate, we cannot yet accurately estimate the relative strengths of these effects.

Evolution, Molecular↗

Antagonists to human and mouse vascular endothelial growth factor receptor 2 generated by directed protein evolution in vitro.

Using directed in vitro protein evolution, we generated proteins that bound and antagonized the function of vascular endothelial growth factor receptor 2 (VEGFR2). Binders to human VEGFR2 (KDR) with 10-200 nM affinities were selected by using mRNA display from a library (10(13) variants) based on the tenth human fibronectin type III domain (10Fn3) scaffold. Subsequently, a single KDR binding clone (K(d) = 11 nM) was subjected to affinity maturation. This yielded improved KDR binding molecules with affinities ranging from 0.06 to 2 nM. Molecules with dual binding specificities (human/mouse) were also isolated by using both KDR and Flk-1 (mouse VEGFR2) as targets in selection. Proteins encoded by the selected clones bound VEGFR2-expressing cells and inhibited their VEGF-dependent proliferation. Our results demonstrate the potential of these inhibitors in the development of anti-angiogenesis therapeutics.

Amino Acid Sequence↗

Thermodynamics of neutral protein evolution.

Naturally evolving proteins gradually accumulate mutations while continuing to fold to stable structures. This process of neutral evolution is an important mode of genetic change and forms the basis for the molecular clock. We present a mathematical theory that predicts the number of accumulated mutations, the index of dispersion, and the distribution of stabilities in an evolving protein population from knowledge of the stability effects (delta deltaG values) for single mutations. Our theory quantitatively describes how neutral evolution leads to marginally stable proteins and provides formulas for calculating how fluctuations in stability can overdisperse the molecular clock. It also shows that the structural influences on the rate of sequence evolution observed in earlier simulations can be calculated using just the single-mutation delta deltaG values. We consider both the case when the product of the population size and mutation rate is small and the case when this product is large, and show that in the latter case the proteins evolve excess mutational robustness that is manifested by extra stability and an increase in the rate of sequence evolution. All our theoretical predictions are confirmed by simulations with lattice proteins. Our work provides a mathematical foundation for understanding how protein biophysics shapes the process of evolution.

Computer Simulation↗

Simulation of protein evolution: evidence for a non-linear aminoacidic substitution rate.

Protein evolution is characterized by several processes. In the theory of neutral evolution the rate of mutation is considered a linear process in which the amount of aminoacidic substitutions in proteins is constant in time. A simulation approach has been developed by using a model of amino-acidic substitution. The frequency of spontaneous mutations has been assumed to be equal to about 10(-9)/base/year. The aim of the present work is to show that starting from a constant mutation rate (nucleotide substitutions) the corresponding process of aminoacidic substitutions becomes non-linear if some criteria of mutation selection are introduced. The basic criteria used are the physical-chemical characteristics of aminoacids, the same criteria that have made it possible to classify aminoacids. Different classifications based on differences in such criteria give different results, indicating that the degenerate nature of the genetic code determines a non-linear behaviour of protein evolution. Simulations have been performed on short protein subsequences of five aminoacids. A further analysis has been made to verify, on the basis of the code structure and of accepted selection criteria, the mechanisms of aminoacidic substitutions and the existence of preferential paths. We have concluded that aminoacidic substitution is not a simple stochastic process, but that complex Markov's chains are involved. The consequences are important, although generally ignored.

Amino Acid Substitution↗

Protein evolution viewed through Escherichia coli protein sequences: introducing the notion of a structural segment of homology, the module.

Paralogous genes are genes which descend from a progenitor gene which has duplicated as an ancestral gene, each copy having diverged prior to speciation. With comprehensive information available on functions of Escherichia coli proteins, analysis of sequence-related E. coli paralogous proteins can give information on the early ancestors of families of proteins now residing in many contemporary organisms, such as the enzymes of metabolism, some kinds of transport mechanisms and some kinds of regulatory mechanisms. In the first step, we have confirmed that E. coli contains a very high proportion of paralogous proteins. Next, we have defined two main classes of paralogous proteins. One class is formed of proteins which contain a unique structural segment homologous to a single set of related proteins. The other class corresponds to proteins which contain more than one structural segment of homology, each segment homologous to unrelated sets of proteins. We define such an independent structural segment of homology as a module. This modular structure (mean length equivalent to 209 amino acids) corresponds often to entire proteins, but there are also proteins that appear to be assembled from two or three independent modules having independent origins. Most multimodular proteins appear to have been formed early in their history, a minority appear to be relatively recent fusions of independent modules. Examining 1404 independent structural segments of homology, composed of both modules and entire proteins, we found that the segments of homology fell into 352 sequence-related groups or families. The majority of these families (ranging from 2 to 62 members) are functionally homogeneous. This strongly suggests that the 1404 present-day modules and proteins derive from a minimal set of 352 ancestral modules, each one being already of the same size and having a function similar to all members of its progeny.

Bacterial Proteins↗