Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

The emergence of scaling in sequence-based physical models of protein evolution.

It has recently been discovered that many biological systems, when represented as graphs, exhibit a scale-free topology. One such system is the set of structural relationships among protein domains. The scale-free nature of this and other systems has previously been explained using network growth models that, although motivated by biological processes, do not explicitly consider the underlying physics or biology. In this work we explore a sequence-based model for the evolution protein structures and demonstrate that this model is able to recapitulate the scale-free nature observed in graphs of real protein structures. We find that this model also reproduces other statistical feature of the protein domain graph. This represents, to our knowledge, the first such microscopic, physics-based evolutionary model for a scale-free network of biological importance and as such has strong implications for our understanding of the evolution of protein structures and of other biological networks.

Algorithms↗

ProtTest: selection of best-fit models of protein evolution.

SUMMARY: Using an appropriate model of amino acid replacement is very important for the study of protein evolution and phylogenetic inference. We have built a tool for the selection of the best-fit model of evolution, among a set of candidate models, for a given protein sequence alignment. AVAILABILITY: ProtTest is available under the GNU license from http://darwin.uvigo.es

Algorithms↗

Correlation between the substitution rate and rate variation among sites in protein evolution.

It is well known that the rate of amino acid substitution varies among different proteins and among different sites of a protein. It is, however, unclear whether the extent of rate variation among sites of a protein and the mean substitution rate of the protein are correlated. We used two approaches to analyze orthologous protein sequences of 51 nuclear genes of vertebrates and 13 mitochondrial genes of mammals. In the first approach, no assumptions of the distribution of the rate variation among sites were made, and in the second approach, the gamma distribution was assumed. Through both approaches, we found a negative correlation between the extent of among-site rate variation and the average substitution rate of a protein. That is, slowly evolving proteins tend to have a high level of rate variation among sites, and vice versa. We found this observation consistent with a simple model of the neutral theory where most sites are either invariable or neutral. We conclude that the correlation is a general feature of protein evolution and discuss its implications in statistical tests of positive Darwinian selection and molecular time estimation of deep divergences.

Animals↗

A single determinant dominates the rate of yeast protein evolution.

A gene's rate of sequence evolution is among the most fundamental evolutionary quantities in common use, but what determines evolutionary rates has remained unclear. Here, we carry out the first combined analysis of seven predictors (gene expression level, dispensability, protein abundance, codon adaptation index, gene length, number of protein-protein interactions, and the gene's centrality in the interaction network) previously reported to have independent influences on protein evolutionary rates. Strikingly, our analysis reveals a single dominant variable linked to the number of translation events which explains 40-fold more variation in evolutionary rate than any other, suggesting that protein evolutionary rate has a single major determinant among the seven predictors. The dominant variable explains nearly half the variation in the rate of synonymous and protein evolution. We show that the two most commonly used methods to disentangle the determinants of evolutionary rate, partial correlation analysis and ordinary multivariate regression, produce misleading or spurious results when applied to noisy biological data. We overcome these difficulties by employing principal component regression, a multivariate regression of evolutionary rate against the principal components of the predictor variables. Our results support the hypothesis that translational selection governs the rate of synonymous and protein sequence evolution in yeast.

Amino Acid Substitution↗

Two types of amino acid substitutions in protein evolution.

The frequency of amino acid substitutions, relative to the frequency expected by chance, decreases linearly with the increase in physico-chemical differences between amino acid pairs involved in a substitution. This correlation does not apply to abnormal human hemoglobins. Since abnormal hemoglobins mostly reflect the process of mutation rather than selection, the correlation manifest during protein evolution between substitution frequency and physico-chemical difference in amino acids can be attributed to natural selection. Outside of 'abnormal' proteins, the correlation also does not apply to certain regions of proteins characterized by rapid rates of substitution. In these cases again, except for the largest physico-chemical differences between amino acid pairs, the substitution frequencies seem to be independent of the physico-chemical parameters. The limination of the substituents involving the largest physico-chemical differences can once more be attributed to natural selection. For smaller physico-chemical differences, natural selection, if it is operating in the polypeptide regions, must be based on parameters other than those examined.

Amino Acid Sequence↗

Structural determinants of the rate of protein evolution in yeast.

We investigate how a protein's structure influences the rate at which its sequence evolves. Our basic hypothesis is that proteins with highly designable structures (structures that are encoded by many sequences) will evolve more rapidly. Recent theoretical advances argue that structures with a higher density of interresidue contacts are more designable, and we show that high contact density is correlated with an increased rate of sequence evolution in yeast. In addition, we investigate the correlations between the rate of sequence evolution and several other structural descriptors, carefully controlling for the strong effect of expression level on evolutionary rate. Overall, we find that the structural descriptors that we consider appear to explain roughly 10% of the variation in rates of protein evolution in yeast. We also show that despite the well-known trend for buried residues to be more conserved, proteins with a higher fraction of buried residues, nonetheless, tend to evolve their sequences more rapidly. We suggest that this effect is due to the increased designability of structures with more buried residues. Our results provide evidence that protein structure plays an important role in shaping the rate of sequence evolution and provide evidence to support recent theoretical advances linking structural designability to contact density.

Analysis of Variance↗

What amino acid properties affect protein evolution?

We studied 10 protein-coding mitochondrial genes from 19 mammalian species to evaluate the effects of 10 amino acid properties on the evolution of the genetic code, the amino acid composition of proteins, and the pattern of nonsynonymous substitutions. The 10 amino acid properties studied are the chemical composition of the side chain, two polarity measures, hydropathy, isoelectric point, volume, aromaticity, aliphaticity, hydrogenation, and hydroxythiolation. The genetic code appears to have evolved toward minimizing polarity and hydropathy but not the other seven properties. This can be explained by our finding that the presumably primitive amino acids differed much only in polarity and hydropathy, but little in the other properties. Only the chemical composition (C) and isoelectric point (IE) appear to have affected the amino acid composition of the proteins studied, that is, these proteins tend to have more amino acids with typical C and IE values, so that nonsynonymous mutations tend to result in small differences in C and IE. All properties, except for hydroxythiolation, affect the rate of nonsynonymous substitution, with the observed amino acid changes having only small differences in these properties, relative to the spectrum of all possible nonsynonymous mutations.

Amino Acid Substitution↗

Family specific rates of protein evolution.

MOTIVATION: Amino acid changing mutations in proteins are contstrained by purifying selection and accumulate at different rates. We estimate evolutionary rates on multiple alignments of eukaryotic protein families in a maximum likelihood framework and spot sets of slow and fast evolving proteins. RESULTS: We find that the evolution of indispensable proteins is constrained by selection and that protein secretion is coupled to an increased evolutionary rate.

Algorithms↗

Observations of amino acid gain and loss during protein evolution are explained by statistical bias.

The authors of a recent manuscript in "Nature" claim to have discovered "universal trends" of amino acid gain and loss in protein evolution. Here, we show that this universal trend can be simply explained by a bias that is unavoidable with the 3-taxon trees used in the original analysis. We demonstrate that a rigorously reversible equilibrium model, when analyzed with the same methods as the "Nature" manuscript, yields identical (and in this case, clearly erroneous) conclusions. A main source of the bias is the division of the sequence data into "informative" and "noninformative" sites, which favors the observation of certain transitions.

Algorithms↗

A protein evolution model with independent sites that reproduces site-specific amino acid distributions from the Protein Data Bank.

BACKGROUND: Since thermodynamic stability is a global property of proteins that has to be conserved during evolution, the selective pressure at a given site of a protein sequence depends on the amino acids present at other sites. However, models of molecular evolution that aim at reconstructing the evolutionary history of macromolecules become computationally intractable if such correlations between sites are explicitly taken into account. RESULTS: We introduce an evolutionary model with sites evolving independently under a global constraint on the conservation of structural stability. This model consists of a selection process, which depends on two hydrophobicity parameters that can be computed from protein sequences without any fit, and a mutation process for which we consider various models. It reproduces quantitatively the results of Structurally Constrained Neutral (SCN) simulations of protein evolution in which the stability of the native state is explicitly computed and conserved. We then compare the predicted site-specific amino acid distributions with those sampled from the Protein Data Bank (PDB). The parameters of the mutation model, whose number varies between zero and five, are fitted from the data. The mean correlation coefficient between predicted and observed site-specific amino acid distributions is larger than = 0.70 for a mutation model with no free parameters and no genetic code. In contrast, considering only the mutation process with no selection yields a mean correlation coefficient of = 0.56 with three fitted parameters. The mutation model that best fits the data takes into account increased mutation rate at CpG dinucleotides, yielding = 0.90 with five parameters. CONCLUSION: The effective selection process that we propose reproduces well amino acid distributions as observed in the protein sequences in the PDB. Its simplicity makes it very promising for likelihood calculations in phylogenetic studies. Interestingly, in this approach the mutation process influences the effective selection process, i.e. selection and mutation must be entangled in order to obtain effectively independent sites. This interdependence between mutation and selection reflects the deep influence that mutation has on the evolutionary process: The bias in the mutation influences the thermodynamic properties of the evolving proteins, in agreement with comparative studies of bacterial proteomes, and it also influences the rate of accepted mutations.

Algorithms↗

Detecting compensatory covariation signals in protein evolution using reconstructed ancestral sequences.

When protein sequences divergently evolve under functional constraints, some individual amino acid replacements that reverse the charge (e.g. Lys to Asp) may be compensated by a replacement at a second position that reverses the charge in the opposite direction (e.g. Glu to Arg). When these side-chains are near in space (proximal), such double replacements might be driven by natural selection, if either is selectively disadvantageous, but both together restore fully the ability of the protein to contribute to fitness (are together "neutral"). Accordingly, many have sought to identify pairs of positions in a protein sequence that suffer compensatory replacements, often as a way to identify positions near in space in the folded structure. A "charge compensatory signal" might manifest itself in two ways. First, proximal charge compensatory replacements may occur more frequently than predicted from the product of the probabilities of individual positions suffering charge reversing replacements independently. Conversely, charge compensatory pairs of changes may be observed to occur more frequently in proximal pairs of sites than in the average pair. Normally, charge compensatory covariation is detected by comparing the sequences of extant proteins at the "leaves" of phylogenetic trees. We show here that the charge compensatory signal is more evident when it is sought by examining individual branches in the tree between reconstructed ancestral sequences at nodes in the tree. Here, we find that the signal is especially strong when the positions pairs are in a single secondary structural unit (e.g. alpha helix or beta strand) that brings the side-chains suffering charge compensatory covariation near in space, and may be useful in secondary structure prediction. Also, "node-node" and "node-leaf" compensatory covariation may be useful to identify the better of two equally parsimonious trees, in a way that is independent of the mathematical formalism used to construct the tree itself. Further, compensatory covariation may provide a signal that indicates whether an episode of sequence evolution contains more or less divergence in functional behavior. Compensatory covariation analysis on reconstructed evolutionary trees may become a valuable tool to analyze genome sequences, and use these analyses to extract biomedically useful information from proteome databases.

Amino Acid Sequence↗

Gene family phylogenetics: tracing protein evolution on trees.

How have proteins taken on the remarkable diversity of biochemical and physiological functions necessary to create and maintain complex organisms? The majority of proteins are organized hierarchically into families and superfamilies, reflecting an ancient and continuing process of gene duplication and divergence. The techniques of molecular phylogenetics, developed to recover the nested hierarchy of taxa from character information in their gene sequences, can also reconstruct the evolutionary relationships among genes and provide a conceptual foundation for comparative evolutionary analysis of proteins and their functions. In this review, I outline the application of phylogenetic approaches to issues in gene family studies, beginning with the inference of phylogeny and the assessment of the two types of homology by which genes in a family can be related: orthology (common descent from a cladogenetic event) and paralogy (common descent from a gene duplication event). I show how the phylogenetic approach makes possible novel kinds of comparative analysis, including detection of exon shuffling, reconstruction of the evolutionary diversification of gene families, tracing of evolutionary change in protein function at the amino acid level, and prediction of structure-function relationships. A marriage of the principles of phylogenetic systematics with the copious sequence data being generated by molecular biology and genomics promises unprecedented insights into the nature of biological organization and the historical processes that created it.

Animals↗

Models of amino acid substitution and applications to mitochondrial protein evolution.

Models of amino acid substitution were developed and compared using maximum likelihood. Two kinds of models are considered. "Empirical" models do not explicitly consider factors that shape protein evolution, but attempt to summarize the substitution pattern from large quantities of real data. "Mechanistic" models are formulated at the codon level and separate mutational biases at the nucleotide level from selective constraints at the amino acid level. They account for features of sequence evolution, such as transition-transversion bias and base or codon frequency biases, and make use of physicochemical distances between amino acids to specify nonsynonymous substitution rates. A general approach is presented that transforms a Markov model of codon substitution into a model of amino acid replacement. Protein sequences from the entire mitochondrial genomes of 20 mammalian species were analyzed using different models. The mechanistic models were found to fit the data better than empirical models derived from large databases. Both the mutational distance between amino acids (determined by the genetic code and mutational biases such as the transition-transversion bias) and the physicochemical distance are found to have strong effects on amino acid substitution rates. A significant proportion of amino acid substitutions appeared to have involved more than one codon position, indicating that nucleotide substitutions at neighboring sites may be correlated. Rates of amino acid substitution were found to be highly variable among sites.

Amino Acid Substitution↗

Protein evolution in the context of Drosophila development.

The tempo at which a protein evolves depends not only on the rate at which mutations arise but also on the selective effects that those mutations have at the organismal level. It is intuitive that proteins functioning during different stages of development may be predisposed to having mutations of different selective effects. For example, it has been hypothesized that changes to proteins expressed during early development should have larger phenotypic consequences because later stages depend on them. Conversely, changes to proteins expressed much later in development should have smaller consequences at the organismal level. Here we assess whether proteins expressed at different times during Drosophila development vary systematically in their rates of evolution. We find that proteins expressed early in development and particularly during mid-late embryonic development evolve unusually slowly. In addition, proteins expressed in adult males show an elevated evolutionary rate. These two trends are independent of each other and cannot be explained by peculiar rates of mutation or levels of codon bias. Moreover, the observed patterns appear to hold across several functional classes of genes, although the exact developmental time of the slowest protein evolution differs among each class. We discuss our results in connection with data on the evolution of development.

Algorithms↗

Protein evolution: intrinsic preferences in peptide bond formation: a computational and experimental analysis.

Two possibilities exist for the evolution of individual enzymes/proteins from a milieu of amino acids, one based on preference and selectivity and the other on the basis of random events. Logic is overwhelmingly in favour of the former. By protein data base analysis and experiments, we have provided data to show the manifestation of two types of preferences, namely, the choice of the neighbour and its acceptance from the amino end (left) or the carboxyl end (right). The study tends to show that if the 20 proteinous amino acids were made to combine in water, the resulting profile would be nonrandom. Such selectivity could be a factor in protein evolution.

Evolution, Molecular↗

Prion protein: evolution caught en route.

The prion protein displays a unique structural ambiguity in that it can adopt multiple stable conformations under physiological conditions. In our view, this puzzling feature resulted from a sudden environmental change in evolution when the prion, previously an integral membrane protein, got expelled into the extracellular space. Analysis of known vertebrate prions unveils a primordial transmembrane protein encrypted in their sequence, underlying this relocalization hypothesis. Apparently, the time elapsed since this event was insufficient to create a "minimally frustrated" sequence in the new milieu, probably due to the functional constraints set by the importance of the very flexibility that was created in the relocalization. This scenario may explain why, in a structural sense, the prion protein is still en route toward becoming a foldable globular protein.

Evolution, Molecular↗

A universal trend of amino acid gain and loss in protein evolution.

Amino acid composition of proteins varies substantially between taxa and, thus, can evolve. For example, proteins from organisms with (G + C)-rich (or (A + T)-rich) genomes contain more (or fewer) amino acids encoded by (G + C)-rich codons. However, no universal trends in ongoing changes of amino acid frequencies have been reported. We compared sets of orthologous proteins encoded by triplets of closely related genomes from 15 taxa representing all three domains of life (Bacteria, Archaea and Eukaryota), and used phylogenies to polarize amino acid substitutions. Cys, Met, His, Ser and Phe accrue in at least 14 taxa, whereas Pro, Ala, Glu and Gly are consistently lost. The same nine amino acids are currently accrued or lost in human proteins, as shown by analysis of non-synonymous single-nucleotide polymorphisms. All amino acids with declining frequencies are thought to be among the first incorporated into the genetic code; conversely, all amino acids with increasing frequencies, except Ser, were probably recruited late. Thus, expansion of initially under-represented amino acids, which began over 3,400 million years ago, apparently continues to this day.

AT Rich Sequence↗

Extensive amino acid polymorphism at the pgm locus is consistent with adaptive protein evolution in Drosophila melanogaster.

PGM plays a central role in the glycolytic pathway at the branch point leading to glycogen metabolism and is highly polymorphic in allozyme studies of many species. We have characterized the nucleotide diversity across the Pgm gene in Drosophila melanogaster and D. simulans to investigate the role that protein polymorphism plays at this crucial metabolic branch point shared with several other enzymes. Although D. melanogaster and D. simulans share common allozyme mobility alleles, we find these allozymes are the result of many different amino acid changes at the nucleotide level. In addition, specific allozyme classes within species contain several amino acid changes, which may explain the absence of latitudinal clines for PGM allozyme alleles, the lack of association of PGM allozymes with the cosmopolitan In(3L)P inversion, and the failure to detect differences between PGM allozymes in functional studies. We find a significant excess of amino acid polymorphisms within D. melanogaster when compared to the complete absence of fixed replacements with D. simulans. There is also strong linkage disequilibrium across the 2354 bp of the Pgm locus, which may be explained by a specific amino acid haplotype that is high in frequency yet contains an excess of singleton polymorphisms. Like G6pd, Pgm shows strong evidence for a branch point enzyme that exhibits adaptive protein evolution.

Adaptation, Physiological↗