Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Protein evolution within a structural space.

Understanding of the evolutionary origins of protein structures represents a key component of the understanding of molecular evolution as a whole. Here we seek to elucidate how the features of an underlying protein structural "space" might impact protein structural evolution. We approach this question using lattice polymers as a completely characterized model of this space. We develop a measure of structural comparison of lattice structures that is analogous to the one used to understand structural similarities between real proteins. We use this measure of structural relatedness to create a graph of lattice structures and compare this graph (in which nodes are lattice structures and edges are defined using structural similarity) to the graph obtained for real protein structures. We find that the graph obtained from all compact lattice structures exhibits a distribution of structural neighbors per node consistent with a random graph. We also find that subgraphs of 3500 nodes chosen either at random or according to physical constraints also represent random graphs. We develop a divergent evolution model based on the lattice space which produces graphs that, within certain parameter regimes, recapitulate the scale-free behavior observed in similar graphs of real protein structures.

Amino Acid Sequence↗

Evolution of eutherian cytochrome c oxidase subunit II: heterogeneous rates of protein evolution and altered interaction with cytochrome c.

Cytochrome c oxidase subunit II (COII), encoded by the mitochondrial genome, exhibits one of the most heterogeneous rates of amino acid replacement among placental mammals. Moreover, it has been demonstrated that cytochrome c oxidase has undergone a structural change in higher primates which has altered its physical interaction with cytochrome c. We collected a large data set of COII sequences from several orders of mammals with emphasis on primates, rodents, and artiodactyls. Using phylogenetic hypotheses based on data independent of the COII gene, we demonstrated that an increased number of amino acid replacements are concentrated among higher primates. Incorporating approximate divergence dates derived from the fossil record, we find that most of the change occurred independently along the New World monkey lineage and in a rapid burst before apes and Old World monkeys diverged. There is some evidence that Old World monkeys have undergone a faster rate of nonsynonymous substitution than have apes. Rates of substitution at four-fold degenerate sites in primates are relatively homogeneous, indicating that the rate heterogeneity is restricted to nondegenerate sites. Excluding the rate acceleration mentioned above, primates, rodents, and artiodactyls have remarkably similar nonsynonymous replacement rates. A different pattern is observed for transversions at four-fold degenerate sites, for which rodents exhibit a higher rate of replacement than do primates and artiodactyls. Finally, we hypothesize specific amino acid replacements which may account for much of the structural difference in cytochrome c oxidase between higher primates and other mammals.

Amino Acids↗

Evolution of the autosomal chorion cluster in Drosophila. IV. The Hawaiian Drosophila: rapid protein evolution and constancy in the rate of DNA divergence.

Autosomal chorion genes s18, s15, and s19 are shown to diverge at extremely rapid rates in closely related taxa of Hawaiian Drosophila. Their nucleotide divergence rates are at least as fast as those of intergenic regions that are known to evolve more extensively between distantly related species. Their amino acid divergence rates are the fastest known to date. There are two nucleotide replacement substitutions for every synonymous one. The molecular basis for observed length and substitution mutations is analyzed. Length mutations are strongly associated with direct repeats in general, and with tandem repeats in particular, whereas the rate for an average transition is twice that for an average transversion. The DNA sequence of the cluster was used to construct a phylogenetic tree for five taxa of the Hawaiian picture-winged species group of Drosophila. Assignment of observed base substitutions occurring in various branches of the tree reveals an excess of would-be homoplasies in a centrally localized 1.8-kb segment containing the s15 gene. This observation may be a reflection of ancestral excess polymorphisms in the segment. The chorion cluster appears to evolve at a constant rate regardless of whether the central 1.8-kb segment is included or not in the analysis. Assuming that the time of divergence of Drosophila grimshawi and the planitibia subgroup coincides with the emergence of the island of Kauai, the overall rate of base substitution in the cluster is estimated to be 0.8% million years, whereas synonymous sites are substituted at a rate of 1.2% million years.

Amino Acid Sequence↗

Penicillin and beyond: evolution, protein fold, multimodular polypeptides, and multiprotein complexes.

As the protein sequence and structure databases expand, the relationships between proteins, the notion of protein superfamily, and the driving forces of evolution are better understood. Key steps of the synthesis of the bacterial cell wall peptidoglycan are revisited in light of these advances. The reactions through which the D-alanyl-D-alanine depeptide is formed, utilized, and hydrolyzed and the sites of action of the glycopeptide and beta-lactam antibiotics illustrate the concept according to which new enzyme functions evolve as a result of tinkering of existing proteins. This occurs by the acquisition of local structural changes, the fusion into multimodular polypeptides, and the association into multiprotein complexes.

Bacterial Proteins↗

[Protein evolution rate and immunoglobulin induction].

The capacity of proteins to induce the synthesis of specific immunoglobulins was shown to be correlated with their evolution rate. This correlation can be understood in terms of the clonal-selectional theory. The characteristic correlation parameter is the number of differences in the amino acid sequences between the immunogenic protein and the homological protein of the immunized animal. The above correlation was traced most clearly for the evolutionary conservative proteins.

Actins↗

Adaptive protein evolution and regulatory divergence in Drosophila.

Two recent studies demonstrated a positive correlation between divergence in gene expression and protein sequence in Drosophila. This correlation could be driven by positive selection or variation in functional constraint. To distinguish between these alternatives, we compared patterns of molecular evolution for 1,862 genes with two previously reported estimates of expression divergence in Drosophila. We found a slight negative trend (nonsignificant) between positive selection on protein sequence and divergence in expression levels between Drosophila melanogaster and Drosophila simulans. Conversely, shifts in expression patterns during Drosophila development showed a positive association with adaptive protein evolution, though as before the relationship was weak and not significant. Overall, we found no strong evidence for an increase in the incidence of positive selection on protein-coding regions in genes with divergent expression in Drosophila, suggesting that the previously reported positive association between protein and regulatory divergence primarily reflects variation in functional constraint.

Amino Acid Sequence↗

Local-scale repetitiveness in amino acid use in eukaryote protein sequences: a genomic factor in protein evolution.

We showed previously that the use of arginine versus lysine residues in eukaryote proteins is correlated positively with local GC content of the genome within approximately 50 residues. Cumulative analyses show that the tendency for self-clustering (or repetitive use) generally is the case for all types of amino acids except for certain hydrophobic types. The degree to which each of the amino acids is used recurrently is weak for ancient proteins (or protein domains), those that are conserved through both eukaryotes and prokaryotes, but strong for modern proteins, which are unique to organisms of particular phyla. These findings support the idea that repetitiveness occurs due to a propensity of genomic DNA to cause tandem genomic duplication. A protein sequence with high repetitiveness tends to be unique in the homology search, which may indicate the weaker constraints and, hence, more arbitrary use of amino acids. Simulation analyses suggest that tandem gene duplications on a very small scale (1 or 2 codons) is an important causal factor in maintaining repetitiveness in the presence of concomittant occurrence of substitutive point mutation. For yeast proteins, approximately 1.3 duplication events per 1,000 residues on average are likely to occur, whereas 10 events of substitution mutation occur. It also is suggested that duplication enhances the probability of occurrence of some peptide motifs, such as those found in zinc fingers and segments with extreme physicochemical characteristics, and, thus, that local repetitiveness is a genomic factor influencing the evolution of eukaryote proteins.

Amino Acid Motifs↗

Rational evolutionary design: the theory of in vitro protein evolution.

Directed evolution uses a combination of powerful search techniques to generate proteins with improved properties. Part of the success is due to the stochastic element of random mutagenesis; improvements can be made without a detailed description of the complex interactions that constitute function or stability. However, optimization is not a conglomeration of random processes. Rather, it requires both knowledge of the system that is being optimized and a logical series of techniques that best explores the pathways of evolution (Eigen et al., 1988). The weighing of parameters associated with mutation, recombination, and screening to achieve the maximum fitness improvement is the beginning of rational evolutionary design. The optimal mutation rate is strongly influenced by the finite number of mutants that can be screened. A smooth fitness landscape implies that many mutations can be accumulated without disrupting the fitness. This has the effect of lowering the required library size to sample a higher mutation rate. As the sequence ascends the fitness landscape, the optimal mutation rate decreases as the probability of discovering improved mutations also decreases. Highly coupled regions require that many mutations be simultaneously made to generate a positive mutant. Therefore, positive mutations are discovered at uncoupled positions as the fitness of the parent increases. The benefit of recombination is twofold: it combines good mutations and searches more sequence space in a meaningful way. Recombination is most beneficial when the number of mutants that can be screened is limited and the landscape is of an intermediate ruggedness. The structure of schema in proteins leads to the conclusion that many cut points are required. The number of parents and their sequence identity are determined by the balance between exploration and exploitation. Many disparate parents can explore more space, but at the risk of losing information. The required screening effort is related to the number of uphill paths, which decreases more rapidly for rugged landscapes. Noise in the fitness measurements causes a dramatic increase in the required mutant library size, thus implying a smaller optimal mutation rate. Because of strict limitations on the number of mutants that can be screened, there is motivation to optimize the content of the mutant library. By restricting mutations to regions of the gene that are expected to show improvement, a greater return can be made with the same number of mutants. Initial studies with subtilisin E have shown that structurally tolerant positions tend to be where positive activity mutants are made during directed evolution. Mutant fitness information is produced by the screening step that has the potential to provide insight into the structure of the fitness landscape, thus aiding the setting of experimental parameters. By analyzing the mutant fitness distribution and targeting specific regions of the sequence, in vitro evolution can be accelerated. However, when expediting the search, there is a trade-off between rapid improvement and the quality of the long-term solution. The benefit of neutrality has yet to be captured with in vitro protein evolution. Neutral theory predicts the punctuated emergence of novel structure and function, however, with current methods, the required time scale is not feasible. Utilizing neutral evolution to accelerate the discovery of new functional and structural solutions requires a theory that predicts the behavior of mutational pathways between networks. Because the transition from neutral to adaptive evolution requires a multi-mutational switch, increasing the mutation rate decreases the time required for a punctuated change to occur. By limiting the search to the less coupled region of the sequence (smooth portion of the fitness landscape), the required larger mutation rate can be tolerated. Advances in directed evolution will be achieved when the driving forces behind such proce

Evolution, Molecular↗

Protein evolution: structure-function relationships of the oncogene beta-catenin in the evolution of multicellular animals.

Beta-catenin functions as a cytoskeletal linker protein in cadherin-mediated adhesion and as a signal mediator in wnt-signal transduction pathways. We use a novel integrative approach, combining evolutionary, genomic, and three-dimensional structural data to analyze and trace the structural and functional evolution of beta-catenin genes. This approach also enabled us to examine the effects of gene duplication on the structure and function of beta-catenin genes in Drosophila, C. elegans, and vertebrates. By sampling a large number of different taxa, we identified both ancestral and derived motifs and residues within the different regions of the beta-catenin proteins. Projecting amino acid substitutions onto the three- dimensional structure established for mouse beta-catenin, we identified specific domains that exhibit loss and gain of selective constraints during beta catenin evolution. Structural changes, changes in the amino acid substitution rate, and the appearance of novel functional domains in beta-catenin can be mapped to specific branches on the metazoan tree. Together, our analyses suggest that a single, beta-catenin gene fulfilled both adhesion and signaling functions in the last common ancestor of metazoans some 700 million years ago. In addition, gene duplications facilitated the evolution of beta-catenins with novel functions and allowed the evolution of multiple, single-function proteins (cell adhesion or wnt-signaling) from the ancestral, dual-function protein. Integrative methods such as those we have applied here, utilizing the 'natural experiments' present in animal diversity, can be employed to identify novel and shared functional motifs and residues in virtually any protein among the proteomes of model systems and humans.

Amino Acid Motifs↗

Combining protein evolution and secondary structure.

An evolutionary model that combines protein secondary structure and amino acid replacement is introduced. It allows likelihood analysis of aligned protein sequences and does not require the underlying secondary (or tertiary) structures of these sequences to be known. One component of the model describes the organization of secondary structure along a protein sequence and another specifies the evolutionary process for each category of secondary structure. A database of proteins with known secondary structures is used to estimate model parameters representing these two components. Phylogeny, the third component of the model, can be estimated from the data set of interest. As an example, we employ our model to analyze a set of sucrose synthase sequences. For the evolution of sucrose synthase, a parametric bootstrap approach indicates that our model is statistically preferable to one that ignores secondary structure.

Amino Acid Sequence↗

Clustering of tissue-specific genes underlies much of the similarity in rates of protein evolution of linked genes.

Are genes nonrandomly distributed around the genome and might this explain why it was found that, in the mouse genome, proteins of linked genes evolve at similar rates? Anecdotal evidence suggests that the similarity of expression of linked genes might, in part, explain the similarity in their rates of evolution. Immune system genes, for example, are known to evolve at a high rate and sometimes cluster in the genome. Here we develop methods for statistical tests of similarity of expression of linked genes and report that there is a significant tendency for genes of similar expression breadth to be linked. Significantly, when we exclude tissue specific genes from our sample, the similarity in rates of protein evolution of linked genes is greatly diminished, if not abolished. This diminution is not a sampling artifact. In contrast, while half of the immune genes in our sample reside in 1 of 10 immune clusters in the mouse genome, this clustering appears not to affect the extent of local similarity in rates of evolution. The distribution of placentally expressed genes, in contrast, does have an effect.

Animals↗

Structural divergence and distant relationships in proteins: evolution of the globins.

The globin family has long been known from studies of approximately 150-residue proteins such as vertebrate myoglobins and haemoglobins. Recently, this family has been enriched by the investigation of the sequences and structures of truncated globins, which have the same basic topology but are approximately 30 residues shorter and exhibit functions other than the familiar one of binding diatomic ligands. The divergence of protein sequences, structures and functions reveals Nature's exploration of the potential inherent in a folding pattern, that is, the topology of the native structure. The observation of what remains constant and what varies during the evolution of a protein family reveals essential features of structure and function. Study of proteins with a wide range of divergence can therefore sharpen our understanding of how different amino acid sequences can determine similar three-dimensional structures. Globins have provided, and continue to provide, interesting material for such studies.

Amino Acid Sequence↗

Tracing pathways of transport protein evolution.

We have conducted bioinformatic analyses of integral membrane transport proteins belonging to dozens of families. These families rarely include proteins that function in a capacity other than transport. Many transporters have arisen by intragenic duplication, triplication and quadruplication events, in which the numbers of transmembrane alpha-helical hydrophobic segments (TMSs) have increased. The elements multiplied may encode two, three, four, five, six, 10 or 12 TMSs and gave rise to proteins with four, six, seven, eight, nine, 10, 12, 20, 24 and 30 TMSs. Gene fusion, splicing, deletion and insertion events have also contributed to protein topological diversity. Amino acid substitutions have allowed membrane-embedded domains to become hydrophilic domains and vice versa. Some evidence suggests that amino acid substitutions occurring over evolutionary time may in some cases have drastically altered protein topology. The results summarized in this microreview establish the independent origins of many transporter families and allow postulation of the specific pathways taken for their appearance.

Computational Biology↗

Trends in protein evolution inferred from sequence and structure analysis.

Complementary developments in comparative genomics, protein structure determination and in-depth comparison of protein sequences and structures have provided a better understanding of the prevailing trends in the emergence and diversification of protein domains. The investigation of deep relationships among different classes of proteins involved in key cellular functions, such as nucleic acid polymerases and other nucleotide-dependent enzymes, indicates that a substantial set of diverse protein domains evolved within the primordial, ribozyme-dominated RNA world.

Evolution, Molecular↗

Parsimony in protein evolution.

Pohl found the activation enthalpy and entropy for melting of his mesophilic proteins to be linear in the total number of residues and Privalov and colleagues found this same linearity for the standard heat-capacity, enthalpy and entropy changes in the overall melting equilibria. Despite the small samples these results suggest that mesophiles individually, and as a class, are related through a single standard representative. If so, very extensive convergent evolution has provided both great simplification and very sophisticated goals for genome decoding and quantitative description of protein substructures [R. Lumry, The protein primer, http://www.umn.edu.chem. /groupslumry].

Evolution, Molecular↗

Mutational bias affects protein evolution in flowering plants.

Amino acid sequences from several thousand homologous gene pairs were compared for two plant genomes, Oryza sativa and Arabidopsis thaliana. The Arabidopsis genes all have similar G+C (guanine plus cytosine) contents, whereas their homologs in rice span a wide range of G+C levels. The results show that those rice genes that display increased divergence in their nucleotide composition (specifically, increased G+C content) showed a corresponding, predictable change in the amino acid compositions of the encoded proteins relative to their Arabidopsis homologs. This trend was not seen in a "control" set of rice genes that had nucleotide contents closer to their Arabidopsis homologs. In addition to showing an overall difference in the amino acid composition of the homologous proteins, we were also able to investigate the biased patterns of amino acid substitution since the divergence of these two species. We found that the amino acid exchange matrix was highly asymmetric when comparing the High G+C rice genes with their Arabidopsis homologs. Finally, we investigated the possible causes of this biased pattern of sequence evolution. Our results indicate that the biased pattern of protein evolution is the consequence, rather than the cause, of the corresponding changes in nucleotide content. In fact, there is an even more marked asymmetry in the patterns of substitution at synonymous nucleotide sites. Surprisingly, there is a very strong negative correlation between the level of nucleotide bias and the length of the coding sequences within the rice genome. This difference in gene length may provide important clues about the underlying mechanisms.

Amino Acids↗