Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Protein evolution with dependence among codons due to tertiary structure.

Markovian models of protein evolution that relax the assumption of independent change among codons are considered. With this comparatively realistic framework, an evolutionary rate at a site can depend both on the state of the site and on the states of surrounding sites. By allowing a relatively general dependence structure among sites, models of evolution can reflect attributes of tertiary structure. To quantify the impact of protein structure on protein evolution, we analyze protein-coding DNA sequence pairs with an evolutionary model that incorporates effects of solvent accessibility and pairwise interactions among amino acid residues. By explicitly considering the relationship between nonsynonymous substitution rates and protein structure, this approach can lead to refined detection and characterization of positive selection. Analyses of simulated sequence pairs indicate that parameters in this evolutionary model can be well estimated. Analyses of lysozyme c and annexin V sequence pairs yield the biologically reasonable result that amino acid replacement rates are higher when the replacements lead to energetically favorable proteins than when they destabilize the proteins. Although the focus here is evolutionary dependence among codons that is associated with protein structure, the statistical approach is quite general and could be applied to diverse cases of evolutionary dependence where surrogates for sequence fitness can be measured or modeled.

Annexin A5↗

Structural convergence during protein evolution.

Several recent protein crystallographic structure determinations have demonstrated the existence of considerable tertiary structural similarity among proteins otherwise having little similarity in either amino acid sequence or biological function. In order to assess the possibility that such proteins may have arisen through processes of divergent evolution from a common ancestor, a graphical presentation is given which correlates the pattern of allowed single base substitutions defined by the genetic code with the associated changes in the structural properties of the encoded amino acids. The results show that while a large degree of structural conservation is evident due to codon synonomy, there is, in general, little tendency for the code to be structurally conservative in the majority of the cases where codon single-base changes result in amino acid substitutions. The possible consequences of this pattern of potential amino acid substitutions are discussed in relation to protein evolutionary processes.

Amino Acids↗

Extreme differences in charge changes during protein evolution.

The maintenance of a proper distribution of charged amino acid residues might be expected to be an important factor in protein evolution. We therefore compared the inferred changes in charge during the evolution of 43 protein families with the changes expected on the basis of random base substitutions. It was found that certain proteins, like the eye lens crystallins and most histones, display an extreme avoidance of changes in charge. Other proteins, like phospholipase A2 and ferredoxin, apparently have sustained more charged replacements than expected, suggesting a positive selection for changes in charge. Depending on function and structure of a protein, charged residues apparently can be important targets for selective forces in protein evolution. It appears that actual biased codon usage tends to decrease the proportion of charged amino acid replacements. The influence of nonrandomness of mutations is more equivocal. Genes that use the mitochondrial instead of the universal code lower the probability that charge changes will occur in the encoded proteins.

Biological Evolution↗

Early protein evolution: building domains from ligand-binding polypeptide segments.

It has been suggested that in the early evolution of proteins, segments of polypeptide, unable to fold in isolation, may have collapsed together to form folded proto-domains. We wondered whether the incorporation of segments with a pre-existing binding activity into a folded domain could, by fixing the ligand binding conformation and/or providing additional contacts, lead to large affinity improvements and provide an evolutionary advantage. As a model, we took a segment of polypeptide from hen egg lysozyme that in the native protein forms the binding interface with the monoclonal antibodies HyHEL5 and F10 (KD=60 pM). When expressed in bacteria the isolated segment was unfolded, readily proteolysed and only bound weakly to the antibodies (KD>1 microM). We then combined the segment with random genomic segments to create a repertoire of chimaeric polypeptides displayed on filamentous bacteriophage. By use of proteolysis (to select folded polypeptide) and anti-lysozyme antibodies (to select an active conformation) we isolated a folded dimeric protein with an enhanced antibody affinity (KD=400 pM). Unexpectedly the dimer also incorporated a single heme molecule (KD=33 nM) that stabilised the dimer (Tm=59 degrees C with heme, 35 degrees C without heme). These results show that the binding affinities of flexible polypeptide segments can be greatly enhanced on protein folding, and that the folding can be stabilised by prosthetic groups. This supports the hypothesis that sub-domain polypeptide segments with functional activities may have contributed to domain creation in early evolution.

Amino Acid Sequence↗

Protein evolution. On the ancestry of barrels.

Most proteins consist of several domains linked together in a single polypeptide chain, and many of these proteins have evolved by gene duplication and fusion. Miles and Davies discuss the study by Lang et al., who show that this type of protein evolution may also occur in b/a barrel proteins, a common single-domain protein fold. Other single domain proteins may have arisen from similar evolutionary mechanisms.

Aldose-Ketose Isomerases↗

X chromosomes and autosomes evolve at similar rates in Drosophila: no evidence for faster-X protein evolution.

Recent data from Drosophila suggest that a substantial fraction of amino acid substitutions observed between species are beneficial. If these beneficial mutations are on average partially recessive, then the rate of protein evolution is predicted to be faster for X-linked genes compared to autosomal genes (the "faster-X" hypothesis). We test this prediction by comparing rates of protein substitutions between orthologous genes, taking advantage of variations in chromosome fusions within the genus Drosophila. In members of the Drosophila melanogaster species group, the chromosomal arm 3L segregates as an ordinary autosome (i.e., two homologous copies in both males and females). However, in the Drosophila pseudoobscura species group, this chromosomal arm has become fused to the ancestral X chromosome and is hemizygous in males. The faster-X hypothesis predicts that protein evolution should be faster for genes on this chromosomal arm in the D. pseudoobscura lineage, relative to the D. melanogaster lineage. Here we combine new sequence data for 202 gene fragments in Drosophila miranda (in the pseudoobscura species group) with the completed genomes of D. melanogaster, D. pseudoobscura, and Drosophila yakuba to show that there are no detectable differences in rates of amino acid evolution for orthologous X-linked and autosomal genes. Our results imply that the contribution of the faster-X (if any) to the large-X effect on reproductive isolation in Drosophila is not due to a generally faster rate of protein evolution. The lack of a detectable faster-X effect in these species suggests either that beneficial amino acids are not partially recessive on average, or that adaptive evolution does not often use newly arising amino acid mutations.

Amino Acid Substitution↗

cis-Regulatory and protein evolution in orthologous and duplicate genes.

The relationship between protein and regulatory sequence evolution is a central question in molecular evolution. It is currently not known to what extent changes in gene expression are coupled with the evolution of protein coding sequences, or whether these changes differ among orthologs (species homologs) and paralogs (duplicate genes). Here, we develop a method to measure the extent of functionally relevant cis-regulatory sequence change in homologous genes, and validate it using microarray data and experimentally verified regulatory elements in different eukaryotic species. By comparing the genomes of Caenorhabditis elegans and C. briggsae, we found that protein and regulatory evolution is weakly coupled in orthologs but not paralogs, suggesting that selective pressure on gene expression and protein evolution is quite similar and persists for a significant amount of time following speciation but not gene duplication. Additionally, duplicates of both species exhibit a dramatic acceleration of both regulatory and protein evolution compared to orthologs, suggesting increased directional selection and/or relaxed selection on both gene expression patterns and protein function in duplicate genes.

Animals↗

Genome architecture drives protein evolution in ciliates.

Studies of microbial eukaryotes have been pivotal in the discovery of biological phenomena, including RNA editing, self-splicing RNA, and telomere addition. Here we extend this list by demonstrating that genome architecture, namely the extensive processing of somatic (macronuclear) genomes in some ciliate lineages, is associated with elevated rates of protein evolution. Using newly developed likelihood-based procedures for studying molecular evolution, we investigate 6 genes to compare 1) ciliate protein evolution to that of 3 other clades of eukaryotes (plants, animals, and fungi) and 2) protein evolution in ciliates with extensively processed macronuclear genomes to that of other ciliate lineages. In 5 of the 6 genes, ciliates are estimated to have a higher ratio of nonsynonymous/synonymous substitution rates, consistent with an increase in the rate of protein diversification in ciliates relative to other eukaryotes. Even more striking, there is a significant effect of genome architecture within ciliates as the most divergent proteins are consistently found in those lineages with the most highly processed macronuclear genomes. We propose a model whereby genome architecture-specifically chromosomal processing, amitosis within macronuclei, and epigenetics-allows ciliates to explore protein space in a novel manner. Further, we predict that examination of diverse eukaryotes will reveal additional evidence of the impact of genome architecture on molecular evolution.

Animals↗

Frustration and hydrophobicity interplay in protein folding and protein evolution.

A lattice model is used to study mutations and compacting effects on protein folding rates and folding temperature. In the context of protein evolution, we address the question regarding the best scenario for a polypeptide chain to fold: either a fast nonspecific collapse followed by a slow rearrangement to form the native structure or a specific collapse from the unfolded state with the simultaneous formation of the native state. This question is investigated for optimized sequences, whose native state has no frustrated contacts between monomers, and also for mutated sequences, whose native state has some degree of frustration. It is found that the best scenario for folding may depend on the amount of frustration of the native structure. The implication of this result on protein evolution is discussed.

Algorithms↗

Molecular clock in neutral protein evolution.

BACKGROUND: A frequent observation in molecular evolution is that amino-acid substitution rates show an index of dispersion (that is, ratio of variance to mean) substantially larger than one. This observation has been termed the overdispersed molecular clock. On the basis of in silico protein-evolution experiments, Bastolla and coworkers recently proposed an explanation for this observation: Proteins drift in neutral space, and can temporarily get trapped in regions of substantially reduced neutrality. In these regions, substitution rates are suppressed, which results in an overall substitution process that is not Poissonian. However, the simulation method of Bastolla et al. is representative only for cases in which the product of mutation rate micro and population size Ne is small. How the substitution process behaves when micro Ne is large is not known. RESULTS: Here, I study the behavior of the molecular clock in in silico protein evolution as a function of mutation rate and population size. I find that the index of dispersion decays with increasing micro Ne, and approaches 1 for large micro Ne. This observation can be explained with the selective pressure for mutational robustness, which is effective when micro Ne is large. This pressure keeps the population out of low-neutrality traps, and thus steadies the ticking of the molecular clock. CONCLUSIONS: The molecular clock in neutral protein evolution can fall into two distinct regimes, a strongly overdispersed one for small micro Ne, and a mostly Poissonian one for large micro Ne. The former is relevant for the majority of organisms in the plant and animal kingdom, and the latter may be relevant for RNA viruses.

Amino Acid Substitution↗

Structural constraints and emergence of sequence patterns in protein evolution.

The aim of this work was to study the relationship between structure conservation and sequence divergence in protein evolution. To this end, we developed a model of structurally constrained protein evolution (SCPE) in which trial sequences, generated by random mutations at gene level, are selected against departure from a reference three-dimensional structure. Since at the mutational level SCPE is completely unbiased, any emergent sequence pattern will be due exclusively to structural constraints. In this first report, it is shown that SCPE correctly predicts the characteristic hexapeptide motif of the left-handed parallel beta helix (LbetaH) domain of UDP-N-acetylglucosamine acyltransferases (LpxA).

Acyltransferases↗

The structurally constrained protein evolution model accounts for sequence patterns of the LbetaH superfamily.

BACKGROUND: Structure conservation constrains evolutionary sequence divergence, resulting in observable sequence patterns. Most current models of protein evolution do not take structure into account explicitly, being unsuitable for investigating the effects of structure conservation on sequence divergence. To this end, we recently developed the Structurally Constrained Protein Evolution (SCPE) model. The model starts with the coding sequence of a protein with known three-dimensional structure. At each evolutionary time-step of an SCPE simulation, a trial sequence is generated by introducing a random point mutation in the current coding DNA sequence. Then, a "score" for the trial sequence is calculated and the mutation is accepted only if its score is under a given cutoff, lambda. The SCPE score measures the distance between the trial sequence and a given reference sequence, given the structure. In our first brief report we used a "global score", in which the same reference sequence, the ancestral one, was used at each evolutionary step. Here, we introduce a new scoring function, the "local score", in which the sequence accepted at the previous evolutionary time-step is used as the reference. We assess the model on the UDP-N-acetylglucosamine acyltransferase (LPXA) family, as in our previous report, and we extend this study to all other members of the left-handed parallel beta helix fold (LbetaH) superfamily whose structure has been determined. RESULTS: We studied site-dependent entropies, amino acid probability distributions, and substitution matrices predicted by SCPE and compared with experimental data for several members of the LbetaH superfamily. We also evaluated structure conservation during simulations. Overall, SCPE outperforms JTT in the description of sequence patterns observed in structurally constrained sites. Maximum Likelihood calculations show that the local-score and global-score SCPE substitution matrices obtained for LPXA outperform the JTT model for the LPXA family and for the structurally constrained sites of class i of other members within the LbetaH superfamily. CONCLUSION: We extended the SCPE model by introducing a new scoring function, the local score. We performed a thorough assessment of the SCPE model on the LPXA family and extended it to all other members of known structure of the LbetaH superfamily.

Acyltransferases↗

Relationship between mutability, polarity and exteriority of amino acid residues in protein evolution.

A systematic study was carried out on mutability of amino acid residues in evolving proteins in relation to their polarity and location within three-dimensional structure of proteins. Exteriority of residue sites is quantitatively defined as accessibility based on their static accessible surface area to solvent water molecule. Residue sites are classified into interior and exterior depending on their accessibility. More frequent substitution on exterior sites is confirmed to be general in eight sets of homologous protein families regardless of their biological functions and of presence or absence of a prosthetic group. Virtually all types of amino acid residues are found to have higher mutabilities on the exterior than in the interior. No correlation between mutability and polarity was observed of amino acid residues in the interior and on the exterior, respectively. Amino acid residues are classified into three depending on their polarity, polar (Arg, Lys, His, Gln, Asn, Asp and Glu), weak polar (Ala, Pro, Gly, Thr and Ser) and nonpolar (Cys, Val, Met, Ile, Leu, Phe, Tyr and Trp). Amino acid replacements during protein evolution are very conservative; 88% and 76% of them in the interior and on the exterior, respectively, are within the same group of the three. Inter-group replacements are such that weak polar residues are replaced more often by nonpolar residues in the interior and more often by polar residues on the exterior.

Amino Acids↗

Simulating protein evolution in sequence and structure space.

Naturally occurring proteins comprise a special subset of all plausible sequences and structures selected through evolution. Simulating protein evolution with simplified and all-atom models has shed light on the evolutionary dynamics of protein populations, the nature of evolved sequences and structures, and the extent to which today's proteins are shaped by selection pressures on folding, structure and function. Extensive mapping of the native structure, stability and folding rate in sequence space using lattice proteins has revealed organizational principles of the sequence/structure map important for evolutionary dynamics. Evolutionary simulations with lattice proteins have highlighted the importance of fitness landscapes, evolutionary mechanisms, population dynamics and sequence space entropy in shaping the generic properties of proteins. Finally, evolutionary-like simulations with all-atom models, in particular computational protein design, have helped identify the dominant selection pressures on naturally occurring protein sequences and structures.

Algorithms↗

Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.

BACKGROUND: The functional repertoire of the human proteome is an incremental collection of functions accomplished by protein domains evolved along the Homo sapiens lineage. Therefore, knowledge on the origin of these functionalities provides a better understanding of the domain and protein evolution in human. The lack of proper comprehension about such origin has impelled us to study the evolutionary origin of human proteome in a unique way as detailed in this study. RESULTS: This study reports a unique approach for understanding the evolution of human proteome by tracing the origin of its constituting domains hierarchically, along the Homo sapiens lineage. The uniqueness of this method lies in subtractive searching of functional and conserved domains in the human proteome resulting in higher efficiency of detecting their origins. From these analyses the nature of protein evolution and trends in domain evolution can be observed in the context of the entire human proteome data. The method adopted here also helps delineate the degree of divergence of functional families occurred during the course of evolution. CONCLUSION: This approach to trace the evolutionary origin of functional domains in the human proteome facilitates better understanding of their functional versatility as well as provides insights into the functionality of hypothetical proteins present in the human proteome. This work elucidates the origin of functional and conserved domains in human proteins, their distribution along the Homo sapiens lineage, occurrence frequency of different domain combinations and proteome-wide patterns of their distribution, providing insights into the evolutionary solution to the increased complexity of the human proteome.

Animals↗

Adaptive protein evolution at the Adh locus in Drosophila.

Proteins often differ in amino-acid sequence across species. This difference has evolved by the accumulation of neutral mutations by random drift, the fixation of adaptive mutations by selection, or a mixture of the two. Here we propose a simple statistical test of the neutral protein evolution hypothesis based on a comparison of the number of amino-acid replacement substitutions to synonymous substitutions in the coding region of a locus. If the observed substitutions are neutral, the ratio of replacement to synonymous fixed differences between species should be the same as the ratio of replacement to synonymous polymorphisms within species. DNA sequence data on the Adh locus (encoding alcohol dehydrogenase, EC 1.1.1.1) in three species in the Drosophila melanogaster species subgroup do not fit this expectation; instead, there are more fixed replacement differences between species than expected. We suggest that these excess replacement substitutions result from adaptive fixation of selectively advantageous mutations.

Alcohol Dehydrogenase↗

Accelerated rates of intron gain/loss and protein evolution in duplicate genes in human and mouse malaria parasites.

Very little is known about molecular evolution in the human malaria parasite Plasmodium falciparum. Given the potentially important role that introns play in directing transcription and the posttranscriptional control of gene expression, we compare rates of intron/gain loss and intronic substitution in P. falciparum and the rodent malaria P. y. yoelii in both orthologous and duplicate genes. Specifically, we test the hypothesis that intron gain/loss and protein evolution is accelerated in duplicate genes versus orthologous genes in both parasites using the genome sequence of both species. We find that duplicate genes in both P. falciparum and P. y. yoelii exhibit a dramatic acceleration of both intron gain/loss and protein evolution in comparison with orthologs, suggesting increased directional and/or relaxed selection in duplicate genes. Further, we find that rates of intron gain/loss and protein evolution are weakly coupled in orthologs but not paralogs, supporting the hypothesis that selection acts on genes as functionally integrated units after speciation but not necessarily after gene duplication. In contrast, we find that rates of nucleotide substitution do not differ significantly between intronic sites and synonymous sites among duplicate genes, implying that a large fraction of intronic sites in Plasmodium evolve under little or no selective constraint.

Animals↗

The effect of high-frequency random mutagenesis on in vitro protein evolution: a study on TEM-1 beta-lactamase.

For a number of years a major limitation in genetic analysis of protein function has been the inability to introduce multiple substitutions at distant sites that would enable the selection of clusters of mutations required for improved or novel biological functions. In order to achieve this, we have recently developed a novel mutagenesis procedure in which the triphosphate derivatives of a pyrimidine (6-(2-deoxy-beta-d-ribofuranosyl)-3, 4-dihydro-8H-pyrimido-[4,5-c][1,2]oxazin-7-one; dP) and a purine (8-oxo-2'-deoxyguanosine; 8-oxodG) nucleoside analogue are employed in DNA synthesis reactions in vitro. The procedure allows control of the mutational load and can yield frequencies of amino acid residue substitutions at least one order of magnitude greater than those previously achieved. Here we report the results of an experiment in which we have hypermutated the bacterial enzyme TEM-1 beta-lactamase and selected small pools (<1.5x10(5)) of clones for enzymatic activity against the beta-lactam antibiotic cefotaxime. The experiment resulted in the isolation of a number of TEM-1 mutants with greatly improved activity against cefotaxime. Among these, clone 3D.5 (E104K:M182T:G238S) exhibited a minimum inhibitory concentration for cefotaxime 20,000-fold higher than wild-type TEM-1 and a catalytic efficiency (kcat/Km) 2383 times higher than the wild-type enzyme. Thus, small pools of hypermutated sequences enabled the selection of one of the most active extended beta-lactamases described so far. These results argue against the accepted view that multiple rounds of low-rate mutagenesis and stepwise selection are essential for in vitro protein evolution and extend the scope of directed molecular evolution to proteins for which no genetic selection is available.

Cefotaxime↗