Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Advances in directed protein evolution by recursive genetic recombination: applications to therapeutic proteins.

Recent developments in directed evolution technologies combined with innovations in robotics and screening methods have revolutionized protein engineering. These methods are being applied broadly to many fields of biotechnology, including chemical engineering, agriculture and human therapeutics. More specifically, DNA shuffling and other methods of genetic recombination and mutation have resulted in the improvement of proteins of therapeutic interest. Optimizing genetic diversity and fitness through iterative directed evolution will accelerate improvements in engineered protein therapeutics.

Antibodies↗

Analysis of sequence periodicity in E. coli proteins: empirical investigation of the "duplication and divergence" theory of protein evolution.

Periodicity was quantified in 4289 Escherichia coli K12 confirmed and putative protein sequences, using a simple chi-square technique previously shown to reveal triplet period periodicity in coding DNA. Periodicities were calculated from period n = 2 to period n = 50 in nine different alphabetic representations of the proteins. By comparison with a randomly generated proteome of the same compositional content, the E. coli proteome does not contain a significant excess of periodic proteins. However, 60 proteins do appear to be significantly periodic in at least one alphabetic representation, after Bonferroni correction, at p < 0.01, and 30 at p < 0.001. These are compared with significantly periodic proteins of solved three-dimensional structure, detected by an identical analysis of the sequences from a protein structure database. It is concluded that there is no evidence for the presence of a proteome-wide quasi-periodicity as predicted by the "duplication and divergence" model of protein evolution and that the major periodicity detected is a consequence of the repetitive tendencies within alpha-helices. However, it is not possible to explain all sequence periodicities in terms of observable secondary structure, as in cases where sequence periodicity can be compared to solved structure, there is often no structural regularity that would provide an obvious explanation in terms of natural selection on protein function.

Databases, Protein↗

Modeling mitochondrial protein evolution using structural information.

We present two new models of protein sequence evolution based on structural properties of mitochondrial proteins. We compare these models with others currently used in phylogenetic analyses, investigating their performance over both short and long evolutionary distances. We find that our models that incorporate secondary structure information from mitochondrial proteins are statistically comparable with existing models when studying 13 mitochondrial protein data sets from eutherian mammals. However, our models give a significantly improved description of the evolutionary process when used with 12 mitochondrial proteins from a broader range of organisms including fungi, plants, protists, and bacteria. Our models may thus be of use in estimating mitochondrial protein phylogenies and for the study of processes of mitochondrial protein evolution, in particular for distantly related organisms.

Evolution, Molecular↗

Assessing the impact of secondary structure and solvent accessibility on protein evolution.

Empirically derived models of amino acid replacement are employed to study the association between various physical features of proteins and evolution. The strengths of these associations are statistically evaluated by applying the models of protein evolution to 11 diverse sets of protein sequences. Parametric bootstrap tests indicate that the solvent accessibility status of a site has a particularly strong association with the process of amino acid replacement that it experiences. Significant association between secondary structure environment and the amino acid replacement process is also observed. Careful description of the length distribution of secondary structure elements and of the organization of secondary structure and solvent accessibility along a protein did not always significantly improve the fit of the evolutionary models to the data sets that were analyzed. As indicated by the strength of the association of both solvent accessibility and secondary structure with amino acid replacement, the process of protein evolution-both above and below the species level-will not be well understood until the physical constraints that affect protein evolution are identified and characterized.

Databases, Factual↗

On the PAM matrix model of protein evolution.

The internal consistency of the PAM matrix model of protein evolution is here investigated. The 1 PAM matrix has been constructed from amino acid replacements observed in closely related sequences. Such replacements are of two types, those that do not require an intermediate amino acid replacement and those that do. The second type of replacement must generally be produced by a repetition of the first. This allows data on the first type to be used in predicting data on the second type so that some elements of the 1 PAM matrix may be used to predict others. A discrepancy of more than two orders of magnitude is found between the predictions and the data when this is carried out. This is partly accounted for by an error in constructing the matrix. However, it also seems necessary that the basic model be modified. Several possibilities are considered. One of these is to incorporate a site-dependent spectrum of mutabilities associated with each amino acid.

Amino Acid Sequence↗

Rates of protein evolution are positively correlated with developmental timing of expression during mouse spermatogenesis.

Male reproductive genes often evolve very rapidly, and sexual selection is thought to be a primary force driving this divergence. We investigated the molecular evolution of 987 genes expressed at different times during mouse spermatogenesis to determine if the rate of evolution and the intensity of positive selection vary across stages of male gamete development. Using mouse-rat orthologs, we found that rates of protein evolution were positively correlated with the developmental timing of expression. Genes expressed early in spermatogenesis had rates of divergence similar to the genome median, while genes expressed after the onset of meiosis were found to evolve much more quickly. Rates of protein evolution were fastest for genes expressed during the dramatic morphogenesis of round spermatids into spermatozoa. Late-expressed genes were also more likely to be specific to the male germline. To test for evidence of positive selection, we analyzed the ratio of nonsynonymous to synonymous changes using a maximum likelihood framework in comparisons among mouse, rat, and human. Many genes showed evidence of positive selection, and most of these genes were expressed late in spermatogenesis and were testis specific. Overall, these data suggest that the intensity of positive selection associated with the evolution of male gametes varies considerably across development and acts primarily on phenotypes that develop late in spermatogenesis.

Animals↗

A simple dependence between protein evolution rate and the number of protein-protein interactions.

BACKGROUND: It has been shown for an evolutionarily distant genomic comparison that the number of protein-protein interactions a protein has correlates negatively with their rates of evolution. However, the generality of this observation has recently been challenged. Here we examine the problem using protein-protein interaction data from the yeast Saccharomyces cerevisiae and genome sequences from two other yeast species. RESULTS: In contrast to a previous study that used an incomplete set of protein-protein interactions, we observed a highly significant correlation between number of interactions and evolutionary distance to either Candida albicans or Schizosaccharomyces pombe. This study differs from the previous one in that it includes all known protein interactions from S. cerevisiae, and a larger set of protein evolutionary rates. In both evolutionary comparisons, a simple monotonic relationship was found across the entire range of the number of protein-protein interactions. In agreement with our earlier findings, this relationship cannot be explained by the fact that proteins with many interactions tend to be important to yeast. The generality of these correlations in other kingdoms of life unfortunately cannot be addressed at this time, due to the incompleteness of protein-protein interaction data from organisms other than S. cerevisiae. CONCLUSIONS: Protein-protein interactions tend to slow the rate at which proteins evolve. This may be due to structural constraints that must be met to maintain interactions, but more work is needed to definitively establish the mechanism(s) behind the correlations we have observed.

Candida albicans↗

The designability hypothesis and protein evolution.

The usage of protein folds in nature is known to be non-uniform: a few folds are used often, while most others are used relatively rarely. What makes one fold more successful than another? The designability explanation, which posits that successful folds have an exponentially larger number of compatible sequences, is critically reviewed, and compared with other structural and functional explanations. It is argued that designability is one component of fold fitness, but most likely not a dominant one.

Evolution, Molecular↗

Protein evolution on rugged landscapes.

We analyze a mathematical model of protein evolution in which the evolutionary process is viewed as hill-climbing on a random fitness landscape. In studying the structure of such landscapes, we note that a large number of local optima exist, and we calculate the time and number of mutational changes until a protein gets trapped at a local optimum. Such a hill-climbing process may underlie the evolution of antibody molecules by somatic hypermutation.

Biological Evolution↗

Protein evolution on partially correlated landscapes.

We extend an earlier model of protein evolution on a rugged landscape to the case in which the landscape exhibits a variable degree of correlation (i.e., smoothness). Correlation is introduced by assuming that a protein is composed of a set of independent blocks or domains and that mutation in one block affects the contribution of that block alone to the overall fitness of the protein. We study the statistical structure of such landscapes and apply our theory to the evolution by somatic hypermutation of antibody molecules composed of framework and complementarity-determining regions. We predict the expected number of replacement mutations in each region.

Adaptation, Biological↗

Simulation of protein evolution by random fixation of allowed codons.

Computer simulation of protein evolution is based on a simple model consisting of random fixation of allowed codons (RFAC). Random replacement of single nucleotides occurs in a DNA sequence. If this results in any of the synonomous codons for allowed amino acids the mutation is fixed, if not, there is no change in the DNA and the cycle is repeated. Multiple fixations at the same nucleotide site, back mutations, degenerate fixations and coincidental identity of amino acids all occur. RFAC simulation begins with a single DNA sequence and follows a phylogeny based on the fossil record. The rate of fixation at the level of DNA is constant. The model upon which RFAC simulation is based is the same as the neutral theory of molecular evolution. The simulation is therefore a test of this theory. The results of simulated and real evolution are compared for fibrinopeptides A in mammals and cytochromes C and hemoglobin alpha and beta chains in vertebrates. In each case the allowed variation at each site has been set equal to that observed, twice that observed and all protein amino acids. Rates of fixation vary from 2.4 X 10(-10) to 10(-8) accepted nucleotide fixations per codon per year. There is some, although never excellent, agreement between real and simulated evolution, the better fits are obtained in the cases of fibrinopeptides A and cytochromes C. The major source of discrepancy between real evolution and simulation is irregularities in the rates of real evolution. RFAC simulation is compared with the random evolutionary hit (REH) model, augmented maximum parsimony and the accepted point mutations (PAM) approach.

Amino Acids↗

An in vitro DNA virus for in vitro protein evolution.

In vitro virus is a molecular construct for in vitro protein evolution, which requires some mechanism to link phenotype to genotype. The first in vitro virus was realized by bonding a nascent protein with its coding mRNA via puromycin in in vitro translation. We report a new construct of in vitro DNA virus. The virion was a covalent cDNA-protein fusion, and virion formation did not require any modification of mRNA. Due to intactness of mRNA, this type of in vitro DNA virus will take the next step toward in vitro autonomous evolution, just like in vivo viral evolution in a cellstat.

DNA Primers↗

Significant impact of protein dispensability on the instantaneous rate of protein evolution.

The neutral theory of molecular evolution predicts that important proteins evolve more slowly than unimportant ones. High-throughput gene-knockout experiments in model organisms have provided information on the dispensability, and therefore importance, of thousands of proteins in a genome. However, previous studies of the correlation between protein dispensability and evolutionary rate were equivocal, and it has been proposed that the observed correlation is due to the covariation with the level of gene expression or is limited to duplicate genes. We here analyzed the gene dispensability data of the yeast Saccharomyces cerevisiae and estimated protein evolutionary rates by comparing S. cerevisiae with nine species of varying degrees of divergence from S. cerevisiae. The correlation between gene dispensability and evolutionary rate, although low, is highly significant, even when the gene expression level is controlled for or when duplicate genes are excluded. Our results thus support the hypothesis of lower evolution rates for more important proteins, a widely used principle in the daily practice of molecular biology. When the evolutionary rate is estimated from closely related species, the ratio between the mean rate of nonessential proteins to that of essential proteins is 1.4. This ratio declines to 1.1 when the evolutionary rate is estimated from distantly related species, suggesting that the importance of a protein may change in evolution, so the dispensability data obtained from a model organism only predicts a short-term rate of protein evolution. A comparison of the fitness contributions of orthologous genes in yeast and nematode supports this conclusion.

Animals↗

A new clustering system for protein sequences and its application to constraints discovery in protein evolution.

A conceptual clustering system, CLUSMOL/S, has been developed to classify protein sequences from a user-defined point of view. Given a grouping of amino acids as a viewpoint, the system constructs taxonomic trees of sequences based on minimum information criterion. Every tree node expresses itself as a generic consensus sequence that consists of specific consensus amino acids insertion/deletion points, and generic amino acids with a specified character. The resulting tree and generic sequences show the similarity-based relationships among sequences and their characteristics. Application to vertebrate cytochromes c yields an acceptable cladrogram only when amino acids are grouped by volume and length of sidechains. The result indicates that the steric factor is the most important constraint in the process of protein evolution.

Amino Acid Sequence↗

Predicting functional divergence in protein evolution by site-specific rate shifts.

Most modern tools that analyze protein evolution allow individual sites to mutate at constant rates over the history of the protein family. However, Walter Fitch observed in the 1970s that, if a protein changes its function, the mutability of individual sites might also change. This observation is captured in the "non-homogeneous gamma model", which extracts functional information from gene families by examining the different rates at which individual sites evolve. This model has recently been coupled with structural and molecular biology to identify sites that are likely to be involved in changing function within the gene family. Applying this to multiple gene families highlights the widespread divergence of functional behavior among proteins to generate paralogs and orthologs.

Amino Acid Sequence↗

Monophyly of class I aminoacyl tRNA synthetase, USPA, ETFP, photolyase, and PP-ATPase nucleotide-binding domains: implications for protein evolution in the RNA.

Protein sequence and structure comparisons show that the catalytic domains of Class I aminoacyl-tRNA synthetases, a related family of nucleotidyltransferases involved primarily in coenzyme biosynthesis, nucleotide-binding domains related to the UspA protein (USPA domains), photolyases, electron transport flavoproteins, and PP-loop-containing ATPases together comprise a distinct class of alpha/beta domains designated the HUP domain after HIGH-signature proteins, UspA, and PP-ATPase. Several lines of evidence are presented to support the monophyly of the HUP domains, to the exclusion of other three-layered alpha/beta folds with the generic "Rossmann-like" topology. Cladistic analysis, with patterns of structural and sequence similarity used as discrete characters, identified three major evolutionary lineages within the HUP domain class: the PP-ATPases; the HIGH superfamily, which includes class I aaRS and related nucleotidyltransferases containing the HIGH signature in their nucleotide-binding loop; and a previously unrecognized USPA-like group, which includes USPA domains, electron transport flavoproteins, and photolyases. Examination of the patterns of phyletic distribution of distinct families within these three major lineages suggests that the Last Universal Common Ancestor of all modern life forms encoded 15-18 distinct alpha/beta ATPases and nucleotide-binding proteins of the HUP class. This points to an extensive radiation of HUP domains before the last universal common ancestor (LUCA), during which the multiple class I aminoacyl-tRNA synthetases emerged only at a late stage. Thus, substantial evolutionary diversification of protein domains occurred well before the modern version of the protein-dependent translation machinery was established, i.e., still in the RNA world.

Adenosine Triphosphatases↗

Stability constraints and protein evolution: the role of chain length, composition and disulfide bonds.

Stability of the native state is an essential requirement in protein evolution and design. Here we investigated the interplay between chain length and stability constraints using a simple model of protein folding and a statistical study of the Protein Data Bank. We distinguish two types of stability of the native state: with respect to the unfolded state (unfolding stability) and with respect to misfolded configurations (misfolding stability). Several contributions to stability are evaluated and their correlations are disentangled through principal components analysis, with the following main results. (1) We show that longer proteins can fulfil more easily the requirements of unfolding and misfolding stability, because they have a higher number of native interactions per residue. Consistently, in longer proteins native interactions are weaker and they are less optimized with respect to non-native interactions. (2) Stability against misfolding is negatively correlated with the strength of native interactions, which is related to hydrophobicity. Hence there is a trade-off between unfolding and misfolding stability. This trade-off is influenced by protein length: less hydrophobic sequences are observed in very long proteins. (3) The number of disulfide bonds is positively correlated with the deficit of free energy stabilizing the native state. Chain length and the number of disulfide bonds per residue are negatively correlated in proteins with short chains and uncorrelated in proteins with long chains. (4) The number of salt bridges per residue and per native contact increases with chain length. We interpret these observations as an indication that the constraints imposed by unfolding stability are less demanding in long proteins and they are further reduced by the competing requirement for stability against misfolding. In particular, disulfide bonds appear to be positively selected in short proteins, whereas they evolve in an effectively neutral way in long proteins.

Amino Acids↗

No accelerated rate of protein evolution in male-biased Drosophila pseudoobscura genes.

Sexually dimorphic traits are often subject to diversifying selection. Genes with a male-biased gene expression also are probably affected by sexual selection and have a high rate of protein evolution. We used SAGE to measure sex-biased gene expression in Drosophila pseudoobscura. Consistent with previous results from D. melanogaster, a larger number of genes were male biased (402 genes) than female biased (138 genes). About 34% of the genes changed the sex-related expression pattern between D. melanogaster and D. pseudoobscura. Combining gene expression with protein divergence between both species, we observed a striking difference in the rate of evolution for genes with a male-biased gene expression in one species only. Contrary to expectations, D. pseudoobscura genes in this category showed no accelerated rate of protein evolution, while D. melanogaster genes did. If sexual selection is driving molecular evolution of male-biased genes, our data imply a radically different selection regime in D. pseudoobscura.

Amino Acid Substitution↗