Evolutionary processes in influenza viruses: divergence, rapid evolution, and stasis.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The animal talins are large, modular proteins that link the actin cytoskeleton to the extracellular environment through interactions with beta-integrins and actin. Dictyostelium discoideum has two talins, TalA and TalB, which have distinct physiological roles in cell adhesion, cell differentiation, and cytokinesis. We previously identified a second talin gene in vertebrates. Thus, talin function in vertebrates is also due to the action of multiple proteins. Using a phylogenomic approach we have determined that D. discoideum TalA/B and the animal talins are related by descent from a common ancestral talin and that duplication of TLN2 early in the chordate lineage produced TLN1. An additional duplication subsequently produced a second Talin-2 in teleost fishes and a second Talin-1 in Xenopus laevis. We also show that vertebrate Talin-2 mRNA is alternatively processed. In the invertebrate Drosophila melanogaster and in the non-vertebrate chordate Ciona intestinalis, which each have only one talin gene, alternative processing of talin mRNA also produces multiple talin species. Thus, in these organisms, talin function may be due to the action of more than one protein. To identify isoform-specific functions of vertebrate talins we have shown through proteomic analysis that mammalian Talin-1 and Talin-2 bind to different protein partners. Further characterization of the differences between animal talins, especially the direct comparison of talins in the model urochordate C. intestinalis, which has one talin gene that produces two talins through alternative mRNA splicing, with Talin-1 and Talin-2 in model vertebrates, will provide an experimental system for studying neofunctionalization or subfunctionalization of talin following the vertebrate talin gene duplication.
To understand how protein segments are inserted and deleted during divergent evolution, a set of pairwise alignments contained exactly one gap, and therefore arising from the first insertion-deletion (indel) event in the time separating the homologs, was examined. The alignments showed that "structure breaking" amino acids (PGDNS) were preferred within and flanking gapped regions, as are two residues with hydrophilic side-chains (QE) that frequently occur at the surface of protein folds. Conversely, hydrophobic residues (FMILYVW) occur infrequently within and flanking the gapped region. These preferences are modestly different in protein pairs separated by an episode of adaptive evolution, than in pairs diverging under strong functional constraints. Surprisingly, regions near an indel have not evolved more rapidly than the sequence pair overall, showing no evidence that an indel event must be compensated by local amino acid replacement. The gap-lengths are best approximated by a Zipfian distribution, with the probability of a gap of length L decreasing as a function of L(-1.8). These features are largely independent of the length of the gap and the extent of divergence (measured by both silent and non-silent sequence changes) separating the two proteins. Surprisingly, amino acid repeats were discovered in more than a third of the polypeptide segments in and around the gap. These correspond to repeats in the DNA sequence. This suggests that a signature of the mechanism by which indels occur in the DNA sequence remains in the encoded protein sequences. These data suggest specific tools to score gap placement in an alignment. They also suggest tools that distinguish true indels from gaps created by mistaken gene finding, including under-predicted and over-predicted introns. By providing mechanisms to identify errors, the tools will enhance the value of genome sequence databases in support of integrated paleogenomics strategies used to extract functional information in a post-genomic environment.
HSP90 proteins are important molecular chaperones. Transcriptome and genome analyses revealed that the human HSP90 family includes 17 genes that fall into four classes. A standardized nomenclature for each of these genes is presented here. Classes HSP90AA, HSP90AB, HSP90B, and TRAP contain 7, 6, 3, and 1 genes, respectively. HSP90AA genes mapped onto chromosomes 1, 3, 4, and 11; HSP90AB genes mapped onto 3, 4, 6, 13 and 15; HSP90B genes mapped onto 1, 12, and 15; and the TRAP1 gene mapped onto 16. Six genes, HSP90AA1, HSP90AA2, HSP90N, HSP90AB1, HSP90B1 and TRAP1, were recognized as functional, and the remaining 11 genes were considered putative pseudogenes. Amino acid polymorphic variants were detected for genes HSP90AA1, HSP90AA2, HSP90AB1, HSP90B1, and TRAP1. The structures of these genes and the functional motifs and polymorphic variants of their proteins were documented and the features and functions of their proteins were discussed. Phylogenetic analyses based on both nucleotide and protein data demonstrated that HSP90(AA+AB+B) formed a monophyletic clade, whereas TRAP is a relatively distant paralogue of this clade.
The use of positional approaches for the isolation of genes from most crop species is difficult due to the large size of their genomes. If the order of genes in segments of the genomes is similar in different plants, it might be feasible to use smaller genomes as templates upon which to base strategies for the positional cloning of genes from other species. Comparative genetic mapping, using markers such as restriction-fragment length polymorphisms, has revealed extensive conservation of long-range genome organization (macrostructure) between related species. But is the organization of the tens or hundreds of genes between the genetic markers also conserved? Recent results suggest that the fine-scale structure (microstructure) of plant genomes is more dynamic than previously assumed from investigations of the macrostructure.
The recently discovered human DNA polymerase lambda (DNA pol lambda) has been implicated in translesion DNA synthesis across abasic sites. One remarkable feature of this enzyme is its preference for Mn(2+) over Mg(2+) as the activating metal ion, but the molecular basis for this preference is not known. Here, we present a kinetic and thermodynamic analysis of the DNA polymerase reaction catalyzed by full length human DNA pol lambda, showing that Mn(2+) favors specifically the catalytic step of nucleotide incorporation. Besides acting as a poor coactivator for catalysis, Mg(2+) appeared to bind also to an allosteric site, resulting in the inhibition of the synthetic activity of DNA pol lambda and in an increased sensitivity to end product (pyrophosphate) inhibition. Comparison with the closely related enzyme human DNA pol beta, as well as with other DNA synthesising enzymes (mammalian DNA pol alpha and DNA pol delta, Escherichia coli DNA pol I, and HIV-1 reverse transcriptase) indicated that these features are unique to DNA pol lambda. A deletion mutant of DNA pol lambda, which contained the highly conserved catalytic core only representing the C-terminal half of the protein, showed biochemical properties comparable to the full length enzyme but clearly different from the close homologue DNA pol beta, highlighting the existence of important differences between DNA pol lambda and DNA pol beta, despite a high degree of sequence similarity.
In an investigation of the evolution of the third hypervariable loop of gp120 (V3), the principal neutralization determinant of human immunodeficiency virus type 1, we have analyzed 89 V3 sequences of plasma viral RNA purified from peripheral blood samples donated over 7 years by an infected hemophiliac. Considerable sequence diversity in the V3 region was found at all time points after seroconversion. Phylogenetic analysis revealed that an important diversification had occurred by 3 years postinfection and that, subsequently, most sequences could be allocated to either one of two major lineages that persisted throughout the remainder of the infection. Rapid changes in frequency of the most common sequences and the observation that the same hexapeptide motif (GPGSAV) at the crown of the V3 loop has evolved convergently provide strong evidence that selective processes determine the evolutionary fate of sequence variants in this region.
Explore the source record for details and available documents.
Recent studies of varicella-zoster virus (VZV) DNA sequence variation, involving large numbers of globally distributed clinical isolates, suggest that this virus has diverged into at least three distinct genotypes designated European (E), Japanese (J), and mosaic (M). In the present study, we determined and analyzed the complete genomic sequences of two M VZV strains and compared them to the sequences of three E strains and two J strains retrieved from GenBank (including the Oka vaccine preparation, V-Oka). Except for a few polymorphic tandem repeat regions, the whole genome, representing approximately 125,000 nucleotides, is highly conserved, presenting a genetic similarity between the E and J genotypes of approximately 99.85%. These analyses revealed that VZV strains distinctly segregate into at least four genotypes (E, J, M1, and M2) in phylogenetic trees supported by high bootstrap values. Separate analyses of informative sites revealed that the tree topology was dependent on the region of the VZV genome used to determine the phylogeny; collectively, these results indicate the observed strain variation is likely to have resulted, at least in part, from interstrain recombination. Recombination analyses suggest that strains belonging to the M1 and M2 genotypes are mosaic recombinant strains that originated from ancestral isolates belonging to the E and J genotypes through recombination on multiple occasions. Furthermore, evidence of more recent recombination events between M1 and M2 strains is present in six segments of the VZV genome. As such, interstrain recombination in dually infected cells seems to figure prominently in the evolutionary history of VZV, a feature it has in common with other herpesviruses. In addition, we report here six novel genomic targets located in open reading frames 51 to 58 suitable for genotyping of clinical VZV isolates.
Over 35 years ago, Susumu Ohno stated that gene duplication was the single most important factor in evolution. He reiterated this point a few years later in proposing that without duplicated genes the creation of metazoans, vertebrates, and mammals from unicellular organisms would have been impossible. Such big leaps in evolution, he argued, required the creation of new gene loci with previously nonexistent functions. Bold statements such as these, combined with his proposal that at least one whole-genome duplication event facilitated the evolution of vertebrates, have made Ohno an icon in the literature on genome evolution. However, discussion on the occurrence and consequences of gene and genome duplication events has a much longer, and often neglected, history. Here we review literature dealing with the occurrence and consequences of gene duplication, beginning in 1911. We document conceptual and technological advances in gene duplication research from this early research in comparative cytology up to recent research on whole genomes, "transcriptomes," and "interactomes."
Explore the source record for details and available documents.
Divergent evolution can explain how many proteins containing structurally similar domains, which perform a variety of related functions, have evolved from a relatively small number of modules or protein domains. However, it cannot explain how protein domains with similar, but distinguishable, functions and similar, but distinguishable, structures have evolved. Examples of this are the RNA-binding protein containing the RNA-binding domain (RBD), and a newly established protein group, the cold-shock domain (CSD) protein family. Both protein domains contain conserved RNP motifs on similar single-stranded nucleic acid-binding surfaces. Apart from the RNP motifs, which have a similar function, the two families show little similarity in topology or amino acid sequence. This can be considered an interesting example of convergent evolution at the molecular level. Previously, a beta-sheet surface was found to interact with RNA in non-homologous proteins from yeast, phage and man, revealing that this mode of RNA binding may be a widely recurring theme.
Many methods have been developed to analyse protein sequences and structures, although less work has been undertaken describing and comparing protein surfaces. Evolution can lead sequences to diverge or structures to change topology; nevertheless, surface determinants that are essential to protein function itself may be mantained. Moreover, different molecules could converge to similar functions by gaining specific surface determinants. In such cases, sequence or structure comparisons are likely to be inadequate in describing or identifying protein functions and evolutionary relationships among proteins. Surface analysis can identify function determinants that are independent of sequence or secondary structure and can therefore be a powerful tool to highlight cases of possible convergent or divergent evolution. This kind of approach can be useful for a better understanding of protein molecular and biochemical mechanisms of catalysis or interaction with a ligand, which are usually surface dependent. Protein surface comparison, when compared to sequence or structure comparison methods, is a hard computational challenge and evaluated methods allowing the comparison of protein surfaces are difficult to find. In this review, we will survey the current knowledge about protein surface similarity and the techniques to detect it.
Paralogous regions are duplicated segments of chromosomal DNA that have been acquired during the evolution of the genome. Subsequent divergent evolution of the genes within paralogous regions can lead to the formation of gene families. Here, we report the identification of a region on Chromosome (Chr) 6 at 6p21.3 that is paralogous with the Spinal Muscular Atrophy (SMA) gene region on Chr 5 at 5q13.1. Partial characterization of this region identified nine sequences all of which are highly homologous to DNA sequences of the SMA gene region at 5q13.1. These sequences include four beta-glucuronidase sequences, two retrotransposon sequences, a novel cDNA, a Sequence Tagged Site (STS), and one that is homologous to exon 9 of the Neuronal Apoptosis Inhibitor Protein (NAIP) gene. The 6p21.3 paralogous SMA region may contain genes that are related to those in the SMA region at 5q13.1; however, a direct association of this region with SMA is unlikely given that no linkage of SMA with Chr 6 has been reported.
The divergent evolution of proteins in cellular signaling pathways requires ligands and their receptors to co-evolve, creating new pathways when a new receptor is activated by a new ligand. However, information about the evolution of binding specificity in ligand-receptor systems is difficult to glean from sequences alone. We have used phosphoglycerate kinase (PGK), an enzyme that forms its active site between its two domains, to develop a standard for measuring the co-evolution of interacting proteins. The N-terminal and C-terminal domains of PGK form the active site at their interface and are covalently linked. Therefore, they must have co-evolved to preserve enzyme function. By building two phylogenetic trees from multiple sequence alignments of each of the two domains of PGK, we have calculated a correlation coefficient for the two trees that quantifies the co-evolution of the two domains. The correlation coefficient for the trees of the two domains of PGK is 0. 79, which establishes an upper bound for the co-evolution of a protein domain with its binding partner. The analysis is extended to ligands and their receptors, using the chemokines as a model. We show that the correlation between the chemokine ligand and receptor trees' distances is 0.57. The chemokine family of protein ligands and their G-protein coupled receptors have co-evolved so that each subgroup of chemokine ligands has a matching subgroup of chemokine receptors. The matching subfamilies of ligands and their receptors create a framework within which the ligands of orphan chemokine receptors can be more easily determined. This approach can be applied to a variety of ligand and receptor systems.
Evolutionary history of tRNA is studied by comparative sequence analysis of two specified tRNA's at various phylogenetic levels and of tRNA families within four different species. Criteria are developed that allow 1) to distinguish between convergent and divergent evolution, 2) to determine the mechanism of divergence and 3) to estimate the degree of randomization of the variable parts of the sequences. The conclusion of these investigations is that tRNA's represent ancient molecules that existed in the form of a mutant distribution prior to their integration into genomes.
In silico genome sequence analyses suggested that mycobacteria are devoid of the highly conserved mutLS-based post-replicative mismatch repair system. Here, we present the first biological evidence for the lack of a classical mismatch repair function in mycobacteria. We found that frameshifts, but not general mutation rates are unusually high in Mycobacterium smegmatis. However, despite the absence of mismatch correction, M. smegmatis establishes a strong barrier to recombination between homeologous DNA sequences. We show that 10-12% of DNA sequence heterology restricts initiation of recombination but not extension of heteroduplex DNA intermediates. Together, the lack of mismatch correction and a high stringency of initiation of homologous recombination provide an adequate strategy for mycobacterial genome evolution, which occurs by gene duplication and divergent evolution.
Male genitalia in Drosophila exemplify strikingly rapid and divergent evolution, whereas female genitalia are relatively invariable. Whereas precopulatory and post-copulatory sexual selection has been invoked to explain this trend, the functional significance of genital structures during copulation remains obscure. We used time-sequence analysis to study the functional significance of external genitalic structures during the course of copulation, between D. melanogaster and D. simulans. This functional analysis has provided new information that reveals the importance of male-driven copulatory mechanics and strategies in the rapid diversification of genitalia. The posterior process, which is a recently evolved sexual character and present only in males of the melanogaster clade, plays a crucial role in mounting as well as in genital coupling. Whereas there is ample evidence for precopulatory and/or post-copulatory female choice, we show here that during copulation there is little or no physical female choice, consequently, males determine copulation duration. We also found subtle differences in copulatory mechanics between very closely related species. We propose that variation in male usage of novel genitalic structures and shifts in copulatory behaviour have played an important role in the diversification of genitalia in species of the Drosophila subgroup.