Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “molecular evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Sensitivity of patterns of molecular evolution to alterations in methodology: a critique of Hughes and Yeager.

Employing a set of 43 othologous mouse and rat genes, Hughes and Yeager (J. Mol. Evol. 45:125-130, 1997) reported (1) no correlation between synonymous and nonsynonymous rates of nucleotide substitution, (2) a positive correlation between intronic GC contents (GCi) and intronic substitution rates (Ki), (3) that the average Ki value was very similar to the average Ks value, and (4) that the compositional correlation between the rat and the mouse genes is stronger at the third codon position (GC3) than at the first and second codon positions (GC12). We have examined the robustness of these results to alterations in substitution rate estimation protocol, alignment protocol, and statistical procedure. We find that a significant correlation between Ka and Ks is observed either if a rank correlation statistic is used instead of regression analysis, if one outlier is excluded from the analysis, or if a regression weighted by gene size is employed. The correlation between Ki and GCi we find to be sensitive to changes in alignment protocol and disappears on the use of weighted means. The finding that Ks and Ki are approximately the same is dependent on the method for estimating Ks values. Finally, the variance around the regression line of rat GC3 versus mouse GC3 we find to be significantly higher than that in GC12. The source of the discrepancy between this and Hughes and Yeager's result is unclear. The variance around the line for GC4 is higher still, as might be expected. Using a methodology that may be considered preferable to that of Hughes and Yeager, we find that all four of their results are contradicted. More importantly this analysis reinforces the need for caution in assembling and analyzing data sets, as the degree of sensitivity to what many might consider minor methodological alterations is unexpected.

Animals↗

Function and molecular evolution of multicopper blue proteins.

Multicopper blue proteins (MCBPs) are multidomain proteins that utilize the distinctive redox ability of copper ions. There are a variety of MCBPs that have been roughly classified into three different groups, based on their domain organization and functions: (i) nitrite reductase-type with two domains, (ii) laccase-type with three domains, and (iii) ceruloplasmin-type with six domains. Together, the second and third group are often commonly called multicopper oxidases (MCOs). The rapid accumulation of genome sequence information in recent years has revealed several new types of proteins containing MCBP domains, mainly from bacteria. In this review, the recent research on the functions and structures of MCBPs is summarized, mainly focusing on the new types. The latter half of this review focusses on the two domain MCBPs, which we propose as the evolutionary intermediate of the MCBP family.

Amino Acid Sequence↗

Molecular evolution within the L-malate and L-lactate dehydrogenase super-family.

The NAD(P)-dependent malate (L-MalDH) and NAD-dependent lactate (L-LDH) form a large super-family that has been characterized in organisms belonging to the three domains of life. In the first part of this study, the group of [LDH-like] L-MalDH, which are malate dehydrogenases resembling lactate dehydrogenase, were analyzed and clearly defined with respect to the other enzymes. In the second part, the phylogenetic relationships of the whole super-family were presented by taking into account the [LDH-like] L-MalDH. The inferred tree unambiguously shows that two ancestral genes duplications, and not one as generally thought, are needed to explain both the distribution into two enzymatic functions and the observation of three main groups within the super-family: L-LDH, [LDH-like] L-MalDH, and dimeric L-MalDH. In addition, various cases of functional changes within each group were observed and analyzed. The direction of evolution was found to always be polarized: from enzymes with a high stringency of substrate recognition to enzymes with a broad substrate specificity. A specific phyletic distribution of the L-LDH, [LDH-like] L-MalDH, and dimeric L-MalDH over the Archaeal, Bacterial, and Eukaryal domains was observed. This was analyzed in the light of biochemical, structural, and genomic data available for the L-LDH, [LDH-like] L-MalDH, and dimeric L-MalDH. This analysis led to the elaboration of a refined evolutionary scenario of the super-family, in which the selection of L-LDH and the fate of L-MalDH during mitochrondrial genesis are presented.

Animals↗

Comparative molecular evolution of primary (Buchnera) and secondary symbionts of aphids based on two protein-coding genes.

A+T content, phylogenetic relationships, codon usage, evolutionary rates, and ratio of synonymous versus non-synonymous substitutions have been studied in partial sequences of the atpD and aroQ/pheA genes of primary ( Buchnera) and secondary symbionts of aphids and a set of selected non-symbiotic bacteria, belonging to the five subdivisions of the Proteobacteria. Compared to the homologous genes of the last group, both genes belonging to Buchnera behave in a similar way, showing a higher A+T content, forming a monophyletic group, a loss in codon bias, especially in third base position, an evolutionary acceleration and an increase in the number of non-synonymous substitutions, confirming previous results reported elsewhere for other genes. When available, these properties have been partly observed with the secondary symbionts, but with values that are intermediate between Buchnera and free living Proteobacteria. They show high A+T content, but not as high as Buchnera, a non-solved phylogenetic position between Buchnera, and the other gamma-Proteobacteria, a loss in codon bias, again not as high as in Buchnera and a significant evolutionary acceleration in the case of the three atpD genes, but not when considering aroQ/pheA genes. These results give support to the hypothesis that they are symbionts at different stages of the symbiotic accommodation to the host.

AT Rich Sequence↗

Molecular evolution of the lysine biosynthetic pathways.

Among the different biosynthetic pathways found in extant organisms, lysine biosynthesis is peculiar because it has two different anabolic routes. One is the diaminopimelic acid pathway (DAP), and the other over the a-aminoadipic acid route (AAA). A variant of the AAA route that includes some enzymes involved in arginine and leucine biosyntheses has been recently reported in Thermus thermophilus (Nishida et al. 1999). Here we describe the results of a detailed genomic analysis of each of the sequences involved in the two lysine anabolic routes, as well as of genes from other routes related to them. No evidence was found of an evolutionary relationship between the DAP and AAA enzymes. Our results suggest that the DAP pathway is related to arginine metabolism, since the lysC, asd, dapC, dapE, and lysA genes from lysine biosynthesis are related to the argB, argC, argD, argE, and speAC genes, respectively, whose products catalyze different steps in arginine metabolism. This work supports previous reports on the relationship between AAA gene products and some enzymes involved in leucine biosynthesis and the tricarboxylic acid cycle (Irvin and Bhattacharjee 1998; Miyazaki et al. 2001). Here we discuss the significance of the recent finding that several genes involved in the arginine (Arg) and leucine (Leu) biosynthesis participate in a new alternative route of the AAA pathway (Miyazaki et al. 2001). Our results demonstrate a clear relationship between the DAP and Arg routes, and between the AAA and Leu pathways.

Amino Acid Sequence↗

Unusually expanded SSU ribosomal DNA of primary osmotrophic euglenids: molecular evolution and phylogenetic inference.

Expansion segments within eukaryotic nuclear SSU ribosomal RNA have been characterized in many diverse organisms. So far, only a few studies have examined the evolutionary history of SSU rDNA variable regions for monophyletic groups. A euglenozoan SSU rDNA data set was analyzed combining phylogenetic inference and examination of expansion segment evolution. Although SSU rDNA length expansion could be ascribed to all Euglenozoa, most unusual length variation occurs within primary osmotrophic euglenids, particularly within the genus Distigma. The longest SSU rRNA gene reported to date can be found in D. sennii, comprising more than 4500 bases. RT-PCR analyses revealed that the complete gene is transcribed into RNA without posttranscriptional modifications. Further investigations uncovered that most of the length extension is due to elongated variable regions V2 and V4, but virtually all variable regions except for V3 are extended within primary osmotrophic euglenids. Analyses of secondary structure revealed several insertion points within variable regions, some of which are of phylogenetic importance. Varying GC content has been detected among species and between expansion segments and core regions. Nevertheless, individual expansion segments of one species as well as variable sequence positions within core regions tend to evolve in parallel concerning nucleotide frequencies. The presence of a large internal repeat within V2 of Distigma sennii hints at a possible mechanism for large-scale sequence length expansion.

Animals↗

Molecular evolution of vertebrate goose-type lysozyme genes.

We have found that mammalian genomes contain two lysozyme g genes. To better understand the function of the lysozyme g genes we have examined the evolution of this small gene family. The lysozyme g gene structure has been largely conserved during vertebrate evolution, except at the 5' end of the gene, which varies in number of exons. The expression pattern of the lysozyme g gene varies between species. The fish lysozyme g sequences, unlike bird and mammalian lysozyme g sequences, do not predict a signal peptide, suggesting that the encoded proteins are not secreted. The fish sequences also do not conserve cysteine residues that generate disulfide bridges in the secreted bird enzymes, supporting the hypothesis that the fish enzymes have an intracellular function. The signal peptide found in bird and mammalian lysozyme g genes may have been acquired as an exon in the ancestor of birds and mammals, or, alternatively, an exon encoding the signal peptide has been lost in fish. Both explanations account for the change in gene structure between fish and tetrapods. The mammalian lysozyme g sequences were found to have evolved at an accelerated rate, and to have not perfectly conserved the known active site catalytic triad of the bird enzymes. This observation suggests that the mammalian enzymes may have altered their biological function, as well.

Amino Acid Sequence↗

Molecular evolution of amelogenin in mammals.

An evolutionary analysis of mammalian amelogenin, the major protein of forming enamel, was conducted by comparison of 26 sequences (including 14 new ones) representative of the main mammalian lineages. Amelogenin shows highly conserved residues in the hydrophilic N- and C-terminal regions. The central hydrophobic region (most of exon 6) is more variable, but it has conserved a high amount of proline and glutamine located in triplets, PXQ, indicating that these residues play an important role. This region evolves more rapidly, and is less constrained, than the other well-conserved regions, which are subjected to strong constraints. The comparison of the substitution rates in relation to the CpG richness confirmed that the highly conserved regions are subjected to strong selective pressures. The amino acids located at important sites and the residues known to lead to amelogenesis imperfecta when substituted were present in all sequences examined. Evolutionary analysis of the variable region of exon 6 points to a particular zone, rich in either amino acid insertion or deletion. We consider this region a hot spot of mutation for the mammalian amelogenin. In this region, numerous triplet repeats (PXQ) have been inserted recently and independently in five lineages, while most of the hydrophobic exon 6 region probably had its origin in several rounds of triplet insertions, early in vertebrate evolution. The putative ancestral DNA sequence of the mammalian amelogenin was calculated using a maximum likelihood approach. The putative ancestral protein was composed of 177 residues. It already contained all important amino acid positions known to date, its hydrophobic variable region was rich in proline and glutamine, and it contained triplet repeats PXQ as in the modern sequences.

Amelogenin↗

Molecular evolution of three avian neurotrophin genes: implications for proregion functional constraints.

Neurotrophin proteins are essential for the survival, differentiation, and maintenance of neurons in the peripheral and central nervous systems. Recent studies have shown that the unprocessed proforms of the neurotrophins are preferential high-affinity ligands for p75NTR and potent inducers of p75(NTR)-mediated cell death. Here, we explore differences in the selective constraints acting on the proregions of the three avian neurotrophin genes--NT-3, BDNF, and NGF--in an explicit phylogenetic context. We found a 50-fold difference in levels of constraint as estimated by dN/ds ratios, with the NGF proregion showing the lowest degree of constraint and BDNF the highest. These patterns suggest that the high conservation exhibited by the BDNF proregion results from intense functional constraints that are relaxed in NGF and somewhat relaxed in NT-3. The proregion of BDNF is likely to have a function that differentiates it from the corresponding regions of the NGF and NT-3 genes, suggesting that BDNF is the avian neurotrophin most likely to be used both in its precursor and mature forms in vivo.

Amino Acid Sequence↗

Secondary structure and molecular evolution of the mitochondrial small subunit ribosomal RNA in Agaricales (Euagarics clade, Homobasidiomycota).

The complete sequences and secondary structures of the mitochondrial small subunit (SSU) ribosomal RNAs of both mostly cultivated mushrooms Agaricus bisporus (1930 nt) and Lentinula edodes (2164 nt) were achieved. These secondary structures and that of Schizophyllum commune (1872 nt) were compared to that previously established for Agrocybe aegerita. The four structures are near the model established for Archae, Bacteria, plastids, and mitochondria; particularly the helices 23 and 37, described as specific to bacteria, are present. Within the four Agaricales (Homobasidiomycota), the SSU-rRNA "core" is conserved in size (966 to 1009 nt) with the exception of an unusual extension of 40 nt in the H17 helix of S. commune. The four core sequences possess 76% of conserved positions and a cluster of C in their 3' end, which could constitute a signal involved in the RNA maturation process. Among the nine putative variable domains, three (V3, V5, V7) do not show significant length variations and possess similar percentages of conserved positions (69%) than the core. The other six variable domains show important length variations, due to independent large size inserted/deleted sequences, and higher rates of nucleotide substitutions than the core (only 31% of conserved positions between the four species). Interestingly, the inserted/deleted sequences are located in few preferential sites (hot spots for insertion/deletion) where they seem to arise or disappear haphazardly during evolution. These sites are located on the surface of the tertiary structure of the 30S ribosomal subunit, at the beginning of hairpin loops; the insertions lead to a lengthening of existing hairpins or to branching loops bearing up to five additional helices.

Agaricales↗

Molecular evolution of cycloidea-like genes in Fabaceae.

The cycloidea (CYC) gene controls floral symmetry in snapdragon (Antirrhinum majus). We investigated the evolution of CYC-like genes in some species of legumes that have zygomorphic flowers. Two to four CYC-like genes were isolated from a single species. The results of NJ and ML analyses indicate that CYC-like genes in legumes group into two monophyletic clades; one group consists of eight CYC-like genes (Clade 1) and the other contains three CYC-like genes and TB1 of maize (Clade 2). These phylogenetic trees and the Shimodaira-Hasegawa test suggest that Clade 1 is a sister of the original CYC group (Clade 3). Moreover, the result of the GeneTree analysis showed that the CYC-like genes experienced repeated duplication events during the evolution of legumes. We herein speculate as to the role of CYC-like genes in legumes and discuss the evolutionary processes that these genes have undergone.

Amino Acid Sequence↗

Molecular evolution in large genetic networks: does connectivity equal constraint?

Genetic networks show a broad-tailed distribution of the number of interaction partners per protein, which is consistent with a power-law. It has been proposed that such broad-tailed distributions are observed because they confer robustness against mutations to the network. We evaluate this hypothesis for two genetic networks, that of the E. coli core intermediary metabolism and that of the yeast protein-interaction network. Specifically, we test the hypothesis through one of its key predictions: highly connected proteins should be more important to the cell and, thus, subject to more severe selective and evolutionary constraints. We find, however, that no correlation between highly connected proteins and evolutionary rate exists in the E. coli metabolic network and that there is only a weak correlation in the yeast protein-interaction network. Furthermore, we show that the observed correlation is function-specific within the protein-interaction network: only genes involved in the cell cycle and transcription show significant correlations. Our work sheds light on conflicting results by previous researchers by comparing data from multiple types of protein-interaction datasets and by using a closely related species as a reference taxon. The finding that highly connected proteins can tolerate just as many amino acid substitutions as other proteins leads us to conclude that power-laws in cellular networks do not reflect selection for mutational robustness.

Energy Metabolism↗

Molecular evolution of hisB genes.

The sixth and eighth steps of histidine biosynthesis are catalyzed by an imidazole glycerol-phosphate (IGP) dehydratase (EC 4.2.1.19) and by a histidinol-phosphate (HOL-P) phosphatase (EC 3.1.3.15), respectively. In the enterobacteria, in Campylobacter jejuni and in Xylella/Xanthomonas the two activities are associated with a single bifunctional polypeptide encoded by hisB. On the other hand, in Archaea, Eucarya, and most Bacteria the two activities are encoded by two separate genes. In this work we report a comparative analysis of the amino acid sequence of all the available HisB proteins, which allowed us to depict a likely evolutionary pathway leading to the present-day bifunctional hisB gene. According to the model that we propose, the bifunctional hisB gene is the result of a fusion event between two independent cistrons joined by domain-shuffling. The fusion event occurred recently in evolution, very likely in the proteobacterial lineage after the separation of the gamma- and the beta-subdivisions. Data obtained in this work established that a paralogous duplication event of an ancestral DDDD phosphatase encoding gene originated both the HOL-P phosphatase moiety of the E. coli hisB gene and the gmhB gene coding for a DDDD phosphatase, which is involved in the biosynthesis of a precursor of the inner core of the outer membrane lipopolysaccharides (LPS).

Amino Acid Sequence↗

Molecular evolution of ldpA, a gene mediating the circadian input signal in cyanobacteria.

The ldpA gene is an element of the cyanobacterial circadian system and mediates input to the clock. Using complete prokaryotic genomes from various public databases, I analyzed the structure and phylogeny of the ldpA genes. This gene belongs to the large superfamily of ferredoxins and has a HycB domain as a core element of its structure. In addition to this domain, ldpA has two conserved terminal domains that are specific to this gene and have no homologs in the databases. All three domains are under different selective constraints. The ldpA tree topology features two very distinct clades that are essentially the same as those in the previously reported trees of the sasA gene and the kaiBC operon, two other elements of the circadian system. The data on the ldpA polymorphism and evolutionary patterns give further support to the existence of two types of the system, kaiABC- and kaiBC-based, respectively. Each type has specific functional and selective constraints, which have likely been attained through highly concordant evolution of the system's components.

Amino Acid Sequence↗

Molecular evolution of the hepatitis delta virus antigen gene: recombination or positive selection?

We present the statistical analysis of diversifying selective pressures on the hepatitis D antigen gene (HDAg). Thirty-three distinct HDAg sequences from subtypes I, II, and III were tested for positive selection using maximum likelihood methods based on models of codon substitution that allow variable selective pressures across sites. Such methods have been shown to be sufficiently accurate and successful in detecting positive selection in a variety of viral and nonviral protein-coding genes. About 11% of codon sites in HDAg were estimated to be under diversifying selection. Remarkably, most of the residues predicted to evolve under positive selection were located in the immunogenic domain and the N-terminus region with reported antigenic activity. These sites are potential targets of the host's immune response. Identification of residues mutating to escape immune recognition may help to distinguish the most virulent strains and aid vaccine design. Possible interplay between positive selection and recombination on the gene is discussed but no significant evidence for recombination was found.

Amino Acid Sequence↗

Molecular evolution and phylogeny of sipunculan hemerythrins.

We sequenced seven new hemerythrin (Hr) and myohemerythrin (myoHr) cDNAs from Sipunculus nudus and Golfingia vulgaris vulgaris, thus providing new comparative data that significantly increase the set of the known Hr and myoHr sequences. Bayesian inference, maximum likelihood, and maximum parsimony phylogenetic analyses were performed to investigate the evolutionary relationships among the sipunculan and annelid Hr and myoHr sequences. Annelid myoHrs and sipunculan Hrs were resolved as monophyletic groups. Conversely sipunculan myoHrs did not form a clade. The Hrs having an octameric quaternary structure were resolved as a monophyletic group. The octameric cluster includes the Hr sequences of G. v. vulgaris, Themiste zostericola, Themiste discriptum, and Phascolopsis gouldii. Siphonosoma cumanense Hr, which has a trimeric quaternary structure, assumes a sister group position of the octameric clade. The S. nudus Hrs, having a quaternary structure that is not well resolved, assume an isolate position within the Hrs clade. Likelihood-based analyses reveal that purifying selection mainly characterized the evolution of Hr and myoHr. We suggest that starting from a common gene ancestor, two distinct quaternary structures evolved in the sipunculan Hrs and this differentiation was probably favored by the acquisition of distinct physiological advantages.

Amino Acid Sequence↗

Episodic molecular evolution of pituitary growth hormone in Cetartiodactyla.

The sequence of growth hormone (GH) is generally strongly conserved in mammals, but episodes of rapid change occurred during the evolution of primates and artiodactyls, when the rate of GH evolution apparently increased substantially. As a result the sequences of higher primate and ruminant GHs differ markedly from sequences of other mammalian GHs. In order to increase knowledge of GH evolution in Cetartiodactyla (Artiodactyla plus Cetacea) we have cloned and characterized GH genes from camel (Camelus dromedarius), hippopotamus (Hippopotamus amphibius), and giraffe (Giraffa camelopardalis), using genomic DNA and a polymerase chain reaction technique. As in other mammals, these GH genes comprise five exons and four introns. Two very similar GH gene sequences (encoding identical proteins) were found in each of hippopotamus and giraffe. The deduced sequence for the mature hippopotamus GH is identical to that of dolphin, in accord with current ideas of a close relationship between Cetacea and Hippopotamidae. The sequence of camel GH is identical to that reported previously for alpaca GH. The sequence of giraffe GH is very similar to that of other ruminants but differs from that of nonruminant cetartiodactyls at about 18 residues. The results demonstrate that the apparent burst of rapid evolution of GH occurred largely after the separation of the line leading to ruminants from other cetartiodactyls.

Amino Acid Sequence↗

Molecular evolution of daphnia immunity genes: polymorphism in a gram-negative binding protein gene and an alpha-2-macroglobulin gene.

Studies of DNA polymorphism have shown that some immune system genes of mammals and plants are exceptionally diverse, indicating that coevolution between these taxa and their parasites mediates positive selective sweeps and/or balancing selection. The genes of the arthropod immune system remain comparatively unstudied. We isolated two putative immune system genes from the cladoceran crustacean Daphnia and examined DNA sequence diversity. For one gene, encoding a putative gram-negative binding protein, we found evidence of only purifying selection, indicating that this gene is under strong functional constraint and that selection acts to eliminate amino acid variation. For another gene, encoding a putative alpha-2-macroglobulin, we found evidence of positive selection, indicating the possible involvement of this gene in a host-parasite arms race. We discuss the assumed function of these genes and offer speculation regarding which components of the arthropod immune system might experience diversifying adaptive evolution.

Amino Acid Sequence↗