Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “structural phylogenetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Molecular evolution of calmodulin-like domain protein kinases (CDPKs) in plants and protists.

Many genes for calmodulin-like domain protein kinases (CDPKs) have been identified in plants and Alveolate protists. To study the molecular evolution of the CDPK gene family, we performed a phylogenetic analysis of CDPK genomic sequences. Analysis of introns supports the phylogenetic analysis; CDPK genes with similar intron/exon structure are grouped together on the phylogenetic tree. Conserved introns support a monophyletic origin for plant CDPKs, CDPK-related kinases, and phosphoenolpyruvate carboxylase kinases. Plant CDPKs divide into two major branches. Plant CDPK genes on one branch share common intron positions with protist CDPK genes. The introns shared between protist and plant CDPKs presumably originated before the divergence of plants from Alveolates. Additionally, the calmodulin-like domains of protist CDPKs have intron positions in common with animal and fungal calmodulin genes. These results, together with the presence of a highly conserved phase zero intron located precisely at the beginning of the calmodulin-like domain, suggest that the ancestral CDPK gene could have originated from the fusion of protein kinase and calmodulin genes facilitated by recombination of ancient introns.

Amino Acid Sequence↗

From a comb to a tree: phylogenetic relationships of the comb-footed spiders (Araneae, Theridiidae) inferred from nuclear and mitochondrial genes.

The family Theridiidae is one of the most diverse assemblages of spiders, from both a morphological and ecological point of view. The family includes some of the very few cases of sociality reported in spiders, in addition to bizarre foraging behaviors such as kleptoparasitism and araneophagy, and highly diverse web architecture. Theridiids are one of the seven largest families in the Araneae, with about 2200 species described. However, this species diversity is currently grouped in half the number of genera described for other spider families of similar species richness. Recent cladistic analyses of morphological data have provided an undeniable advance in identifying the closest relatives of the theridiids as well as establishing the family's monophyly. Nevertheless, the comb-footed spiders remain an assemblage of poorly defined genera, among which hypothesized relationships have yet to be examined thoroughly. Providing a robust cladistic structure for the Theridiidae is an essential step towards the clarification of the taxonomy of the group and the interpretation of the evolution of the diverse traits found in the family. Here we present results of a molecular phylogenetic analysis of a broad taxonomic sample of the family (40 taxa in 33 of the 79 currently recognized genera) and representatives of nine additional araneoid families, using approximately 2.5kb corresponding to fragments of three nuclear genes (Histone 3, 18SrDNA, and 28SrDNA) and two mitochondrial genes (16SrDNA and CoI). Several methods for incorporating indel information into the phylogenetic analysis are explored, and partition support for the different clades and sensitivity of the results to different assumptions of the analysis are examined as well. Our results marginally support theridiid monophyly, although the phylogenetic structure of the outgroup is unstable and largely contradicts current phylogenetic hypotheses based on morphological data. Several groups of theridiids receive strong support in most of the analyses: latrodectines, argyrodines, hadrotarsines, a revised version of spintharines and two clades including all theridiids without trace of a colulus and those without colular setae. However, the interrelationships of these clades are sensitive to data perturbations and changes in the analysis assumptions.

Animals↗

Structure and mechanism of the aberrant ba(3)-cytochrome c oxidase from thermus thermophilus.

Cytochrome c oxidase is a respiratory enzyme catalysing the energy-conserving reduction of molecular oxygen to water. The crystal structure of the ba(3)-cytochrome c oxidase from Thermus thermophilus has been determined to 2.4 A resolution using multiple anomalous dispersion (MAD) phasing and led to the discovery of a novel subunit IIa. A structure-based sequence alignment of this phylogenetically very distant oxidase with the other structurally known cytochrome oxidases leads to the identification of sequence motifs and residues that seem to be indispensable for the function of the haem copper oxidases, e.g. a new electron transfer pathway leading directly from Cu(A) to Cu(B). Specific features of the ba(3)-oxidase include an extended oxygen input channel, which leads directly to the active site, the presence of only one oxygen atom (O(2-), OH(-) or H(2)O) as bridging ligand at the active site and the mainly hydrophobic character of the interactions that stabilize the electron transfer complex between this oxidase and its substrate cytochrome c. New aspects of the proton pumping mechanism could be identified.

Amino Acid Sequence↗

Predicting protein functional sites with phylogenetic motifs.

In this report, we demonstrate that phylogenetic motifs, sequence regions conserving the overall familial phylogeny, represent a promising approach to protein functional site prediction. Across our structurally and functionally heterogeneous data set, phylogenetic motifs consistently correspond to functional sites defined by both surface loops and active site clefts. Additionally, the partially buried prosthetic group regions of cytochrome P450 and succinate dehydrogenase are identified as phylogenetic motifs. In nearly all instances, phylogenetic motifs are structurally clustered, despite little overall sequence proximity, around key functional site features. Based on calculated false-positive expectations and standard motif identification methods, we show that phylogenetic motifs are generally conserved in sequence. This result implies that they can be considered motifs in the traditional sense as well. However, there are instances where phylogenetic motifs are not (overall) well conserved in sequence. This point is enticing, because it implies that phylogenetic motifs are able to identify key sequence regions that traditional motif-based approaches would not. Further, phylogenetic motif results are also shown to be consistent with evolutionary trace results, and bootstrapping is used to demonstrate tree significance.

Amino Acid Motifs↗

Secondary structural analysis of retrovirus integrase: characterization by circular dichroism and empirical prediction methods.

The retrovirus integrase (IN) protein is essential for integration of viral DNA into host DNA. The secondary structure of the purified IN protein from avian myeloblastosis virus was investigated by both circular dichroism (CD) spectroscopy and five empirical prediction methods. The secondary structures determined from the resolving of CD spectra through a least-squares curve fitting procedure were compared with those predicted from four statistical methods, e.g., the Chou-Fasman, Garnier-Osguthorpe-Robson, Nishikawa-Ooi, and a JOINT scheme which combined all three of these methods, plus a pure a priori one, the Ptitsyn-Finkelstein method. Among all of the methods used, the Nishikawa-Ooi prediction gave the closest match in the composition of secondary structure to the CD result, although the other methods each correctly predicted one or more secondary structural group. Most of the alpha-helix and beta-sheet states predicted by the Ptitsyn-Finkelstein method were in accord with the Nishikawa-Ooi method. Secondary structural predictions by the Nishikawa-Ooi method were extended further to include IN proteins from four phylogenetic distinct retroviruses. The structural relationships between the four most conserved amino acid blocks of these IN proteins were compared using sequence homology and secondary structure predictions.

Amino Acid Sequence↗

PASSML: combining evolutionary inference and protein secondary structure prediction.

MOTIVATION: Evolutionary models of amino acid sequences can be adapted to incorporate structure information; protein structure biologists can use phylogenetic relationships among species to improve prediction accuracy. Results : A computer program called PASSML ('Phylogeny and Secondary Structure using Maximum Likelihood') has been developed to implement an evolutionary model that combines protein secondary structure and amino acid replacement. The model is related to that of Dayhoff and co-workers, but we distinguish eight categories of structural environment: alpha helix, beta sheet, turn and coil, each further classified according to solvent accessibility, i.e. buried or exposed. The model of sequence evolution for each of the eight categories is a Markov process with discrete states in continuous time, and the organization of structure along protein sequences is described by a hidden Markov model. This paper describes the PASSML software and illustrates how it allows both the reconstruction of phylogenies and prediction of secondary structure from aligned amino acid sequences. AVAILABILITY: PASSML 'ANSI C' source code and the example data sets described here are available at http://ng-dec1.gen.cam.ac.uk/hmm/Passml.html and 'downstream' Web pages. CONTACT: P.Lio@gen.cam.ac.uk

Adenylate Kinase↗

The modeled structure of the RNA dependent RNA polymerase of GBV-C virus suggests a role for motif E in Flaviviridae RNA polymerases.

BACKGROUND: The Flaviviridae virus family includes major human and animal pathogens. The RNA dependent RNA polymerase (RdRp) plays a central role in the replication process, and thus is a validated target for antiviral drugs. Despite the increasing structural and enzymatic characterization of viral RdRps, detailed molecular replication mechanisms remain unclear. The hepatitis C virus (HCV) is a major human pathogen difficult to study in cultured cells. The bovine viral diarrhea virus (BVDV) is often used as a surrogate model to screen antiviral drugs against HCV. The structure of BVDV RdRp has been recently published. It presents several differences relative to HCV RdRp. These differences raise questions about the relevance of BVDV as a surrogate model, and cast novel interest on the "GB" virus C (GBV-C). Indeed, GBV-C is genetically closer to HCV than BVDV, and can lead to productive infection of cultured cells. There is no structural data for the GBV-C RdRp yet. RESULTS: We show in this study that the GBV-C RdRp is closest to the HCV RdRp. We report a 3D model of the GBV-C RdRp, developed using sequence-to-structure threading and comparative modeling based on the atomic coordinates of the HCV RdRp structure. Analysis of the predicted structural features in the phylogenetic context of the RNA polymerase family allows rationalizing most of the experimental data available. Both available structures and our model are explored to examine the catalytic cleft, allosteric and substrate binding sites. CONCLUSION: Computational methods were used to infer evolutionary relationships and to predict the structure of a viral RNA polymerase. Docking a GTP molecule into the structure allows defining a GTP binding pocket in the GBV-C RdRp, such as that of BVDV. The resulting model suggests a new proposition for the mechanism of RNA synthesis, and may prove useful to design new experiments to implement our knowledge on the initiation mechanism of RNA polymerases.

Binding Sites↗

Up-to-dating of complete sequenced DNA data of Hansenula wingei yeast mitochondria.

To update sequenced data, we determined the 5' and 3' termini of yeast Hansenula wingei (Pichia canadensis) mitochondrial (mt) large subunit ribosomal RNA (LSU) which is encoded in the mt genome. The 5' end position was mapped downstream from a putative transcription starting site which is homologous to a Saccharomyces cerevisiae mitochondrial promoter sequence. This suggests that the primary transcript of LSU is processed from 5' end and then mature transcript is formed. This processing is different from that of S. cerevisiae mt LSU in which processing on its 5' end does not occur. Based on the sequence data of H. wingei mt LSU, we constructed its secondary structure, and compared it with those of the other fungal organisms. Conserved regions of H. wingei LSU were identified and used for subsequent phylogenetic analysis. In genome structure and gene content, H. wingei mt genome has several characteristics similar to those in filamentous fungi, but the phylogenetic analysis indicates closer kinship to yeast S. cerevisiae. This agrees with previous non-sequencing phylogenies and suggests that extraordinary rearrangements have occurred in yeast mt genomes during divergent evolution.

Base Sequence↗

Genetically distinct populations in an Asian soldier-producing aphid, Pseudoregma bambucicola (Homoptera: Aphididae), identified by DNA fingerprinting and molecular phylogenetic analysis.

To estimate genetic structure of a soldier-producing aphid, Pseudoregma bambucicola, samples from natural populations throughout southeastern Asia were analyzed by a DNA fingerprinting technique. We unexpectedly found that P. bambucicola comprises two geographic groups, the northern group and the southern group, which are genetically distinct from each other but morphologically almost indistinguishable. Molecular phylogenetic and statistical analyses based on mitochondrial ribosomal DNA sequences demonstrated that the northern and southern groups of P. bambucicola are not closely related but constitute distinct lineages in the genus Pseudoregma. Detailed morphological reexamination revealed that the two groups could be distinguished by the number of setae on the 8th abdominal tergite of 1st instar nymphs and soldiers. From these results, it was suggested that P. bambucicola should be divided into two species. The northern group from Japan, Taiwan, Hong Kong, and northern Vietnam retains the name P. bambucicola, whereas we suggest that the name P. carolinensis (R. Takahashi, 1941, Tenthredo 3, 208-220) should be used for the southern group from Thailand, Malay Peninsula, Java, Irian Jaya, and Micronesia. The morphological resemblance between P. bambucicola and P. carolinensis might be due to shared ancestral characters of the genus Pseudoregma.

Animals↗

Molecular evolution of dengue 2 virus in Puerto Rico: positive selection in the viral envelope accompanies clade reintroduction.

Dengue virus is a circumtropical, mosquito-borne flavivirus that infects 50-100 million people each year and is expanding in both range and prevalence. Of the four co-circulating viral serotypes (DENV-1 to DENV-4) that cause mild to severe febrile disease, DENV-2 has been implicated in the onset of dengue haemorrhagic fever (DHF) in the Americas in the early 1980s. To identify patterns of genetic change since DENV-2's reintroduction into the region, molecular evolution in DENV-2 from Puerto Rico (PR) and surrounding countries was examined over a 20 year period of fluctuating disease incidence. Structural genes (over 20 % of the viral genome), which affect viral packaging, host-cell entry and immune response, were sequenced for 91 DENV-2 isolates derived from both low- and high-prevalence years. Phylogenetic analyses indicated that DENV-2 outbreaks in PR have been caused by viruses assigned to subtype IIIb, originally from Asia. Variation amongst DENV-2 viruses in PR has since largely arisen in situ, except for a lineage-replacement event in 1994 that appears to have non-PR New World origins. Although most structural genes have remained relatively conserved since the 1980s, strong evidence was found for positive selection acting on a number of amino acid sites in the envelope gene, which have also been important in defining phylogenetic structure. Some of these changes are exhibited by the multiple lineages present in 1994, during the largest Puerto Rican outbreak of dengue, suggesting that they may have altered disease dynamics, although their functional significance will require further investigation.

Dengue↗

Structure-sensitive RNA footprinting of yeast nuclear ribonuclease P.

Several enzymatic and chemical reagents were used to probe the secondary structure of Saccharomyces cerevisiae nuclear RNase P RNA in the presence and absence of its protein components. Double-stranded regions were detected with RNase V1 and single-stranded regions with RNase ONE (Escherichia coli RNase I). Nucleotides not paired at Watson-Crick positions were monitored with dimethyl sulfate, kethoxal, and 1-cyclohexyl-3-[2-(N-methylmorpholinio)ethyl]carbodiimide p-toluenesulfonate. The results supported most aspects of the previously proposed, phylogenetically-derived RNA secondary structure, although minor refinements allowed incorporation of both the biochemical and phylogenetic data. Digestion of the RNase P protein(s) with proteinase K gave enhanced reactivities to structure probes at selected positions, indicating regions of the RNA made inaccessible by the presence of the protein subunit(s). The regions of RNA protected in the yeast nuclear holoenzyme were considerably more extensive than that seen in the Escherichia coli holoenzyme, consistent with the observation that the protein moiety generally comprises a larger percentage of the RNase P holoenzyme in eukaryotes than in eubacteria.

Base Composition↗

A common structural core in the internal ribosome entry sites of picornavirus, hepatitis C virus, and pestivirus.

Cap-independent translations of viral RNAs of enteroviruses and rhinoviruses, cardioviruses and aphthoviruses, hepatitis A and C viruses (HAV and HCV), and pestivirus are initiated by the direct binding of 40S ribosomal subunits to a cis-acting genetic element termed the internal ribosome entry site (IRES) or ribosome landing pad (RLP) in the 5' noncoding region (5'NCR). RNA higher ordered structure models for these IRES elements were derived by a combined approach using thermodynamic RNA folding, Monte Carlo simulation, and phylogenetic comparative analysis. The structural differences among the three groups of picornaviruses arise not only from point mutations, but also from the addition or deletion of structural domains. However, a common core can be identified in the proposed structural models of these IRES elements from enteroviruses and rhinoviruses, cardioviruses and aphthoviruses, and HAV. The common structural core identified within the picornavirus IRES is also conserved in the 5'NCR of the divergent viruses, HCV, and pestiviruses. Furthermore, the proposed structural motif shares a structural feature similar to that observed in the catalytic core of the group 1 intron. The conserved structural motif from these divergent sequences that looks like the common core region of group 1 introns is probably a crucial element involved in the IRES-dependent translation.

Animals↗

Predicted secondary structure for 28S and 18S rRNA from Ichneumonoidea (Insecta: Hymenoptera: Apocrita): impact on sequence alignment and phylogeny estimation.

We utilize the secondary structural properties of the 28S rRNA D2-D10 expansion segments to hypothesize a multiple sequence alignment for major lineages of the hymenopteran superfamily Ichneumonoidea (Braconidae, Ichneumonidae). The alignment consists of 290 sequences (originally analyzed in Belshaw and Quicke, Syst Biol 51:450-477, 2002) and provides the first global alignment template for this diverse group of insects. Predicted structures for these expansion segments as well as for over half of the 18S rRNA are given, with highly variable regions characterized and isolated within conserved structures. We demonstrate several pitfalls of optimization alignment and illustrate how these are potentially addressed with structure-based alignments. Our global alignment is presented online at (http://hymenoptera.tamu.edu/rna) with summary statistics, such as basepair frequency tables, along with novel tools for parsing structure-based alignments into input files for most commonly used phylogenetic software. These resources will be valuable for hymenopteran systematists, as well as researchers utilizing rRNA sequences for phylogeny estimation in any taxon. We explore the phylogenetic utility of our structure-based alignment by examining a subset of the data under a variety of optimality criteria using results from Belshaw and Quicke (2002) as a benchmark.

Animals↗

PALI: a database of alignments and phylogeny of homologous protein structures.

PALI is a database of structure-based sequence alignments and phylogenetic relationships derived on the basis of three-dimensional structures of homologous proteins. This database enables grouping of pairs of homologous protein structures on the basis of their sequence identity calculated from the structure-based alignment and PALI also enables association of a new sequence to a family and automatic generation of a dendrogram combining the query sequence and homologous protein structures.

Databases, Factual↗

Structural rRNA characters support monophyly of raptorial limbs and paraphyly of limb specialization in water fleas.

The evolutionary success of arthropods has been attributed partly to the diversity of their limb morphologies. Large morphological diversity and increased specialization are observed in water flea (Cladocera) limbs, but it is unclear whether the increased limb specialization in different cladoceran orders is the result of shared ancestry or parallel evolution. We inferred a robust among-order cladoceran phylogeny using small-subunit and large-subunit rRNA nuclear gene sequences, signature sequence regions, novel stem-loops and secondary structure morphometrics to assess the phylogenetic distribution of limb specialization. The sequence-based and structural rRNA morphometric phylogenies were congruent and suggested monophyly of orders with raptorial limbs, but paraphyly of orders with reduced numbers of specialized limbs. These results highlight the utility of complex molecular structural characters in resolving ancient rapid radiations.

Animals↗

Multiple maternal origins and weak phylogeographic structure in domestic goats.

Domestic animals have played a key role in human history. Despite their importance, however, the origins of most domestic species remain poorly understood. We assessed the phylogenetic history and population structure of domestic goats by sequencing a hypervariable segment (481 bp) of the mtDNA control region from 406 goats representing 88 breeds distributed across the Old World. Phylogeographic analysis revealed three highly divergent goat lineages (estimated divergence >200,000 years ago), with one lineage occurring only in eastern and southern Asia. A remarkably similar pattern exists in cattle, sheep, and pigs. These results, combined with recent archaeological findings, suggest that goats and other farm animals have multiple maternal origins with a possible center of origin in Asia, as well as in the Fertile Crescent. The pattern of goat mtDNA diversity suggests that all three lineages have undergone population expansions, but that the expansion was relatively recent for two of the lineages (including the Asian lineage). Goat populations are surprisingly less genetically structured than cattle populations. In goats only approximately 10% of the mtDNA variation is partitioned among continents. In cattle the amount is >/=50%. This weak structuring suggests extensive intercontinental transportation of goats and has intriguing implications about the importance of goats in historical human migrations and commerce.

Animals↗

Venezuelan equine encephalomyelitis virus structure and its divergence from old world alphaviruses.

Although alphaviruses have been extensively studied as model systems for the structural organization of enveloped viruses, no structures exist for the phylogenetically distinct eastern equine encephalomyelitis (EEE)-Venezuelan equine encephalomyelitis (VEE) lineage of New World alphaviruses. Here we report the 25-A structure of VEE virus, obtained from electron cryomicroscopy and image reconstruction. The envelope spike glycoproteins of VEE virus have a T=4 icosahedral arrangement, similar to that observed in Old World Sindbis, Semliki Forest, and Ross River alphaviruses. However, VEE virus has pronounced differences in its nucleocapsid structure relative to nucleocapsid structures repeatedly observed in Old World alphaviruses.

Alphavirus↗

Improved alignment quality by combining evolutionary information, predicted secondary structure and self-organizing maps.

BACKGROUND: Protein sequence alignment is one of the basic tools in bioinformatics. Correct alignments are required for a range of tasks including the derivation of phylogenetic trees and protein structure prediction. Numerous studies have shown that the incorporation of predicted secondary structure information into alignment algorithms improves their performance. Secondary structure predictors have to be trained on a set of somewhat arbitrarily defined states (e.g. helix, strand, coil), and it has been shown that the choice of these states has some effect on alignment quality. However, it is not unlikely that prediction of other structural features also could provide an improvement. In this study we use an unsupervised clustering method, the self-organizing map, to assign sequence profile windows to "structural states" and assess their use in sequence alignment. RESULTS: The addition of self-organizing map locations as inputs to a profile-profile scoring function improves the alignment quality of distantly related proteins slightly. The improvement is slightly smaller than that gained from the inclusion of predicted secondary structure. However, the information seems to be complementary as the two prediction schemes can be combined to improve the alignment quality by a further small but significant amount. CONCLUSION: It has been observed in many studies that predicted secondary structure significantly improves the alignments. Here we have shown that the addition of self-organizing map locations can further improve the alignments as the self-organizing map locations seem to contain some information that is not captured by the predicted secondary structure.

Computational Biology↗