Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “structural phylogenetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Structure-function relationships in a self-splicing group II intron: a large part of domain II of the mitochondrial intron aI5 is not essential for self-splicing.

An oligonucleotide-directed deletion of 156 nucleotides has been introduced into the yeast mitochondrial group II intron al5 (887 nt). The deletion comprises almost all of domain II, which is one of the six phylogenetically conserved structural elements of group II introns. This mutant displays reduced self-splicing activity, but results of chemical probing with dimethylsulphate suggest that sequences at the site of the deletion interfere with the normal folding of the intron. This is supported by computer analyses, which predict a number of alternative structures involving conserved intron sequences. Splicing activity could be restored by insertion of a 10-nucleotide palindromic sequence into the unique Smal site of the deletion mutant, resulting in the formation of a small stable stem-loop element at the position of domain II. These results provide a direct correlation between folding of the RNA and its activity. We conclude that at least a large part of domain II of the group II intron al5 is not required for self-splicing activity. This deletion mutant with a length of 731 nucleotides represents the smallest self-splicing group II intron so far known.

Base Sequence↗

The influence of downstream protein-coding sequence on internal ribosome entry on hepatitis C virus and other flavivirus RNAs.

Some studies suggest that the hepatitis C virus (HCV) internal ribosome entry site (IRES) requires downstream 5' viral polyprotein-coding sequence for efficient initiation of translation, but the role of this RNA sequence in internal ribosome entry remains unresolved. We confirmed that the inclusion of viral sequence downstream of the AUG initiator codon increased IRES-dependent translation of a reporter RNA encoding secretory alkaline phosphatase, but found that efficient translation of chloramphenicol acetyl transferase (CAT) required no viral sequence downstream of the initiator codon. However, deletion of an adenosine-rich domain near the 5' end of the CAT sequence, or the insertion of a small stable hairpin structure (deltaG = -18 kcal/mol) between the HCV IRES and CAT sequences (hpCAT) substantially reduced IRES-mediated translation. Although translation could be restored to both mutants by the inclusion of 14 nt of the polyprotein-coding sequence downstream of the AUG codon, a mutational analysis of the inserted protein-coding sequence demonstrated no requirement for either a specific nucleotide or amino acid-coding sequence to restore efficient IRES-mediated translation to hpCAT. Similar results were obtained with the structurally and phylogenetically related IRES elements of classical swine fever virus and GB virus B. We conclude that there is no absolute requirement for viral protein-coding sequence with this class of IRES elements, but that there is a requirement for an absence of stable RNA structure immediately downstream of the AUG initiator codon. Stable RNA structure immediately downstream of the initiator codon inhibits internal initiation of translation but, in the case of hpCAT, did not reduce the capacity of the RNA to bind to purified 40S ribosome subunits. Thus, stable RNA structure within the 5' proximal protein-coding sequence does not alter the capacity of the IRES to form initial contacts with the 40S subunit, but appears instead to prevent the formation of subsequent interactions between the 40S subunit and viral RNA in the vicinity of the initiator codon that are essential for efficient internal ribosome entry.

Base Sequence↗

Phylogenetic comparative analysis and the secondary structure of ribonuclease P RNA--a review.

The most incisive a priori approach to inferring the higher order structure of large RNAs has proven to be the use of phylogenetic comparisons. This article provides guidelines to the method, using as an illustration the elucidation of the secondary structure of the catalytic RNA subunit of ribonuclease P (RNase P). The resultant structure is compared to the possibilities that are predicted thermodynamically for the RNase P RNA sequences of nine eubacteria.

Bacillus subtilis↗

A phylogenetic approach to target selection for structural genomics: solution structure of YciH.

Structural genomics presents an enormous challenge with up to 100 000 protein targets in the human genome alone. At current rates of structure deter-mination, judicious selection of targets is necessary. Here, a phylogenetic approach to target selection is described which makes use of the National Center for Biotechnology Information database of Clusters of Orthologous Groups (COGS). The strategy is designed so that each new protein structure is likely to provide novel sequence-fold information. To demonstrate this approach, the NMR solution structure of YciH (COG0023), a putative translation initiation factor from Escherichia coli, has been determined and its fold classified. YciH is an ortholog of eIF-1/SUI1, an integral component of the translation initiation complex in eukaryotes. The structure consists of two antiparallel alpha-helices packed against the same side of a five-stranded beta-sheet. The first 31 residues of the 11.5 kDa protein are unstructured in solution. Comparative analysis indicates that the folded portion of YciH resembles a number of structures with the alpha-beta plait topology, though its sequence is not homologous to any of them. Thus, the phylogenetic approach to target selection described here was used successfully to identify a new homologous superfamily within this topology.

Amino Acid Sequence↗

Secondary structure model for bacterial 16S ribosomal RNA: phylogenetic, enzymatic and chemical evidence.

We have derived a secondary structure model for 16S ribosomal RNA on the basis of comparative sequence analysis, chemical modification studies and nuclease susceptibility data. Nucleotide sequences of the E. coli and B. brevis 16S rRNA chains, and of RNAse T1 oligomer catalogs from 16S rRNAs of over 100 species of eubacteria were used for phylogenetic comparison. Chemical modification of G by glyoxal, A by m-chloroperbenzoic acid and C by bisulfite in naked 16S rRNA, and G by kethoxal in active and inactive 30S ribosomal subunits was taken as an indication of single stranded structure. Further support for the structure was obtained from susceptibility to RNases A and T1. These three approaches are in excellent agreement. The structure contains fifty helical elements organized into four major domains, in which 46 percent of the nucleotides of 16S rRNA are involved in base pairing. Phylogenetic comparison shows that highly conserved sequences are found principally in unpaired regions of the molecule. No knots are created by the structure.

Bacillus↗

Conformational changes involved in initiation of minus-strand synthesis of a virus-associated RNA.

Synthesis of wild-type levels of turnip crinkle virus (TCV)-associated satC complementary strands by purified, recombinant TCV RNA-dependent RNA polymerase (RdRp) in vitro was previously determined to require 3' end pairing to the large symmetrical internal loop of a phylogenetically conserved hairpin (H5) located upstream from the hairpin core promoter. However, wild-type satC transcripts, which fold into a single detectable conformation in vitro as determined by temperature-gradient gel electrophoresis, do not contain either the phylogenetically inferred H5 structure or the 3' end/H5 interaction. This implies that conformational changes are required to produce the phylogenetically inferred H5 structure for its pairing with the 3' end, which takes place subsequent to the initial conformation assumed by the RNA and prior to transcription initiation. The DR region, located 140 nucleotides upstream from the 3' end and previously determined to be important for transcription in vitro and replication in vivo, is proposed to have a role in the conformational switch, since stabilizing the phylogenetically inferred H5 structure decreases the negative effects of a DR mutation in vivo. In addition, high levels of aberrant transcription correlate with a specific conformational change in the Pr while maintaining the same conformation of the 3' terminus. These results suggest that a series of events that promote conformational changes is needed to expose the 3' terminus to the RdRp for accurate synthesis of wild-type levels of complementary strands in vitro.

Base Sequence↗

Phylogenetic invariants and geometry.

The method of invariants is an important approach in biology for determining phylogenetic information which avoids the problems involving long branch lengths that plague some other methods. In this paper, we present a geometric framework underlying the method of invariants. This perspective sheds new lights on problems in the field. It has recently enabled the solution of questions on the number and structure of phylogenetic invariants and suggests possible avenues for future empirical and theoretical research.

Animals↗

Identification of cytosolic Mg2+-dependent soluble inorganic pyrophosphatases in potato and phylogenetic analysis.

Using polyclonal antibodies raised against a previously cloned potato Mg2+-dependent soluble inorganic pyrophosphatase (ppa1 gene) [8], a second gene, called ppa2, could be isolated. A single locus homologous to ppa2 was mapped on potato chromosomes, unlinked to the two loci identified for ppa1. From a phylogenetic and structural point of view, the PPA1 and PPA2 polypeptides are more closely related to prokaryotic than to eukaryotic Mg2+-dependent soluble inorganic pyrophosphatases (soluble PPases). Subcellular localization by immunogold electron microscopy, using sections from leaf parenchyma cells, showed that PPA and PPA2 are localized to the cytosol. Based on these observations, the likely phylogenetic origin and the physiological significance of the cytosolic soluble pyrophosphatases are discussed.

Amino Acid Sequence↗

Absence of phylogenetic signal in the niche structure of meadow plant communities.

A significant proportion of the global diversity of flowering plants has evolved in recent geological time, probably through adaptive radiation into new niches. However, rapid evolution is at odds with recent research which has suggested that plant ecological traits, including the beta- (or habitat) niche, evolve only slowly. We have quantified traits that determine within-habitat alpha diversity (alpha niches) in two communities in which species segregate on hydrological gradients. Molecular phylogenetic analysis of these data shows practically no evidence of a correlation between the ecological and evolutionary distances separating species, indicating that hydrological alpha niches are evolutionarily labile. We propose that contrasting patterns of evolutionary conservatism for alpha- and beta-niches is a general phenomenon necessitated by the hierarchical filtering of species during community assembly. This determines that species must have similar beta niches in order to occupy the same habitat, but different alpha niches in order to coexist.

Base Sequence↗

Phylogenetic analysis of the cadherin superfamily allows identification of six major subfamilies besides several solitary members.

Cadherins play an important role in specific cell-cell adhesion events. Their expression appears to be tightly regulated during development and each tissue or cell type shows a characteristic pattern of cadherin molecules. Inappropriate regulation of their expression levels or functionality has been observed in human malignancies, in many cases leading to aggravated cancer cell invasion and metastasis. The cadherins form a superfamily with at least six subfamilies, which can be distinguished on the basis of protein domain composition, genomic structure, and phylogenetic analysis of the protein sequences. These subfamilies comprise classical or type-I cadherins, atypical or type-II cadherins, desmocollins, desmogleins, protocadherins and Flamingo cadherins. In addition, several cadherins clearly occupy isolated positions in the cadherin superfamily (cadherin-13, -15, -16, -17, Dachsous, RET, FAT, MEGF1 and most invertebrate cadherins). We suggest a different evolutionary origin of the protocadherin and Flamingo cadherin genes versus the genes encoding desmogleins, desmocollins, classical cadherins, and atypical cadherins. The present phylogenetic analysis may accelerate the functional investigation of the whole cadherin superfamily by allowing focused research of prototype cadherins within each subfamily.

Amino Acid Sequence↗

recA-like genes from three archaean species with putative protein products similar to Rad51 and Dmc1 proteins of the yeast Saccharomyces cerevisiae.

The process of homologous recombination has been documented in bacterial and eucaryotic organisms. The Escherichia coli RecA and Saccharomyces cerevisiae Rad51 proteins are the archetypal members of two related families of proteins that play a central role in this process. Using the PCR process primed by degenerate oligonucleotides designed to encode regions of the proteins showing the greatest degree of identity, we examined DNA from three organisms of a third phylogenetically divergent group, Archaea, for sequences encoding proteins similar to RecA and Rad51. The archaeans examined were a hyperthermophilic acidophile, Sulfolobus sofataricus (Sso); a halophile, Haloferax volcanii (Hvo); and a hyperthermophilic piezophilic methanogen, Methanococcus jannaschii (Mja). The PCR generated DNA was used to clone a larger genomic DNA fragment containing an open reading frame (orf), that we refer to as the radA gene, for each of the three archaeans. As shown by amino acid sequence alignments, percent amino acid identities and phylogenetic analysis, the putative proteins encoded by all three are related to each other and to both the RecA and Rad51 families of proteins. The putative RadA proteins are more similar to the Rad51 family (approximately 40% identity at the amino acid level) than to the RecA family (approximately 20%). Conserved sequence motifs, putative tertiary structures and phylogenetic analysis implied by the alignment are discussed. The 5' ends of mRNA transcripts to the Sso radA were mapped. The levels of radA mRNA do not increase after treatment with UV irradiation as do recA and RAD51 transcripts in E.coli and S.cerevisiae. Hence it is likely that radA in this organism is a constitutively expressed gene and we discuss possible implications of the lack of UV-inducibility.

Amino Acid Sequence↗

Relative age of proviral porcine endogenous retrovirus sequences in Sus scrofa based on the molecular clock hypothesis.

Porcine endogenous retroviruses (PERV) are discussed as putative infectious agents in xenotransplantation. PERV classes A, B, and C harbor different envelope proteins. Two different types of long terminal repeat (LTR) structures exist, of which both are present only in PERV-A. One type of LTR contains a distinct repeat structure in U3, while the other is repeatless, conferring a lower level of transcriptional activity. Since the different LTR structures are distributed unequally among the proviruses and, apparently, PERV is the only virus harboring two different LTR structures, we were interested in determining which LTR is the ancestor. Replication-competent viruses can still be found today, suggesting an evolutionary recent origin. Our studies revealed that the age of PERV is at most 7.6 x 10(6) years, whereas the repeatless LTR type evolved approximately 3.4 x 10(6) years ago, being the phylogenetically younger structure. The age determined for PERV correlates with the time of separation between pigs (Suidae, Sus scrofa) and their closest relatives, American-born peccaries (Tayassuidae, Pecari tajacu), 7.4 x 10(6) years ago.

Animals↗

DYRK gene structure and erythroid-restricted features of DYRK3 gene expression.

DYRKs are an emerging family of dual-specificity kinases that play key roles in cell proliferation, survival, and development. Up to seven mammalian DYRK isoforms have been reported, but only the DYRK1A gene (and its products) has been well characterized. Defined here are the genomic structures (and phylogenetics) of DYRK3, four additional murine and/or human DYRKs, and two related HIPK genes. For murine DYRK3, direct BAC sequences are provided, a basis for differential processing of murine versus human transcripts is defined, and tissue expression profiles plus a functional DYRK3 promoter are characterized. Unlike complex DYRKs 1A, 1B, and 4A, DYRK3 and DYRK2 possess simple 4-exon structures. For DYRK3, expression is strong in erythroid cells and testis, but is also detected in adult kidney and liver in situ. In addition, a 1930-bp DYRK3 promoter drives erythroid expression and transcript expression is inhibited sharply on GATA1-induced G1E cell differentiation.

Animals↗

Fold-recognition analysis predicts that the Tag protein family shares a common domain with the helix-hairpin-helix DNA glycosylases.

The Escherichia coli protein Tag is traditionally regarded as an archetype of one of four classes of N-alkylpurine DNA glycosylases. However, its structure and phylogenetic relationship to other glycosylases remains a mystery. Fold-recognition and sequence profile analyses suggest that Tag shares the catalytic domain with helix-hairpin-helix (HhH) glycosylases such as MutY, AlkA and EndoIII, but its N- and C-termini together form a unique His2Cys2 cluster. The findings presented in this paper provide insight into sequence-structure-function relationships in the Tag family and should aid in a more precise definition of the common core of the HhH superfamily of glycosylases involved in DNA repair.

Amino Acid Sequence↗

Applying hybrid reasoning to mine for associative features in biological data.

We develop the means to mine for associative features in biological data. The hybrid reasoning schema for deterministic machine learning and its implementation via logic programming is presented. The methodology of mining for correlation between features is illustrated by the prediction tasks for protein secondary structure and phylogenetic profiles. The suggested methodology leads to a clearer approach to hierarchical classification of proteins and a novel way to represent evolutionary relationships. Comparative analysis of Jasmine and other statistical and deterministic systems (including Explanation-Based Learning and Inductive Logic Programming) are outlined. Advantages of using deterministic versus statistical data mining approaches for high-level exploration of correlation structure are analyzed.

Algorithms↗

RNA sequence evolution with secondary structure constraints: comparison of substitution rate models using maximum-likelihood methods.

We test models for the evolution of helical regions of RNA sequences, where the base pairing constraint leads to correlated compensatory substitutions occurring on either side of the pair. These models are of three types: 6-state models include only the four Watson-Crick pairs plus GU and UG; 7-state models include a single mismatch state that combines all of the 10 possible mismatches; 16-state models treat all mismatch states separately. We analyzed a set of eubacterial ribosomal RNA sequences with a well-established phylogenetic tree structure. For each model, the maximum-likelihood values of the parameters were obtained. The models were compared using the Akaike information criterion, the likelihood-ratio test, and Cox's test. With a high significance level, models that permit a nonzero rate of double substitutions performed better than those that assume zero double substitution rate. Some models assume symmetry between GC and CG, between AU and UA, and between GU and UG. Models that relaxed this symmetry assumption performed slightly better, but the tests did not all agree on the significance level. The most general time-reversible model significantly outperformed any of the simplifications. We consider the relative merits of all these models for molecular phylogenetics.

Base Pairing↗

The analysis of Circe, an LTR retrotransposon of Drosophila melanogaster, suggests that an insertion of non-LTR retrotransposons into LTR elements can create chimeric retroelements.

Circe is a transposable element recently identified in Drosophila melanogaster which appears to be mostly associated with the constitutive heterochromatin. This element shows the structural features of a long terminal repeat (LTR)-containing retrotransposon: It is flanked by 240-bp-long terminal repeats, and its two open reading frames encode putative proteins resembling the gag and pol polyproteins of retroviruses. However, Circe displays striking similarities of both LOA and Ulysses, a non-LTR element and an LTR element, respectively. The result of its phylogenetic and structural analysis has allowed us to propose a new mechanism for non-LTR retrotransposon evolution.

Amino Acid Sequence↗

The natural evolutionary relationships among prokaryotes.

Two contrasting and very different proposals have been put forward to account for the evolutionary relationships among prokaryotes. The currently widely accepted three domain proposal by Woese et al. (Proc. Natl. Acad. Sci. USA (1990) 87: 4576-4579) calls for the division of prokaryotes into two primary groups or domains, termed archaebacteria (Archaea) and eubacteria (Bacteria), both of which are suggested to have originated independently from a universal ancestor. However, this proposal, which is based primarily on genes involved in the information transfer processes, is inconsistent with the ultrastructural characteristics of prokaryotes as well as with many gene phylogenies and provides no explanation as to how the structural and molecular differences seen between these groups arose and how other prokaryotic taxa are related or evolved from the common ancestor. It also postulates that the last common ancestor of all organisms was a hypothetical entity lacking a cell membrane, which is contrary to the basic requirement of a cell membrane to define and separate all forms of life from the surrounding environment. A second alternate proposal for the evolutionary relationships among prokaryotes has emerged from extensive analyses of numerous conserved inserts and deletions found in various proteins (Gupta, R. S., Microbiol. Mol. Biol. Rev. (1998)62: 1435-1491; FEMS Microbiol. Rev. (2000) 24: in press. This proposal points to a specific relationship between archaebacteria and gram-positive bacteria, both of which are prokaryotes bounded by a single cell membrane (monoderm prokaryotes). Gram-negative bacteria, which are bounded by two different membranes (diderm prokaryotes), are indicated to comprise a structurally and phylogenetically distinct taxa originating from gram-positive bacteria. This proposal postulates that the earliest prokaryote was a gram-positive bacteria from which both archaebacteria and diderm prokaryotes evolved by normal evolutionary mechanisms in response to the strong selection pressure exerted by antibiotics produced by certain groups of gram-positive bacteria. This proposal accounts for both the molecular as well structural differences seen among the main groups of prokaryotes by known evolutionary mechanisms without invoking any hypothetical process or entity and thus is a closer representation of the natural relationships among prokaryotes than the proposal for two distinct domains. Based on this new proposal, it is now possible to logically deduce the branching order of different prokaryotic taxa from the common ancestor, which is as follows: Gram-positive bacteria (Low G + C) (<=> Archaebacteria) => Gram-positive bacteria (High G + C) (<=> Archaebacteria)=> Deinococcus-Thermus => Green nonsulfur bacteria => Cyanobacteria => Spirochetes => Chlamydia- Cytophaga-Green sulfur bacteria => Proteobacteria-1 (epsilon, delta)=> Proteobacteria-2 (alpha) => Proteobacteria-3 (beta) => Proteobacteria-4 (gamma). A surprising but very important aspect of the relationship deduced here is that the main eubacterial phyla are related to each other linearly rather than in a tree-like manner, suggesting that the major evolutionary changes within prokaryotes (bacteria) have occurred in a directional manner.

Archaea↗