Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “structural phylogenetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

A phylogenetically based secondary structure for the yeast telomerase RNA.

BACKGROUND: Telomerase is a ribonucleoprotein complex whose RNA moiety dictates the addition of specific simple sequences onto chromosomes ends. While relevant for certain human genetic diseases, the contribution of the essential telomerase RNA to RNP assembly still remains unclear. Phylogenetic analyses of vertebrate and ciliate telomerase RNAs revealed conserved elements that potentially organize protein subunits for RNP function. In contrast, the yeast telomerase RNA could not be fitted to any known structural model, and the limited number of known sequences from Saccharomyces species did not permit the prediction of a yeast specific conserved structure. RESULTS: We cloned and analyzed the complete telomerase RNA loci (TLC1) from all known Saccharomyces species belonging to the "sensu stricto" group. Complementation analyses in S. cerevisiae and end mappings of mature RNAs ensured the relevance of the cloned sequences. By using phylogenetic comparative analysis coupled with in vitro enzymatic probing, we derived a secondary structure prediction of the Saccharomyces cerevisiae TLC1 RNA. This conserved secondary structure prediction includes a central domain that is likely to orchestrate DNA synthesis and at least two accessory domains important for RNA stability and telomerase recruitment. The structure also reveals a potential tertiary interaction between two loops in the central core. CONCLUSIONS: The predicted secondary structure of the TLC1 RNA of S. cerevisiae reveals a distinct folding pattern featuring well-separated but conserved functional elements. The predicted structure now allows for a detailed and rationally designed study to the structure-function relationships within the telomerase RNP-complex in a genetically tractable system.

Base Sequence↗

The puf operon of the purple sulfur bacterium Amoebobacter purpureus: structure, transcription and phylogenetic analysis.

The puf operon, encoding photosynthetic reaction center and light-harvesting genes, of the purple sulfur phototrophic bacterium Amoebobacter purpureus was cloned and sequenced. This revealed an unusual operon structure of the genes pufB1 A1 LMCB2 A2 B3 A3. The sequence represents the second complete puf operon available for Chromatiaceae. So far, additional sets of light-harvesting 1 (LH1) genes, pufB2 A2 and pufB3 A3 in the region downstream of pufC have only been described for Allochromatium vinosum. Along with reports of multiple LH1 polypeptides found in some Ectothiorhodospiraceae by direct protein sequencing, our results indicate that multiple LH1 genes may occur frequently in phototrophic gamma-proteobacteria. Phylogenetic analyses suggested a coevolution of the core puf genes pufB1 A1 LM. Separate analysis of the LH1 alpha and beta polypeptides revealed a high intraspecies relatedness for the secondary LH1beta polypeptides, possibly caused by functional constraints. In contrast, LH1alpha subunits of Amb. purpureus and Alc. vinosum are closely related (85% sequence identity) which could reflect horizontal gene transfer. RNA analyses suggested co-transcription of all puf genes in Amb. purpureus as a 5.5 kb primary transcript which appears to be more stable than the puf operon primary transcripts of purple non-sulfur bacteria. The 5' end of the transcript mapped to a putative promoter, which contains a -35 region located in an inverted repeat DNA sequence.

Amino Acid Sequence↗

The evolution of SMC proteins: phylogenetic analysis and structural implications.

The SMC proteins are found in nearly all living organisms examined, where they play crucial roles in mitotic chromosome dynamics, regulation of gene expression, and DNA repair. We have explored the phylogenetic relationships of SMC proteins from prokaryotes and eukaryotes, as well as their relationship to similar ABC ATPases, using maximum-likelihood analyses. We have also investigated the coevolution of different domains of eukaryotic SMC proteins and attempted to account for the evolutionary patterns we have observed in terms of available structural data. Based on our analyses, we propose that each of the six eukaryotic SMC subfamilies originated through a series of ancient gene duplication events, with the condensins evolving more rapidly than the cohesins. In addition, we show that the SMC5 and SMC6 subfamily members have evolved comparatively rapidly and suggest that these proteins may perform redundant functions in higher eukaryotes. Finally, we propose a possible structure for the SMC5/SMC6 heterodimer based on patterns of coevolution.

ATP-Binding Cassette Transporters↗

A comparison of optimal and suboptimal RNA secondary structures predicted by free energy minimization with structures determined by phylogenetic comparison.

This article describes the latest version of an RNA folding algorithm that predicts both optimal and suboptimal solutions based on free energy minimization. A number of RNA's with known structures deduced from comparative sequence analysis are folded to test program performance. The group of solutions obtained for each molecule is analysed to determine how many of the known helixes occur in the optimal solution and in the best suboptimal solution. In most cases, a structure about 80% correct is found with a free energy within 2% of the predicted lowest free energy structure.

Algorithms↗

Islet hormones from the African bullfrog Pyxicephalus adspersus (Anura:Ranidae): structural characterization and phylogenetic implications.

The African bullfrog Pyxicephalus adspersus is generally classified along with frogs of the genus Rana in the subfamily Raninae of the family Ranidae but precise phylogenetic relationships between species are unclear. Pancreatic polypeptide (PP), insulin, and glucagon-like peptide (GLP-1) were isolated from an extract of P. adspersus pancreas and characterized structurally. A comparison of the amino acid sequence of Pyxicephalus PP (APSEPQHPGG(10)QATPEQLAQY(20)YSDLYQYITF(30)ITRPRF++ +. NH(2)) with those of the known amphibian PP molecules in a maximum parsimony analysis generates a single phylogenetic tree in which Pyxicephalus is the sister to the clade comprising the members of the genus Rana. The three orders of living amphibians form discrete clades with the representative of the Gymnophiona appearing as sister to the Caudata-Anura. In contrast, Pyxicephalus insulin (A chain, GIVEQCCHSA(10)CSLYDLENYC(20)N; B-chain, LANQHLCGSH(10)LVEALYMVCG(20)ERGFFYYPKS(30)) and and GLP-1 (HAEGTFTSDM(10)TSYLEEKAAK(20)EFVDWLIKGR(30)PK) resemble more closely the corresponding peptides from the cane toad Bufo marinus than the peptides from any species of Rana. Cladistic analysis based upon the amino acid sequences of insulin produced a polyphyletic assemblage with the Gymnophiona nesting within an unresolved clade containing the non-ranid frogs. The data support the assertion that the amino acid sequence of PP, but not those of the other islet hormones, is of value as a molecular marker for inferring phylogenetic relationships between early tetrapod species.

Amino Acid Sequence↗

Extreme specificity in epiparasitic Monotropoideae (Ericaceae): widespread phylogenetic and geographical structure.

The Monotropoideae (Ericaceae) are nonphotosynthetic plants that obtain fixed carbon from their fungal mycorrhizal associates. To infer the evolutionary history of this symbiosis we identified both the plant and fungal lineages involved using a molecular phylogenetic approach to screen 331 plants, representing 10 of the 12 described species. For five species no prior molecular data were available; for three species we confirmed prior studies which used limited samples; for five species all previous reports are in conflict with our results, which are supported by sequence analysis of multiple samples and are consistent with the phylogenetic patterns of host plants. The phylogenetic patterns observed indicate that: (i) each of the 13 plant phylogenetic lineages identified is specialized to a different genus or species group within five families of ectomycorrhizal Basidiomycetes; (ii) mycorrhizal specificity is correlated with phylogeny; (iii) in sympatry, there is no overlap in mature plant fungal symbionts even if the fungi and the plants are closely related; and (iv) there are geographical patterns to specificity.

Basidiomycota↗

Structural analysis of phylogenetically conserved J domain protein gene.

Novel cDNAs encoding evolutionarily conserved J Domain Proteins (JDPs) were investigated from Drosophila and mouse. Each of the full coding sequences potentially encodes a conserved J domain, but lacks additional characteristic structures present in DnaJ family proteins. The expression was restricted to head in Drosophila. However, ubiquitous expression was observed in mice with the highest level in kidney.

Amino Acid Sequence↗

The cation/Ca(2+) exchanger superfamily: phylogenetic analysis and structural implications.

Cation/Ca(2+) exchangers are an essential component of Ca(2+) signaling pathways and function to transport cytosolic Ca(2+) across membranes against its electrochemical gradient by utilizing the downhill gradients of other cation species such as H(+), Na(+), or K(+). The cation/Ca(2+) exchanger superfamily is composed of H(+)/Ca(2+) exchangers and Na(+)/Ca(2+) exchangers, which have been investigated extensively in both plant cells and animal cells. Recently, information from completely sequenced genomes of bacteria, archaea, and eukaryotes has revealed the presence of genes that encode homologues of cation/Ca(2+) exchangers in many organisms in which the role of these exchangers has not been clearly demonstrated. In this study, we report a comprehensive sequence alignment and the first phylogenetic analysis of the cation/Ca(2+) exchanger superfamily of 147 sequences. The results present a framework for structure-function relationships of cation/Ca(2+) exchangers, suggesting unique signature motifs of conserved residues that may underlie divergent functional properties. Construction of a phylogenetic tree with inclusion of cation/Ca(2+) exchangers with known functional properties defines five protein families and the evolutionary relationships between the members. Based on this analysis, the cation/Ca(2+) exchanger superfamily is classified into the YRBG, CAX, NCX, and NCKX families and a newly recognized family, designated CCX. These findings will provide guides for future studies concerning structures, functions, and evolutionary origins of the cation/Ca(2+) exchangers.

Amino Acid Sequence↗

Reticulate phylogenetics and phytogeographical structure of Heliosperma (Sileneae, Caryophyllaceae) inferred from chloroplast and nuclear DNA sequences.

The Balkan Peninsula is known to be one of the most diverse and species-rich parts of Europe, but its biota has gained much less attention in phylogenetic and evolutionary studies compared to other southern European mountain systems. We used nuclear ribosomal internal transcribed spacer (ITS) sequences and intron sequences of the chloroplast gene rps16 to examine phylogenetic and biogeographical patterns within the genus Heliosperma (Sileneae, Caryophyllaceae). The ITS and rps16 intron sequences both support monophyly of Heliosperma, but the data are not conclusive with regard to its exact origin. Three strongly supported clades are found in both data sets, corresponding to Heliosperma alpestre, Heliosperma macranthum and the Heliosperma pusillum clade, including all other taxa. The interrelationships among these three differ between the nuclear and the plastid data sets. Hierarchical relationships within the H. pusillum clade are poorly resolved by the ITS data, but the rps16 intron sequences form two well-supported clades which are geographically, rather than taxonomically, correlated. A similar geographical structure is found in the ITS data, when analyzed with the NeighbourNet method. The apparent rate of change within Heliosperma is slightly higher for rps16 as compared to ITS. In contrast, in the Sileneae outgroup, ITS substitution rates are more than twice as high as those for rps16, a situation more in agreement with what has been found in other rate comparisons of noncoding cpDNA and ITS. Unlike most other Sileneae ITS sequences, the H. pusillum group sequences display extensive polymorphism. A possible explanation to these patterns is extensive hybridization and gene flow within Heliosperma, which together with concerted evolution may have eradicated the ancient divergence suggested by the rps16 data. The morphological differentiation into high elevation, mainly widely distributed taxa, and low elevation narrow endemics is not correlated with the molecular data, and is possibly a result of ecological differentiation.

Base Sequence↗

Phylogenetic conservation of RNA secondary and tertiary structure in the trpEDCFBA operon leader transcript in Bacillus.

Expression of the trpEDCFBA operon of Bacillus subtilis is regulated by transcription attenuation and translation control mechanisms. We recently determined that the B. subtilis trp leader readthrough transcript can adopt a Mg(2+)-dependent tertiary structure that appears to interfere with TRAP-mediated translation control of trpE. In the present study, sequence comparisons to trp leaders from three other Bacillus sp. were made, suggesting that RNA secondary and tertiary structures are phylogenetically conserved. To test this hypothesis, experiments were carried out with the trp leader transcript from Bacillus stearothermophilus. Structure mapping experiments confirmed the predicted secondary structure. Native gel experiments identified a faster mobility species in the presence of Mg(2+), suggesting that a Mg(2+)-dependent tertiary structure forms. Mg(2+)-dependent protection of residues within the first five triplet repeats of the TRAP binding target and a pyrimidine-rich internal loop were observed, consistent with tertiary structure formation between these regions. Structure mapping in the presence of a competitor DNA oligonucleotide allowed the interacting partners to be identified as a single-stranded portion of the purine-rich TRAP binding target and the large downstream pyrimidine-rich internal loop. Thermal denaturation experiments revealed a Mg(2+)- and pH-dependent unfolding transition that was absent for a transcript missing the first five triplet repeats. The stability of several mutant transcripts allowed a large portion of the base-pairing register for the tertiary interaction to be determined. These data indicate that RNA secondary and tertiary structures involved in TRAP-mediated translation control are conserved in at least four Bacillus species.

Base Sequence↗

A combined analysis of genomic and primary protein structure defines the phylogenetic relationship of new members if the T-box family.

T-box genes form an ancient family of putative transcriptional regulators characterized by a region of homology to the DNA-binding domain of the murine Brachyury (T) gene product. This T-box domain is conserved from Caenorhabditis elegans to human, and mutations in T-box genes have been associated with developmental defects in Drosophila, zebrafish, mice, and humans. Here we report the identification of three novel murine T-box genes and an investigation of their evolutionary relationship to previously known family members by studying the genomic structure of the T-box. All T-box genes from nematodes to humans possess a characteristic central intron that presumably was inherited from a common ancestral precursor. Two additional intron positions are also conserved with the exception of two nematode T-box genes. Subsequent intron insertions, potential deletions, and/or intron sliding formed a structural basis for the divergence into distinct subfamilies and a substrate for length variations of the T-box domain. In mice, the 11 T-box genes known to date can be grouped into seven subfamilies. Genes assigned to the same subfamily by genomic structure show related expression patterns. We propose a model for the phylogenetic relationships within the gene family that provides a rationale for classifying new T-box genes and facilitates interspecific comparisons.

Amino Acid Sequence↗

The C terminus of the nuclear protein NuMA: phylogenetic distribution and structure.

The C terminus of the nuclear protein NuMA, NuMA-CT, has a well-known function in mitosis via its proximal segment, but it seems also involved in the control of differentiation. To further investigate the structure and function of NuMA, we exploited established computational techniques and tools to collate and characterize proteins with regions similar to the distal portion of NuMA-CT (NuMA-CTDP). The phylogenetic distribution of NuMA-CTDP was examined by PSI-BLAST- and TBLASTN-based analysis of genome and protein sequence databases. Proteins and open reading frames with a NuMA-CTDP-like region were found in a diverse set of vertebrate species including mammals, birds, amphibia, and early teleost fish. The potential structure of NuMA-CTDP was investigated by searching a database of protein sequences of known three-dimensional structure with a hidden Markov model (HMM) estimated using representative (human, frog, chicken, and pufferfish) sequences. The two highest scoring sequences that aligned to the HMM were the extracellular domains of beta3-integrin and Her2, suggesting that NuMA-CTDP may have a primarily beta fold structure. These data indicate that NuMA-CTDP may represent an important functional sequence conserved in vertebrates, where it may act as a receptor to coordinate cellular events.

Amino Acid Sequence↗

Statistical modeling, phylogenetic analysis and structure prediction of a protein splicing domain common to inteins and hedgehog proteins.

Inteins, introns spliced at the protein level, and the hedgehog family of proteins involved in eucaryotic development both undergo autocatalytic proteolysis. Here, a specific and sensitive hidden Markov model (HMM) of protein splicing domain shared by inteins and the hedgehog proteins has been trained and employed for further analysis. The HMM characterizes the common features of this domain including the position where a site-specific DNA endonuclease domain is inserted in the majority of the inteins. The HMM was used to identify several new putative inteins, such as that in the Methanococcus jannaschii klbA protein, and to generate a multiple sequence alignment of sequences possessing this domain. Phylogenetic analysis suggests that hedgehog proteins evolved from inteins. Secondary and tertiary structure predictions suggest that the domain has a structure similar to a beta-sandwich. Similarities between the serine protease cleavage mechanism and the protein splicing reaction mechanism are discussed. Examination of the locations of inteins indicates that they are not inserted randomly in an extein, but are often inserted at functionally important positions in the host proteins. A specific and sensitive HMM for a domain present in klbA proteins identified several additional bacterial and archaeal family members, and analysis of the site of insertion of the intein suggests residues that may be functionally important. This domain may play a role in formation of surface-associated protein complexes.

Algorithms↗

Structure of the phylogenetically most conserved domain of SRP RNA.

The signal recognition particle (SRP) is a phylogenetically conserved ribonucleoprotein required for cotranslational targeting of proteins to the membrane of the endoplasmic reticulum of the bacterial plasma membrane. Domain IV of SRP RNA consists of a short stem-loop structure with two internal loops that contain the most conserved nucleotides of the molecule. All known essential interactions of SRP occur in that moiety containing domain IV. The solution structure of a 43-nt RNA comprising the complete Escherichia coli domain IV was determined by multidimensional NMR and restrained molecular dynamics refinement. Our data confirm the previously determined rigid structure of a smaller subfragment containing the most conserved, symmetric internal loop A (Schmitz et al., Nat Struct Biol, 1999, 6:634-638), where all conserved nucleotides are involved in nucleotide-specific structural interactions. Asymmetric internal loop B provides a hinge in the RNA molecule; it is partially flexible, yet also uniquely structured. The longer strand of internal loop B extends the major groove by creating a ledge-like arrangement; for loop B however, there is no obvious structural role for the conserved nucleotides. The structure of domain IV suggests that loop A is the initial site for the RNA/protein interaction creating specificity, whereas loop B provides a secondary interaction site.

Base Sequence↗

Granule-bound starch synthase: structure, function, and phylogenetic utility.

Interest in the use of low-copy nuclear genes for phylogenetic analyses of plants has grown rapidly, because highly repetitive genes such as those commonly used are limited in number. Furthermore, because low-copy genes are subject to different evolutionary processes than are plastid genes or highly repetitive nuclear markers, they provide a valuable source of independent phylogenetic evidence. The gene for granule-bound starch synthase (GBSSI or waxy) exists in a single copy in nearly all plants examined so far. Our study of GBSSI had three parts: (1) Amino acid sequences were compared across a broad taxonomic range, including grasses, four dicotyledons, and the microbial homologs of GBSSI. Inferred structural information was used to aid in the alignment of these very divergent sequences. The informed alignments highlight amino acids that are conserved across all sequences, and demonstrate that structural motifs can be highly conserved in spite of marked divergence in amino acid sequence. (2) Maximum-likelihood (ML) analyses were used to examine exon sequence evolution throughout grasses. Differences in probabilities among substitution types and marked among-site rate variation contributed to the observed pattern of variation. Of the parameters examined in our set of likelihood models, the inclusion of among-site rate variation following a gamma distribution caused the greatest improvement in likelihood score. (3) We performed cladistic parsimony analyses of GBSSI sequences throughout grasses, within tribes, and within genera to examine the phylogenetic utility of the gene. Introns provide useful information among very closely related species, but quickly become difficult to align among more divergent taxa. Exons are variable enough to provide extensive resolution within the family, but with low bootstrap support. The combined results of amino acid sequence comparisons, maximum-likelihood analyses, and phylogenetic studies underscore factors that might affect phylogenetic reconstruction. In this case, accommodation of the variable rate of evolution among sites might be the first step in maximizing the phylogenetic utility of GBSSI.

Amino Acid Sequence↗

Conserved features of Y RNAs: a comparison of experimentally derived secondary structures.

In this study, phylogenetically conserved structural features of the Ro RNP associated Y RNAs were investigated. The human, iguana, and frog Y3 and Y4 RNA sequences have been determined previously and the respective RNAs were subjected to enzymatic and chemical probing to obtain structural information. For all of the analyzed RNAs, the probing data were used to compose secondary structures, which partly deviate from previously predicted structures. Our results confirm the existence of two stem structures, which are also found at similar positions in hY1 and hY5 RNA. For the remaining parts of hY3 and hY4 RNA the secondary structures differ from those previously proposed based upon computer predictions. What might be more important is that certain parts of the RNAs appear to be flexible, i.e., to adopt several conformations. Another striking feature is that a characteristic pyrimidine-rich region, present in every Y RNA known, is single-stranded in all secondary structures. This may suggest that this region is readily available for base pairing inter-actions with other cellular nucleic acids, which might be important for the as yet unknown function of the RNAs.

Animals↗

Ribosomal RNA sequences of Enterocytozoon bieneusi, Septata intestinalis and Ameson michaelis: phylogenetic construction and structural correspondence.

The microsporidian species Enterocytozoon bieneusi, Septata intestinalis and Ameson michaelis were compared by using sequence data of their rRNA gene segments, which were amplified by polymerized chain reaction and directly sequenced. The forward primer 530f (5'-GTGCCATCCAGCCGCGG-3') was in the small subunit rRNA (SSU-rRNA) and the reverse primer 580r (5'-GGTCCGTGTTTCAAGACGG-3') was in the large subunit rRNA (LSU-rRNA). We have utilized these sequence data, the published data on Encephalitozoon cuniculi and Encephalitozoon hellem and our cloned SSU-rRNA genes from E. bieneusi and S. intestinalis to develop a phylogenetic tree for the microsporidia involved in human infection. The higher sequence similarities demonstrated between S. intestinalis and E. cuniculi support the placement of S. intestinalis in the family Encephalitozoonidae. This method of polymerized chain reaction rRNA phylogeny allows the establishment of phylogenetic relationships on limiting material where culture and electron microscopy are difficult or impossible and can be applied to archival material to expand the molecular phylogenetic analysis of the phylum Microspora. In addition, the highly variable region (E. coli numbering 590-650) and intergenic spacer regions in the microsporidia were noted to have structural correspondence, suggesting the possibility that they are coevolving.

AIDS-Related Opportunistic Infections↗

Structural, functional, and phylogenetic characterization of a large CBF gene family in barley.

CBFs are key regulators in the Arabidopsis cold signaling pathway. We used Hordeum vulgare (barley), an important crop and a diploid Triticeae model, to characterize the CBF family from a low temperature tolerant cereal. We report that barley contains a large CBF family consisting of at least 20 genes (HvCBFs) comprising three multigene phylogenetic groupings designated the HvCBF1-, HvCBF3-, and HvCBF4-subgroups. For the HvCBF1- and HvCBF3-subgroups, there are comparable levels of phylogenetic diversity among rice, a cold-sensitive cereal, and the cold-hardy Triticeae. For the HvCBF4-subgroup, while similar diversity levels are observed in the Triticeae, only a single ancestral rice member was identified. The barley CBFs share many functional characteristics with dicot CBFs, including a general primary domain structure, transcript accumulation in response to cold, specific binding to the CRT motif, and the capacity to induce cor gene expression when ectopically expressed in Arabidopsis. Individual HvCBF genes differed in response to abiotic stress types and in the response time frame, suggesting different sets of HvCBF genes are employed relative to particular stresses. HvCBFs specifically bound monocot and dicot cor gene CRT elements in vitro under both warm and cold conditions; however, binding of HvCBF4-subgroup members was cold dependent. The temperature-independent HvCBFs activated cor gene expression at warm temperatures in transgenic Arabidopsis, while the cold-dependent HvCBF4-subgroup members of three Triticeae species did not. These results suggest that in the Triticeae - as in Arabidopsis - members of the CBF gene family function as fundamental components of the winter hardiness regulon.

Amino Acid Sequence↗