Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “structural phylogenetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Postsynaptic alpha-neurotoxin gene of the spitting cobra, Naja naja sputatrix: structure, organization, and phylogenetic analysis.

The venom of the spitting cobra, Naja naja sputatrix contains highly potent alpha-neurotoxins (NTXs) in addition to phospholipase A2 (PLA2) and cardiotoxin (CTX). In this study, we report the complete characterization of three genes that are responsible for the synthesis of three isoforms of alpha-NTX in the venom of a single spitting cobra. DNA amplification by long-distance polymerase chain reaction (LD-PCR) and genome walking have provided information on the gene structure including their promoter and 5' and 3' UTRs. Each NTX isoform is approximately 4 kb in size and contains three exons and two introns. The sequence homology among these isoforms was found to be 99%. Two possible transcription sites were identified by primer extension analysis and they corresponded to the adenine (A) nucleotide at positions +1 and -45. The promoter also contains two TATA boxes and a CCAAT box. Putative binding sites for transcriptional factors AP-2 and GATA are also present. The high percentage of similarity observed among the NTX gene isoforms of N. n. sputatrix as well as with the alpha-NTX and kappa-NTX genes from other land snakes suggests that the NTX gene has probably evolved from a common ancestral gene.

3' Untranslated Regions↗

A model for phylogenetic inference using structural and chemical covariates.

We investigated whether or not evolutionary change in DNA sequence data was homogeneous across different classes of base pairs. DNA sequences for eight protein-coding mitochrondrial genes were obtained for 38 vertebrate taxa from GenBank. Each nucleotide site in the alignment was classified according to a number of covariates, including its codon position, genetic code degeneracy, and hydrophobicity. The evolutionary transition matrix for each base was estimated by tracing implied character changes under parsimony on a known phylogenetic tree. Canonical variates analyses of the inferred transition matrices were performed for each gene to determine whether or not different classes of bases behaved similarly. We found five distinct clusters of transition matrices that could be roughly defined by combinations of codon position and degeneracy. This pattern was consistent among all genes. A stochastic model of rate variation based on the interaction of the covariates was developed to assess the statistical significance of the clusters. The five-group classification was found to explain significantly more sequence variation than did a codon only classification, a codon degeneracy classification, or a codon and degeneracy classification. The same five-group classification was found for all genes tested, suggesting a common process underlying the molecular evolution of the mitochondrial genome. These results confirm that there are classes of base pairs that evolve differently, and suggest that models of sequence evolution that incorporate covariate information may be useful in developing nucleotide substitution models that more accurately reflect evolutionary history.

Animals↗

'Candidatus mycoplasma haemodidelphidis' sp. nov., 'Candidatus mycoplasma haemolamae' sp. nov. and Mycoplasma haemocanis comb. nov., haemotrophic parasites from a naturally infected opossum (Didelphis virginiana), alpaca (Lama pacos) and dog (Canis familiaris): phylogenetic and secondary structural relatedness of their 16S rRNA genes to other mycoplasmas.

The 16S rRNA sequence of newly characterized haemotrophic bacteria in an opossum (Didelphis virginiana) and alpaca (Lama pacos) was determined. In addition, the 16S rRNA sequence of a haemotrophic parasite in the dog (Canis familiaris) was determined. Sequence alignment and evolutionary analysis as well as secondary structural similarity and signature nucleotide sequence motifs of their 16S rRNA genes, positioned these organisms in the genus Mycoplasma. The highest scoring sequence similarities were 16S rRNA genes from haemotrophic mycoplasma species (Haemobartonella and Eperythrozoon spp.). However, the lack of several higher-order structural idiosyncrasies used to define the pneumoniae group, suggests that these organisms and related haemotrophic mycoplasmas represent a new group of mycoplasmas. It is recommended that the organisms be named 'Candidatus Mycoplasma haemodidelphidis', 'Candidatus Mycoplasma haemolamae' and Mycoplasma haemocanis comb. nov., to provide some indication of the target cell and host species of these parasites, and to reflect their phylogenetic affiliation.

Animals↗

Complete genomes, phylogenetic relatedness, and structural proteins of six strains of the hepatitis B virus, four of which represent two new genotypes.

The genomes of six hepatitis B viral (HBV) strains were sequenced from 10 overlapping amplificates obtained by the polymerase chain reaction. Four of the strains, specifying subtypes ayw4 and adw4q-, represented on the basis of divergency within the S gene two new genomic groups identified by us. The other two strains, encoding adrq- and of Pacific origin, belonged to genomic group C. The relation of these genomes to 21 published human, 1 chimpanzee, and 4 rodent hepadnaviral genomes was analyzed by constructing a phylogenetic dendrogram. Thereby, the segregation of human HBV strains into six genomic groups was confirmed. A consistent grouping of the genomes compared was also obtained in dendrograms based on the P and S genes, although the branching order differed from that based on the entire genomes. Each of the two representatives of genomic groups E and F differed by 8.1 to 13.6% and by 12.8 to 15.5% from the genomes of the other groups and by 1.5 and 3.7% from each other. The two Pacific group C strains differed by 2.7% from each other and by 4.1 to 5.4% from other group C genomes, suggesting that they diverged early from the other group C genomes. The F strains formed the most divergent group of HBV genomes, which may be explained by their representing the original strains of the New World. Within the structural gene products, 17 and 34 amino acids unique for human HBV strains were recorded in the sequenced E and F strains, respectively. Most notable is the Ser81 to Ala81 substitution in an immunodominant region of HBcAg, and the four extra cysteine residues in HBsAg at residues 19, 183, 206, and 220, which might be engaged in additional disulphide bridges. Five residues shared by E and F strains were also unique for human HBV strains. Two of these, Leu127 and Ser140 in HBsAg, were the only substitutions that may explain the w4 reactivity shared by these HBV strains. Interestingly, the Ser140 substitution occurs in an immunodominant loop of the a determinant claimed to be important for the protective immune response to HBV vaccination.

Amino Acid Sequence↗

M13 endopeptidases: New conserved motifs correlated with structure, and simultaneous phylogenetic occurrence of PHEX and the bony fish.

M13 endopeptidase alignments have focused mainly on mammalian sequences and on the active site region defining the catalytic sequence signatures. Aligning all available M13 from bacteria to human on a full-length basis, we have performed a sequence analysis. This enabled us to highlight the origin and function of the M13 PHEX subtype family endopeptidase (phosphate regulating gene with homologies to endopeptidases on the X chromosome). New evolutionary conserved regions in both prokaryotes and eukaryotes have been detected and eukaryotic-specific regions clearly delineated. Using the recently solved neprilysin structure, we have observed that all new motifs, except one, localize in the spatial vicinity of the previously reported catalytic signatures. Interestingly, a highly hydrophobic pocket containing three newly reported motifs is centered by the C-terminal tryptophan residue. Extensive M13 searches in complete and in progress higher eukaryotic genomes have lead to the identification of Danio rerio as the simplest organism having PHEX. Finally, the human PHEX substrate, the parathyroid hormone-related peptide, PTHrP(107-139), is absent in bony fish: this suggests the existence of further PHEX substrates common to both bony fishes and higher vertebrates.

Amino Acid Motifs↗

Structure prediction and phylogenetic analysis of a functionally diverse family of proteins homologous to the MT-A70 subunit of the human mRNA:m(6)A methyltransferase.

MT-A70 is the S-adenosylmethionine-binding subunit of human mRNA:m(6)A methyl-transferase (MTase), an enzyme that sequence-specifically methylates adenines in pre-mRNAs. The physiological importance yet limited understanding of MT-A70 and its apparent lack of similarity to other known RNA MTases combined to make this protein an attractive target for bioinformatic analysis. The sequence of MT-A70 was subjected to extensive in silico analysis to identify orthologous and paralogous polypeptides. This analysis revealed that the MT-A70 family comprises four subfamilies with varying degrees of interrelatedness. One subfamily is a small group of bacterial DNA:m(6)A MTases. The other three subfamilies are paralogous eukaryotic lineages, two of which have not been associated with MTase activity but include proteins having substantial regulatory effects. Multiple sequence alignments and structure prediction for members of all four subfamilies indicated a high probability that a consensus MTase fold domain is present. Significantly, this consensus fold shows the permuted topology characteristic of the b class of MTases, which to date has only been known to include DNA MTases.

Amino Acid Sequence↗

Molecular evolution of biomembranes: structural equivalents and phylogenetic precursors of sterols.

Derivatives of one triterpene family, the hopane family, are widely distributed in prokaryotes; they may be localized in membranes, playing there the same role as sterols play in eukaryotes, as a result of their similar size, rigidity, and amphiphilic character. Their biosynthesis embodies many primitive features compared to that of sterols and could have evolved toward the latter once aerobic conditions had been established. Membrane reinforcement appears to be achieved in other prokaryotes by other mechanisms, involving either approximately 40-A-long rigid hydrocarbon chains terminated by one polar group acting like a peg through the double-layer or similar chains terminated by two polar groups acting like tie-bars across the membrane. These inserts can be tetraterpenes (e.g., carotenoids). The biophysical function of membrane optimizers appears to have evolved toward sterols by changes limited to only a few enzymatic steps of the same fundamental biosynthetic processes.

Biological Evolution↗

Secondary structure probing of the human RNase MRP RNA reveals the potential for MRP RNA subsets.

RNase MRP is a ribonucleoprotein endoribonuclease involved in eukaryotic pre-rRNA processing. The enzyme possesses an RNA subunit, structurally related to that of RNase P RNA, that is thought to be catalytic. RNase MRP RNA sequences from Saccharomycetaceae species are structurally well defined through detailed phylogenetic and structural analysis. In contrast, higher eukaryote MRP RNA structure models are based on comparative sequence analysis of only five sequences and limited probing data. Detailed structural analysis of the Homo sapiens MRP RNA, entailing enzymatic and chemical probing, is reported. The data are consistent with the phylogenetic secondary structure model and demonstrate unequivocally that higher eukaryote MRP RNA structure differs significantly from that reported for Saccharomycetaceae species. Neither model can account for all of the known MRP RNAs and we thus propose the evolution of at least two subsets of RNase MRP secondary structure, differing predominantly in the predicted specificity domain.

Algorithms↗

Improving the precision of the structure-function relationship by considering phylogenetic context.

Understanding the relationship between protein structure and function is one of the foremost challenges in post-genomic biology. Higher conservation of structure could, in principle, allow researchers to extend current limitations of annotation. However, despite significant research in the area, a precise and quantitative relationship between biochemical function and protein structure has been elusive. Attempts to draw an unambiguous link have often been complicated by pleiotropy, variable transcriptional control, and adaptations to genomic context, all of which adversely affect simple definitions of function. In this paper, I report that integrating genomic information can be used to clarify the link between protein structure and function. First, I present a novel measure of functional proximity between protein structures (F-score). Then, using F-score and other entirely automatic methods measuring structure and phylogenetic similarity, I present a three-dimensional landscape describing their inter-relationship. The result is a "well-shaped" landscape that demonstrates the added value of considering genomic context in inferring function from structural homology. A generalization of methodology presented in this paper can be used to improve the precision of annotation of genes in current and newly sequenced genomes.

Journal Article↗

Characterization of novel GPCR gene coding locus in amphioxus genome: gene structure, expression, and phylogenetic analysis with implications for its involvement in chemoreception.

Chemosensation is the primary sensory modality in almost all metazoans. The vertebrate olfactory receptor genes exist as tandem clusters in the genome, so that identifying their evolutionary origin would be useful for understanding the expansion of the sensory world in relation to a large-scale genomic duplication event in a lineage leading to the vertebrates. In this study, I characterized a novel GPCR (G-protein-coupled receptor) gene-coding locus from the amphioxus genome. The genomic DNA contains an intronless ORF whose deduced amino acid sequence encodes a seven-transmembrane protein with some amino acid residues characteristic of vertebrate olfactory receptors (ORs). Surveying counterparts in the Ciona intestinalis (Asidiacea, Urochordata) genome by querying BLAST programs against the Ciona genomic DNA sequence database resulted in the identification of a remotely related gene. In situ hybridization analysis labeled primary sensory neurons in the rostral epithelium of amphioxus adults. Based on these findings, together with comparison of the developmental gene expression between amphioxus and vertebrates, I postulate that chemoreceptive primary sensory neurons in the rostrum are an ancient cell population traceable at least as far back in phylogeny as the common ancestor of amphioxus and vertebrates.

Amino Acid Sequence↗

Intergeneric complementation of a circadian rhythmicity defect: phylogenetic conservation of structure and function of the clock gene frequency.

The Neurospora crassa frequency locus encodes a 989 amino acid protein that is a central component, a state variable, of the circadian biological clock. We have determined the sequence of all or part of this protein and surrounding regulatory regions from additional fungi representing three genera and report that there is distinct, preferential conservation of the frequency open reading frame (ORF) as compared with non-coding sequences. Within the coding region, many of the domain hallmarks of the N. crassa protein are highly conserved, especially an internal region bearing the causative mutations in frq1 and frq7, the most extreme alleles in the frequency allelic series. Despite considerable diversity among the strains analyzed in terms of morphology, growth, circadian clock output and frq sequence, the ORF from the most distantly related fungus included in this study (Sordaria fimicola) rescues rhythmicity in a N.crassa frequency null strain. Both sequence conservation, and the ability of frequency from a genus displaying one developmental program to complement circadian defects in a separate genus with a distinct, clock-regulated developmental program, are consistent with a central role of the frequency gene product in a general circadian oscillator capable of controlling diverse outputs in a variety of systems.

Amino Acid Sequence↗

Early evolutionary relationships among known life forms inferred from elongation factor EF-2/EF-G sequences: phylogenetic coherence and structure of the archaeal domain.

Phylogenies were inferred from both the gene and the protein sequences of the translational elongation factor termed EF-2 (for Archaea and Eukarya) and EF-G (for Bacteria). All treeing methods used (distance-matrix, maximum likelihood, and parsimony), including evolutionary parsimony, support the archaeal tree and disprove the "eocyte tree" (i.e., the polyphyly and paraphyly of the Archaea). Distance-matrix trees derived from both the amino acid and the DNA sequence alignments (first and second codon positions) showed the Archaea to be a monophyletic-holophyletic grouping whose deepest bifurcation divides a Sulfolobus branch from a branch comprising Methanococcus, Halobacterium, and Thermoplasma. Bootstrapped distance-matrix treeing confirmed the monophyly-holophyly of Archaea in 100% of the samples and supported the bifurcation of Archaea into a Sulfolobus branch and a methanogen-halophile branch in 97% of the samples. Similar phylogenies were inferred by maximum likelihood and by maximum (protein and DNA) parsimony. DNA parsimony trees essentially identical to those inferred from first and second codon positions were derived from alternative DNA data sets comprising either the first or the second position of each codon. Bootstrapped DNA parsimony supported the monophyly-holophyly of Archaea in 100% of the bootstrap samples and confirmed the division of Archaea into a Sulfolobus branch and a methanogen-halophile branch in 93% of the bootstrap samples. Distance-matrix and maximum likelihood treeing under the constraint that branch lengths must be consistent with a molecular clock placed the root of the universal tree between the Bacteria and the bifurcation of Archaea and Eukarya. The results support the division of Archaea into the kingdoms Crenarchaeota (corresponding to the Sulfolobus branch and Euryarchaeota). This division was not confirmed by evolutionary parsimony, which identified Halobacterium rather than Sulfolobus as the deepest offspring within the Archaea.

Amino Acid Sequence↗

Structure, expression, and phylogenetic relationships of a family of ypt genes encoding small G-proteins in the green alga Volvox carteri.

In addition to the previously described gene yptV1 encoding a small G-protein we have now identified and sequenced four more ras-related ypt genes (yptV2-yptV5) from the green alga Volvox carteri. The four new genes encode polypeptides consisting of 203 to 217 amino-acid residues that contain the typical sequence elements (GTP-binding domains, effector domain) of the ypt/rab subgroup of the Ras superfamily. Comparison of the derived amino-acid sequences from the V. carteri ypt gene products and their Ypt homologs from other species revealed similarity values ranging from 60% to 85%, whereas intraspecies similarities were found to approach only 55%. The coding sequences are interrupted by 5-7 introns of variable size (70-1000 nucleotides) occupying different positions in the genes. Reverse-transcribed samples of stage-specific RNAs were PCR-amplified with primers specific to yptV1, yptV3, yptV4, and yptV5 to determine if yptV transcription might be restricted to either cell type or to a specific stage of the life cycle. These experiments demonstrated that each of these genes is expressed throughout the entire Volvox life cycle and in both the somatic and the reproductive cells of the alga. The transcription start sites of yptV1 and yptV5 were mapped by primer extension. Expression of recombinant yptV cDNA in E. coli yielded recombinant proteins that bound GTP specifically, demonstrating a property which is typical for small G-proteins. The derived YptV polypeptide sequences were used to group them into four distinct classes of Ras-like proteins. These are the first proteins of the Ras superfamily to be identified in a green alga. We discuss the possible role of the YptV-proteins in the intracellular vesicle transport of Volvox.

Amino Acid Sequence↗

Pectin methylesterases: sequence-structural features and phylogenetic relationships.

Pectin methylesterases (PMEs) are enzymes produced by bacteria, fungi and higher plants. They belong to the carbohydrate esterase family CE-8. This study deals with comparison of 127 amino acid sequences of this family containing the five characteristic sequence segments: 44_GxYxE, 113_QAVAL, 135_QDTL, 157_DFIFG, 223_LGRPW (Daucus carota numbering). Six strictly conserved residues (Gly44, Gly154, Asp157, Gly161, Arg225 and Trp227) and six conservative ones (Ile39, Ser86, Ser137, Ile152, Ile159 and Leu223) were identified. A set of 70 representative PMEs was created. The sequences were aligned and the evolutionary tree based on the alignment was calculated. The tree reflected the taxonomy: the fungal and bacterial PMEs formed their own clusters and the plant enzymes were grouped into eight separate clades. The plant PME from Vitis riparia was placed in a common clade with fungi. Three plant clades (Plant 1, 2 and 3) were relatively homogenous reflecting high degree of mutual sequence identity. The clade Plant 4 contained PMEs from flower parts (mostly form pollen) and was heterogenous, like the clades Plant 1a and 2a, which moreover exhibit an intermediate character. The clades Plant X1 and X2 were situated in the tree close to microbial clades and represented atypical plant PMEs. Taking into account the remaining plant PMEs, an expanded plant alignment and tree (with most Arabidopsis thaliana and Oryza sativa enzymes), were prepared. An exclusive Arabidopsis alignment and tree indicated the existence of a new plant clade X3. In the pre pro region of most plant enzymes a longer conserved segment containing basic dipeptide, R(K)/R(K), that precedes the N-terminal end of PME was revealed. This was not observed in the clade Plant X1 and majority of the clade Plant X2. This study brings further the description of occurrence of potential glycosylation sites in pre pro sequences and in mature enzymes as well as important amino acid residues, such as aspartates, cysteines, histidines and other aromatic residues (Tyr, Phe and Trp), with discussion of their possible function in the activity of PMEs.

Amino Acid Sequence↗

Structure, expression and phylogenetic analysis of the glycoprotein gene of Cocal virus.

A cDNA copy of the mRNA of the glycoprotein G of Cocal virus, a rhabdovirus, has been cloned, sequenced and expressed in mammalian cells. The deduced amino acid sequence shows a typical transmembrane glycoprotein, 512 amino acids in length, containing two potential N-linked glycosylation sites. The amino acid sequence showed a high degree of identity with that of the prototype vesicular stomatitis virus serotype Indiana [VSV (IND)] G protein. In addition, phylogenetic analysis of amino acid sequence differences among the G proteins of vesiculoviruses indicated that Cocal virus represents a distinct lineage within the VSV (IND) serotype. Expression of the cloned Cocal G gene in mammalian cells produced a glycoprotein of mol.wt 71000 which was not palmitylated but induced cell fusion at acid pH.

Amino Acid Sequence↗

Predicting the three-dimensional folding of transfer RNA with a computer modeling protocol.

We have developed a computer modeling protocol that can be used to predict the three-dimensional folding of a ribonucleic acid on the basis of limited amounts of secondary and tertiary data. This protocol extends the use of distance geometry beyond the domain of NMR data in which it is usually applied. The use of this algorithm to fold the molecule eliminates operator subjectivity and reproducibly predicts the overall dimensions and shape of the transfer RNA molecule. By use of a replacement pseudoatom set based on helical substructures, a series of transfer RNA foldings have been completed that utilize only the primary structure, the phylogenetically deduced secondary structure, and five long-range interactions that were determined without reference to the crystal structure. In a control set of foldings, all the interactions suspected to exist in 1969 have been included. In all cases, the modeling process consistently predicts the global arrangement of the helical domains and to a lesser extent the general path of the backbone of transfer RNA.

Computer Simulation↗

Saccharomyces cerevisiae U1 small nuclear RNA secondary structure contains both universal and yeast-specific domains.

The five small nuclear RNAs (snRNAs) involved in mammalian pre-mRNA splicing (U1, U2, U4, U5, and U6) are well conserved in length, sequence, and especially secondary structure. These five snRNAs from Saccharomyces cerevisiae show notable size and sequence differences from their metazoan counterparts. This is most striking for the large S. cerevisiae U1 and U2 snRNAs, for which no secondary structure models currently exist. Because of the importance of U1 snRNA in the early steps of "spliceosome" assembly, we wanted to compare the highly conserved secondary structure of metazoan U1 snRNA (approximately 165 nucleotides) with that of S. cerevisiae U1 snRNA (568 nucleotides). To this end, we have cloned and sequenced the U1 gene from two other yeast species possessing large U1 RNAs. Using computer-derived structure predictions, phylogenetic comparisons, and structure probing, we have arrived at a secondary structure model for S. cerevisiae U1 snRNA. The results show that most elements of higher eukaryotic U1 snRNA secondary structure are conserved in S. cerevisiae. The hundreds of "extra" nucleotides of yeast U1 RNA, also highly structured, suggest that large insertions and/or deletions have occurred during the evolution of the U1 gene.

Base Sequence↗