Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

A putative Drosophila homolog of the Huntington's disease gene.

The Huntington's disease (HD) gene encodes a protein, huntingtin, with no known function and no detectable sequence similarity to other proteins in current databases. To gain insight into the normal biological role of huntingtin, we isolated and sequenced a cDNA encoding a protein that is a likely homolog of the HD gene product in Drosophila melanogaster. We also determined the complete sequence of 43 125 contiguous base pairs of genomic DNA that encompass the Drosophila HD gene, allowing the intron-exon structure and 5'- and 3'-flanking regions to be delineated. The predicted Drosophila huntingtin protein has 3583 amino acids, which is several hundred amino acids larger than any other previously characterized member of the HD family. Analysis of the genomic and cDNA sequences indicates that Drosophila HD has 29 exons, compared with the 67 exons present in vertebrate HD genes, and that Drosophila huntingtin lacks the polyglutamine and polyproline stretches present in its mammalian counterparts. The Drosophila HD mRNA is expressed in a broad range of developmental stages and in the adult, a temporal pattern of expression similar to that observed for mammalian HD transcripts. We can discern five regions of high similarity from multiple sequence alignments between Drosophila and vertebrate huntingtins. These regions may define functionally important domains within the protein.

Amino Acid Sequence↗

Analysis of genetic diversity in cytotoxin-producing and non-cytotoxin-producing Helicobacter pylori strains.

Analysis of 32 Helicobacter pylori strains indicated a strong association between the presence of the cagA gene and a specific type of vacA allele found predominantly in cytotoxin-producing strains (P < .001). To determine whether tox+/CagA+ and tox-/CagA- strains constituted two separate noncombining lineages, sequences of the H. pylori ureC gene, cysS homologue, and the intergenic region between cysS and vacA were determined for multiple strains. The mean levels of nucleotide identity in the three regions were 96.7% +/- 0.5%, 95.0% +/- 1.0%, and 89.0% +/- 2.9%, respectively. Multiple sequence alignments and dendrograms based on these three regions failed to identify two clonal populations of organisms for which cagA and vacA genotypes were markers. The presence of a 63- to 64-bp insertion in the cysS-vacA intergenic region was unrelated to the vacA genotype of the strains. These data suggest that recombination between Helicobacter genomes may occur in vivo.

Base Sequence↗

Arsenite oxidase, an ancient bioenergetic enzyme.

Operons coding for the enzyme arsenite oxidase have been detected in the genomes from Archaea and Bacteria by Blast searches using the amino acid sequences of the respective enzyme characterized in two different beta-proteobacteria as templates. Sequence analyses show that in all these species, arsenite oxidase is transported over the cytoplasmic membrane via the tat system and most probably remains membrane attached by an N-terminal transmembrane helix of the Rieske subunit. The biochemical and biophysical data obtained for arsenite oxidase in the green filamentous bacterium Chloroflexus aurantiacus allow a structural model of the enzyme's membrane association to be proposed. Phylogenies for the two constituent subunits (i.e., the molybdopterin-containing and the Rieske subunit) of the heterodimeric enzyme and their respective homologs in DMSO-reductase, formate dehydrogenase, nitrate reductase, and the Rieske/cytb complexes were calculated from multiple sequence alignments. The obtained phylogenetic trees indicate an early origin of arsenite oxidase before the divergence of Archaea and Bacteria. Evolutionary implications of these phylogenies are discussed.

Amino Acid Sequence↗

Assessment of protein distance measures and tree-building methods for phylogenetic tree reconstruction.

Distance-based methods are popular for reconstructing evolutionary trees of protein sequences, mainly because of their speed and generality. A number of variants of the classical neighbor-joining (NJ) algorithm have been proposed, as well as a number of methods to estimate protein distances. We here present a large-scale assessment of performance in reconstructing the correct tree topology for the most popular algorithms. The programs BIONJ, FastME, Weighbor, and standard NJ were run using 12 distance estimators, producing 48 tree-building/distance estimation method combinations. These were evaluated on a test set based on real trees taken from 100 Pfam families. Each tree was used to generate multiple sequence alignments with the ROSE program using three evolutionary models. The accuracy of each method was analyzed as a function of both sequence divergence and location in the tree. We found that BIONJ produced the overall best results, although the average accuracy differed little between the tree-building methods (normally less than 1%). A noticeable trend was that FastME performed poorer than the rest on long branches. Weighbor was several orders of magnitude slower than the other programs. Larger differences were observed when using different distance estimators. Protein-adapted Jukes-Cantor and Kimura distance correction produced clearly poorer results than the other methods, even worse than uncorrected distances. We also assessed the recently developed Scoredist measure, which performed equally well as more complex methods.

Base Sequence↗

Automated phylogenetic detection of recombination using a genetic algorithm.

The evolution of homologous sequences affected by recombination or gene conversion cannot be adequately explained by a single phylogenetic tree. Many tree-based methods for sequence analysis, for example, those used for detecting sites evolving nonneutrally, have been shown to fail if such phylogenetic incongruity is ignored. However, it may be possible to propose several phylogenies that can correctly model the evolution of nonrecombinant fragments. We propose a model-based framework that uses a genetic algorithm to search a multiple-sequence alignment for putative recombination break points, quantifies the level of support for their locations, and identifies sequences or clades involved in putative recombination events. The software implementation can be run quickly and efficiently in a distributed computing environment, and various components of the methods can be chosen for computational expediency or statistical rigor. We evaluate the performance of the new method on simulated alignments and on an array of published benchmark data sets. Finally, we demonstrate that prescreening alignments with our method allows one to analyze recombinant sequences for positive selection.

Algorithms↗

Functional classification of amino acid decarboxylases from the alanine racemase structural family by phylogenetic studies.

Arginine decarboxylase (ADC) and ornithine decarboxylase (ODC) are involved in the biosynthesis of putrescine, which is the precursor of other polyamines in animals, plants, and bacteria. These pyridoxal-5'-phosphate-dependent decarboxylases belong to the alanine racemase (AR) structural family together with diaminopimelate decarboxylase (DapDC), which catalyzes the final step of lysine biosynthesis in bacteria. We have constructed a multiple-sequence alignment of decarboxylases in the AR structural family and, based on the alignment, inferred phylogenetic trees. The phylogenetic tree consists of 3 distinct clades formed by ADC, DapDC, and ODC that diverged from an ancestral decarboxylase. The ancestral decarboxylase probably was able to recognize several substrates, and in archaea and bacteria, ODC may have retained the ability to bind other amino acids. Previously, a paralogue of ODC has been proposed to account for ADC activity detected in mammalian cells. According to our results, this appears unlikely, emphasizing the need for more caution in functional assignment made using sequence data and illustrating the continuing value of phylogenetic analysis in clarifying relationships and putative functions.

Alanine Racemase↗

Molecular cloning of the cDNA for the catalytic subunit of human DNA polymerase delta.

The cDNA of human DNA polymerase delta was cloned. The cDNA had a length of 3.5 kb and encoded a protein of 1107 amino acid residues with a calculated molecular mass of 124 kDa. Northern blot analysis showed that the cDNA hybridized to a mRNA of 3.4 kb. Monoclonal and polyclonal antibodies to the C-terminal 20 residues specifically immunoblotted the human pol delta catalytic polypeptide. A multiple sequence alignment was constructed. This showed that human pol delta is closely related to yeast pol delta and the herpes virus DNA polymerases. The levels of pol delta message were found to be induced concomitantly with DNA pol delta activity and DNA synthesis in serum restimulated proliferating IMR90 cultured cells. The human pol delta gene was localized to chromosome 19 by Southern blotting of EcoRI digested DNA from a panel of rodent/human cell hybrids.

Amino Acid Sequence↗

Molecular evolution of SRP cycle components: functional implications.

Signal recognition particle (SRP) is a cytoplasmic ribonucleoprotein that targets a subset of nascent presecretory proteins to the endoplasmic reticulum membrane. We have considered the SRP cycle from the perspective of molecular evolution, using recently determined sequences of genes or cDNAs encoding homologs of SRP (7SL) RNA, the Srp54 protein (Srp54p), and the alpha subunit of the SRP receptor (SR alpha) from a broad spectrum of organisms, together with the remaining five polypeptides of mammalian SRP. Our analysis provides insight into the significance of structural variation in SRP RNA and identifies novel conserved motifs in protein components of this pathway. The lack of congruence between an established phylogenetic tree and size variation in 7SL homologs implies the occurrence of several independent events that eliminated more than half the sequence content of this RNA during bacterial evolution. The apparently non-essential structures are domain I, a tRNA-like element that is constant in archaea, varies in size among eucaryotes, and is generally missing in bacteria, and domain III, a tightly base-paired hairpin that is present in all eucaryotic and archeal SRP RNAs but is invariably absent in bacteria. Based on both structural and functional considerations, we propose that the conserved core of SRP consists minimally of the 54 kDa signal sequence-binding protein complexed with the loosely base-paired domain IV helix of SRP RNA, and is also likely to contain a homolog of the Srp68 protein. Comparative sequence analysis of the methionine-rich M domains from a diverse array of Srp54p homologs reveals an extended region of amino acid identity that resembles a recently identified RNA recognition motif. Multiple sequence alignment of the G domains of Srp54p and SR alpha homologs indicates that these two polypeptides exhibit significant similarity even outside the four GTPase consensus motifs, including a block of nine contiguous amino acids in a location analogous to the binding site of the guanine nucleotide dissociation stimulator (GDS) for E. coli EF-Tu. The conservation of this sequence, in combination with the results of earlier genetic and biochemical studies of the SRP cycle, leads us to hypothesize that a component of the Srp68/72p heterodimer serves as the GDS for both Srp54p and SR alpha. Using an iterative alignment procedure, we demonstrate similarity between Srp68p and sequence motifs conserved among GDS proteins for small Ras-related GTPases. The conservation of SRP cycle components in organisms from all three major branches of the phylogenetic tree suggests that this pathway for protein export is of ancient evolutionary origin.

Amino Acid Sequence↗

The Genome Sequence DataBase version 1.0 (GSDB): from low pass sequences to complete genomes.

The Genome Sequence DataBase (GSDB) has completed its conversion to an improved relational database. The new database, GSDB 1.0, is fully operational and publicly available. Data contributions, including both original sequence submissions and community annotation, are being accomplished through the use of a graphical client-server interface tool, the GSDB Annotator, and via GIO (GSDB Input/Output) files. Data retrieval services are being provided through a new Web Query Tool and direct SQL. All methods of data contribution and data retrieval fully support the new data types that have been incorporated into GSDB, including discontiguous sequences, multiple sequence alignments, and community annotation.

Animals↗

Comparative sequence analysis of ribonucleases HII, III, II PH and D.

Escherichia coli ribonucleases (RNases) HII, III, II, PH and D have been used to characterise new and known viral, bacterial, archaeal and eucaryotic sequences similar to these endo- (HII and III) and exoribonucleases (II, PH and D). Statistical models, hidden Markov models (HMMs), were created for the RNase HII, III, II and PH and D families as well as a double-stranded RNA binding domain present in RNase III. Results suggest that the RNase D family, which includes Werner syndrome protein and the 100 kDa antigenic component of the human polymyositis scleroderma (PMSCL) autoantigen, is a 3'-->5' exoribonuclease structurally and functionally related to the 3'-->5' exodeoxyribonuclease domain of DNA polymerases. Polynucleotide phosphorylases and the RNase PH family, which includes the 75 kDa PMSCL autoantigen, possess a common domain suggesting similar structures and mechanisms of action for these 3'-->5' phosphorolytic enzymes. Examination of HMM-generated multiple sequences alignments for each family suggest amino acids that may be important for their structure, substrate binding and/or catalysis.

Amino Acid Sequence↗

The proofreading domain of Escherichia coli DNA polymerase I and other DNA and/or RNA exonuclease domains.

Prior sequence analysis studies have suggested that bacterial ribonuclease (RNase) Ds comprise a complete domain that is found also in Homo sapiens polymyositis-scleroderma overlap syndrome 100 kDa autoantigen and Werner syndrome protein. This RNase D 3'-->5' exoribonuclease domain was predicted to have a structure and mechanism of action similar to the 3'-->5' exodeoxyibonuclease (proofreading) domain of DNA polymerases. Here, hidden Markov model (HMM) and phylogenetic studies have been used to identify and characterise other sequences that may possess this exonuclease domain. Results indicate that it is also present in the RNase T family; Borrelia burgdorferi P93 protein, an immunodominant antigen in Lyme disease; bacteriophage T4 dexA and Escherichia coli exonuclease I, processive 3'-->5' exodeoxyribonucleases that degrade single-stranded DNA; Bacillus subtilis dinG, a probable helicase involved in DNA repair and possibly replication, and peptide synthase 1; Saccharomyces cerevisiae Pab1p-dependent poly(A) nuclease PAN2 subunit, required for shortening mRNA poly(A) tails; Caenorhabditis elegans and Mus musculus CAF1, transcription factor CCR4-associated factor 1; Xenopus laevis XPMC2, prevention of mitotic catastrophe in fission yeast; Drosophila melanogaster egalitarian, oocyte specification and axis determination, and exuperantia, establishment of oocyte polarity; H.sapiens HEM45, expressed in tumour cell lines and uterus and regulated by oestrogen; and 31 open reading frames including one in Methanococcus jannaschii . Examination of a multiple sequence alignment and two three-dimensional structures of proofreading domains has allowed definition of the core sequence, structural and functional elements of this exonuclease domain.

Amino Acid Sequence↗

GPCRDB: an information system for G protein-coupled receptors.

The GPCRDB is a G protein-coupled receptor (GPCR) database system aimed at the collection and dissemination of GPCR related data. It holds sequences, mutant data and ligand binding constants as primary (experimental) data. Computationally derived data such as multiple sequence alignments, three dimensional models, phylogenetic trees and two dimensional visualization tools are added to enhance the database's usefulness. The GPCRDB is an EU sponsored project aimed at building a generic molecular class specific database capable of dealing with highly heterogeneous data. GPCRs were chosen as test molecules because of their enormous importance for medical sciences and due to the availability of so much highly heterogeneous data. The GPCRDB is available via the WWW at http://www.gpcr.org/7tm

Computer Communication Networks↗

Histone Sequence Database: new histone fold family members.

Searches of the major public protein databases with core and linker chicken and human histone sequences have resulted in the compilation of an annotated set of histone protein sequences. In addition, new database searches with two distinct motif search algorithms have identified several members of the histone fold family, including human DRAP1 and yeast CSE4. Database resources include information on conflicts between similar sequence entries in different source databases, multiple sequence alignments, links to the Entrez integrated information retrieval system, structures for histone and histone fold proteins, and the ability to visualize structural data through Cn3D. The database currently contains >1000 protein sequences, which are searchable by protein type, accession number, organism name, or any other free text appearing in the definition line of the entry. All sequences and alignments in this database are available through the World Wide Web at http://www.nhgri.nih. gov/DIR/GTB/HISTONES or http://www.ncbi.nlm.nih. gov/Baxevani/HISTONES

Amino Acid Sequence↗

Concordance analysis of microbial genomes.

The set of proteins which are conserved across families of microbes contain important targets of new anti-microbial agents. We have developed a simple and efficient computational tool which determines concordances of putative gene products that show sets of proteins conserved across one set of user specified genomes and not present in another set of user specified genomes. The thresholds and the homology scoring criterion are selectable to allow the user to decide the stringency of the homologies. The system uses a relational database to store protein coding regions from different genomes, and to store the results of a complete comparison of all sequences against all sequences using the FASTA program. Using Web technology, the display of all the related proteins for a given sequence and calculation of multiple sequence alignments (using CLUSTALW) can be performed with the click of a button. The current database holds 97 365 sequences from 19 complete or partial genomes and 8798905 FASTA comparison results. A example concordance is presented which demonstrates that the target of the quinolone antibiotics could have been identified using this tool.

Bacteria↗

The Ligand Gated Ion Channel Database.

The ligand gated ion channels (LGICs) are ionotropic receptors to neurotransmitters. Their physiological effect is carried out by the opening of an ionic channel upon binding of a particular neurotransmitter. These LGICs constitute superfamilies of receptors formed by homologous subunits. A database has been developed to handle the growing wealth of cloned subunits. This database contains nucleic acid sequences, protein sequences, as well as multiple sequence alignments and phylogenetic studies. This database is accessible via the worldwide web (http://www.pasteur.fr/units/neubiomol/LGIC.h tml), where it is continuously updated. A downloadable version is also available [currently v0.1 (98.06)].

Databases, Factual↗

Identification and characterization of a DNA primase from the hyperthermophilic archaeon Methanococcus jannaschii.

We report the identification and characterisation of a DNA primase from the thermophilic methanogenic archaeon Methanococcus jannaschii (Mjpri). The analysis of the complete genome sequence of this organism has identified an open reading frame coding for a protein with sequence similarity to the small subunit of the eukaryotic DNA primase (the p50 subunit of the polymerase alpha-primase complex). This protein has been overexpressed in Escherichia coli and purified to near homogeneity. Recombinant Mjpri is able to synthesise oligoribonucleotides on various pyrimidine single-stranded DNA templates [poly(dT) and poly(dC)]. This activity requires divalent cations such Mg(2+), Mn(2+)or Zn(2+), and is additionally stimulated by the monovalent cation K(+). A multiple sequence alignment has revealed that most of the regions that are conserved in eukaryotic p50 subunits are also present in the archaeal primases, including the conserved negatively charged residues, which have been shown to be essential for catalysis in the mouse primase. Of the four cysteine residues that have been postulated to make up a putative Zn-binding motif, two are not present in the archaeal homologue. This is the first report on the biochemical characterisation of an archaeal DNA primase.

Amino Acid Sequence↗

Molecular evolution of DNA-(cytosine-N4) methyltransferases: evidence for their polyphyletic origin.

DNA N4-cytosine methyltransferases (N4mC MTases) are a family of S-adenosyl-L-methionine (AdoMet)-dependent MTases. Members of this family were previously found to share nine conserved sequence motifs, but the evolutionary basis of these similarities has never been studied in detail. We performed phylogenetic analysis of 37 known and potential new family members from the multiple sequence alignment using distance matrix, parsimony and maximum likelihood approaches to infer the evolutionary relationship among the N4mC MTases and classify them into groups of orthologs. All the treeing algorithms employed as well as results of exhaustive sequence database searching support a scenario, in which the majority of N4mC MTases, except for M. Bal I and M. Bam HI, arose by divergence from a common ancestor. Interestingly, MTases M. Bal I and M. Bam HI apparently originated from N6-adenine MTases and represent the most recent addendum to the N4mC MTase family. In addition to the previously reported nine sequence motifs, two more conserved sequence patches were detected. Phylogenetic analysis also provided the evidence for massive horizontal transfer of MTase genes, presumably with the whole restriction-modification systems, between Bacteria and Archaea.

Amino Acid Sequence↗

The Pfam protein families database.

Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the WWW in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgr.ki.se/Pfam/ and in the US at http://pfam.wustl.edu/. The latest version (4.3) of Pfam contains 1815 families. These Pfam families match 63% of proteins in SWISS-PROT 37 and TrEMBL 9. For complete genomes Pfam currently matches up to half of the proteins. Genomic DNA can be directly searched against the Pfam library using the Wise2 package.

Databases, Factual↗