Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

ClustalW-MPI: ClustalW analysis using distributed and parallel computing.

ClustalW is a tool for aligning multiple protein or nucleotide sequences. The alignment is achieved via three steps: pairwise alignment, guide-tree generation and progressive alignment. ClustalW-MPI is a distributed and parallel implementation of ClustalW. All three steps have been parallelized to reduce the execution time. The software uses a message-passing library called MPI (Message Passing Interface) and runs on distributed workstation clusters as well as on traditional parallel computers.

Amino Acid Sequence↗

Fast algorithms for large-scale genome alignment and comparison.

We describe a suffix-tree algorithm that can align the entire genome sequences of eukaryotic and prokaryotic organisms with minimal use of computer time and memory. The new system, MUMmer 2, runs three times faster while using one-third as much memory as the original MUMmer system. It has been used successfully to align the entire human and mouse genomes to each other, and to align numerous smaller eukaryotic and prokaryotic genomes. A new module permits the alignment of multiple DNA sequence fragments, which has proven valuable in the comparison of incomplete genome sequences. We also describe a method to align more distantly related genomes by detecting protein sequence homology. This extension to MUMmer aligns two genomes after translating the sequence in all six reading frames, extracts all matching protein sequences and then clusters together matches. This method has been applied to both incomplete and complete genome sequences in order to detect regions of conserved synteny, in which multiple proteins from one organism are found in the same order and orientation in another. The system code is being made freely available by the authors.

Algorithms↗

Automated de novo identification of repeat sequence families in sequenced genomes.

Repetitive sequences make up a major part of eukaryotic genomes. We have developed an approach for the de novo identification and classification of repeat sequence families that is based on extensions to the usual approach of single linkage clustering of local pairwise alignments between genomic sequences. Our extensions use multiple alignment information to define the boundaries of individual copies of the repeats and to distinguish homologous but distinct repeat element families. When tested on the human genome, our approach was able to properly identify and group known transposable elements. The program, should be useful for first-pass automatic classification of repeats in newly sequenced genomes.

Algorithms↗

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗

Combining phylogenetic data with co-regulated genes to identify regulatory motifs.

MOTIVATION: Discovery of regulatory motifs in unaligned DNA sequences remains a fundamental problem in computational biology. Two categories of algorithms have been developed to identify common motifs from a set of DNA sequences. The first can be called a 'multiple genes, single species' approach. It proposes that a degenerate motif is embedded in some or all of the otherwise unrelated input sequences and tries to describe a consensus motif and identify its occurrences. It is often used for co-regulated genes identified through experimental approaches. The second approach can be called 'single gene, multiple species'. It requires orthologous input sequences and tries to identify unusually well conserved regions by phylogenetic footprinting. Both approaches perform well, but each has some limitations. It is tempting to combine the knowledge of co-regulation among different genes and conservation among orthologous genes to improve our ability to identify motifs. RESULTS: Based on the Consensus algorithm previously established by our group, we introduce a new algorithm called PhyloCon (Phylogenetic Consensus) that takes into account both conservation among orthologous genes and co-regulation of genes within a species. This algorithm first aligns conserved regions of orthologous sequences into multiple sequence alignments, or profiles, then compares profiles representing non-orthologous sequences. Motifs emerge as common regions in these profiles. Here we present a novel statistic to compare profiles of DNA sequences and a greedy approach to search for common subprofiles. We demonstrate that PhyloCon performs well on both synthetic and biological data. AVAILABILITY: Software available upon request from the authors. http://ural.wustl.edu/softwares.html

Algorithms↗

Saccharomyces Genome Database (SGD) provides tools to identify and analyze sequences from Saccharomyces cerevisiae and related sequences from other organisms.

The Saccharomyces Genome Database (SGD; http://www.yeastgenome.org/), a scientific database of the molecular biology and genetics of the yeast Saccharomyces cerevisiae, has recently developed several new resources that allow the comparison and integration of information on a genome-wide scale, enabling the user not only to find detailed information about individual genes, but also to make connections across groups of genes with common features and across different species. The Fungal Alignment Viewer displays alignments of sequences from multiple fungal genomes, while the Sequence Similarity Query tool displays PSI-BLAST alignments of each S.cerevisiae protein with similar proteins from any species whose sequences are contained in the non-redundant (nr) protein data set at NCBI. The Yeast Biochemical Pathways tool integrates groups of genes by their common roles in metabolism and displays the metabolic pathways in a graphical form. Finally, the Find Chromosomal Features search interface provides a versatile tool for querying multiple types of information in SGD.

Amino Acid Sequence↗

MulBlast 1.0: a multiple alignment of BLAST output to boost protein sequence similarity analysis.

The protein sequence similarity search has become a major tool for biologists. Various efficient and rapid programs and comparison matrices have been designed and refined in order to perform the scanning task (BLAST, FASTA, Automat, etc.). However, the final step of the search, the analysis of the results, is still tedious and time consuming. In order to optimize true-positive hit screening, we have developed a program which makes a multiple alignment from the BLAST search output. Conserved sequence segments are pointed out. It makes the recognition of already known as well as new sequence patterns easier. It allows at a glance a rapid identification of significant similarities, protein family signature and new sequence motifs. This alignment is written in a compatible format for the GCG programs LineUp and ProfileMake.

Algorithms↗

A space-efficient algorithm for aligning large genomic sequences.

SUMMARY: In the segment-by-segment approach to sequence alignment, pairwise and multiple alignments are generated by comparing gap-free segments of the sequences under study. This method is particularly efficient in detecting local homologies, and it has been used to identify functional regions in large genomic sequences. Herein, an algorithm is outlined that calculates optimal pairwise segment-by-segment alignments in essentially linear space. AVAILABILTIY: The program is available at the Bielefeld Bioinformatics Server (BiBiServ) at http://bibiserv.techfak. uni-bielefeld.de/dialign/

Algorithms↗

Assessment of cry1 gene contents of Bacillus thuringiensis strains by use of DNA microarrays.

A single Bacillus thuringiensis strain can harbor numerous different insecticidal crystal protein (cry) genes from 46 known classes or primary ranks. The cry1 primary rank is the best known and contains the highest number of cry genes which currently totals over 130. We have designed an oligonucleotide-based DNA microarray (cryArray) to test the feasibility of using microarrays to identify the cry gene content of B. thuringiensis strains. Specific 50-mer oligonucleotide probes representing the cry1 primary and tertiary ranks were designed based on multiple cry gene sequence alignments. To minimize false-positive results, a consentaneous approach was adopted in which multiple probes against a specific gene must unanimously produce positive hybridization signals to confirm the presence of a particular gene. In order to validate the cryArray, several well-characterized B. thuringiensis strains including isolates from a Mexican strain collection were tested. With few exceptions, our probes performed in agreement with known or PCR-validated results. In one case, hybridization of primary- but not tertiary-ranked cry1I probes indicated the presence of a novel cry1I gene. Amplification and partial sequencing of the cry1I gene in strains IB360 and IB429 revealed the presence of a cry1Ia gene variant. Since a single microarray hybridization can replace hundreds of individual PCRs, DNA microarrays should become an excellent tool for the fast screening of new B. thuringiensis isolates presenting interesting insecticidal activity.

Bacillus thuringiensis↗

Stress-dependent expression of a polymorphic, charged antigen in the protozoan parasite Entamoeba histolytica.

We have identified a novel stress inducible gene, Ehssp1 in Entamoeba histolytica, the causative agent of amebiasis. Ehssp1 belongs to a polymorphic, multigene family and is present on multiple chromosomes. No homologue of this gene was found in the NCBI database. Sequence alignment of the multiple copies, and genomic PCR data restricted the polymorphism to the central region of the gene. This region contains a polypurine stretch that encodes a domain rich in acidic and basic amino acids. Under normal culture conditions only one copy of this multigene family is expressed, as observed by Northern blot and RT-PCR analysis. The size of this copy of the gene is 1,077 nucleotides, encoding a protein of 359 amino acids. The polymorphic domain in this copy is 64 nucleotides long. However, on exposure of cells to stress conditions such as heat shock or oxidative stress, multiple polymorphic copies of the gene are expressed, suggesting a possible role of this gene in adaptation of cells to stress conditions. The gene copy expressed under normal conditions, and the expression profile of cells under heat stress was identical in two different strains of E. histolytica tested. Interestingly, the extent of polymorphism in this gene was very less in E. dispar, a nonpathogenic sibling species of E. histolytica. Ehssp1 was found to be antigenic in invasive amebiasis patients.

Amino Acid Sequence↗

Multiple structural alignment for distantly related all beta structures using TOPS pattern discovery and simulated annealing.

Topsalign is a method that will structurally align diverse protein structures, for example, structural alignment of protein superfolds. All proteins within a superfold share the same fold but often have very low sequence identity and different biological and biochemical functions. There is often significant structural diversity around the common scaffold of secondary structure elements of the fold. Topsalign uses topological descriptions of proteins. A pattern discovery algorithm identifies equivalent secondary structure elements between a set of proteins and these are used to produce an initial multiple structure alignment. Simulated annealing is used to optimize the alignment. The output of Topsalign is a multiple structure-based sequence alignment and a 3D superposition of the structures. This method has been tested on three superfolds: the beta jelly roll, TIM (alpha/beta) barrel and the OB fold. Topsalign outperforms established methods on very diverse structures. Despite the pattern discovery working only on beta strand secondary structure elements, Topsalign is shown to align TIM (alpha/beta) barrel superfamilies, which contain both alpha helices and beta strands.

Protein Structure, Secondary↗

Alignment of RNA base pairing probability matrices.

MOTIVATION: Many classes of functional RNA molecules are characterized by highly conserved secondary structures but little detectable sequence similarity. Reliable multiple alignments can therefore be constructed only when the shared structural features are taken into account. Since multiple alignments are used as input for many subsequent methods of data analysis, structure-based alignments are an indispensable necessity in RNA bioinformatics. RESULTS: We present here a method to compute pairwise and progressive multiple alignments from the direct comparison of base pairing probability matrices. Instead of attempting to solve the folding and the alignment problem simultaneously as in the classical Sankoff's algorithm, we use McCaskill's approach to compute base pairing probability matrices which effectively incorporate the information on the energetics of each sequences. A novel, simplified variant of Sankoff's algorithms can then be employed to extract the maximum-weight common secondary structure and an associated alignment. AVAILABILITY: The programs pmcomp and pmmulti described in this contribution are implemented in Perl and can be downloaded together with the example datasets from http://www.tbi.univie.ac.at/RNA/PMcomp/. A web server is available at http://rna.tbi.univie.ac.at/cgi-bin/pmcgi.pl

Algorithms↗

Malate dehydrogenase: a model for structure, evolution, and catalysis.

Malate dehydrogenases are widely distributed and alignment of the amino acid sequences show that the enzyme has diverged into 2 main phylogenetic groups. Multiple amino acid sequence alignments of malate dehydrogenases also show that there is a low degree of primary structural similarity, apart from in several positions crucial for nucleotide binding, catalysis, and the subunit interface. The 3-dimensional structures of several malate dehydrogenases are similar, despite their low amino acid sequence identity. The coenzyme specificity of malate dehydrogenase may be modulated by substitution of a single residue, as can the substrate specificity. The mechanism of catalysis of malate dehydrogenase is similar to that of lactate dehydrogenase, an enzyme with which it shares a similar 3-dimensional structure. Substitution of a single amino acid residue of a lactate dehydrogenase changes the enzyme specificity to that of a malate dehydrogenase, but a similar substitution in a malate dehydrogenase resulted in relaxation of the high degree of specificity for oxaloacetate. Knowledge of the 3-dimensional structures of malate and lactate dehydrogenases allows the redesign of enzymes by rational rather than random mutation and may have important commercial implications.

Amino Acid Sequence↗

Vertebrate gene finding from multiple-species alignments using a two-level strategy.

BACKGROUND: One way in which the accuracy of gene structure prediction in vertebrate DNA sequences can be improved is by analyzing alignments with multiple related species, since functional regions of genes tend to be more conserved. RESULTS: We describe DOGFISH, a vertebrate gene finder consisting of a cleanly separated site classifier and structure predictor. The classifier scores potential splice sites and other features, using sequence alignments between multiple vertebrate species, while the structure predictor hypothesizes coding transcripts by combining these scores using a simple model of gene structure. This also identifies and assigns confidence scores to possible additional exons. Performance is assessed on the ENCODE regions. We predict transcripts and exons across the whole human genome, and identify over 10,000 high confidence new coding exons not in the Ensembl gene set. CONCLUSION: We present a practical multiple species gene prediction method. Accuracy improves as additional species, up to at least eight, are introduced. The novel predictions of the whole-genome scan should support efficient experimental verification.

Animals↗

Evolutionary HMMs: a Bayesian approach to multiple alignment.

MOTIVATION: We review proposed syntheses of probabilistic sequence alignment, profiling and phylogeny. We develop a multiple alignment algorithm for Bayesian inference in the links model proposed by Thorne et al. (1991, J. Mol. Evol., 33, 114-124). The algorithm, described in detail in Section 3, samples from and/or maximizes the posterior distribution over multiple alignments for any number of DNA or protein sequences, conditioned on a phylogenetic tree. The individual sampling and maximization steps of the algorithm require no more computational resources than pairwise alignment. METHODS: We present a software implementation (Handel) of our algorithm and report test results on (i) simulated data sets and (ii) the structurally informed protein alignments of BAliBASE (Thompson et al., 1999, Nucleic Acids Res., 27, 2682-2690). RESULTS: We find that the mean sum-of-pairs score (a measure of residue-pair correspondence) for the BAliBASE alignments is only 13% lower for Handelthan for CLUSTALW(Thompson et al., 1994, Nucleic Acids Res., 22, 4673-4680), despite the relative simplicity of the links model (CLUSTALW uses affine gap scores and increased penalties for indels in hydrophobic regions). With reference to these benchmarks, we discuss potential improvements to the links model and implications for Bayesian multiple alignment and phylogenetic profiling. AVAILABILITY: The source code to Handelis freely distributed on the Internet at http://www.biowiki.org/Handel under the terms of the GNU Public License (GPL, 2000, http://www.fsf.org./copyleft/gpl.html).

Algorithms↗

Identifying multiple alignment regions satisfying simple formulas and patterns.

MOTIVATION: When studying multiple alignments of genomic sequences one frequently aims to locate and count regions which satisfy a set of constraints. These regions may be putatively functional, but researchers may also be interested in quantifying the frequency of occurrences of certain patterns. RESULTS: We have developed a program that applies simple formulas and pattern specifications to multiple alignments, reporting the positions and counts of conforming regions. As an example, we have navigated a 15-species alignment of the CAV2-CAV1 region and outlined some findings regarding PPARgamma binding sites. AVAILABILITY: Our software and the accompanying documentation can be obtained at no charge by contacting the authors. It can also be accessed at http://ranger.uta.edu/~nick/compgen

Algorithms↗

Computing TaqMan probes for multiplex PCR detection of E. coli O157 serotypes in water.

Diarrheagenic E. coli strains contribute to water related diseases in urban and rural environment in developing and developed world. E. coli pathotype and pathogenicity varies due to complex multifactorial mechanism involving a large number of virulence factors. Rapid assessment of the virulence pattern of E. coli isolates is possible by Real-Time PCR probes like TaqMan. For designing TaqMan probes and primers for multiplex PCR selected E. coli gene sequences: stx1, stx2, hlyA, chuA, eae, lacZ, lamB and fimA were retrieved from NCBI's GenBank database. The alignment of the multiple sequences and analysis of conserved sequences was carried out using ClustalW and BLAST programs. The primers and Taqmen probes were designed using Beacon Designer software version 2.1 for two multiplexed PCR assays. In silico PCR simulation of these assays showed PCR products for stx2 (248bp) stx1 (102 bp), lacZ (228bp) and lamB (86 bp) in multiplex #1 and eae (200bp), chuA (147 bp), hlyA (141bp) and fimA (79 bp) in multiplex #2, respectively. These multiplexed PCR amplification products and probes can be used to identify and confirm presence of O157:H7/ H7-, O157:H43/45 and O26:H-/H11 serotypes. In conclusion, multiplex Real-Time Polymerase Chain Reaction oligomers and TaqMan probes designed and validated in silico will be helpful in management of water quality and outbreaks, by improving specificity and minimizing time needed for in vitro verification work.

Base Sequence↗

Sequence determinants for the reaction specificity of murine (12R)-lipoxygenase: targeted substrate modification and site-directed mutagenesis.

Mammalian lipoxygenases (LOXs) are categorized with respect to their positional specificity of arachidonic acid oxygenation. Site-directed mutagenesis identified sequence determinants for the positional specificity of these enzymes, and a critical amino acid for the stereoselectivity was recently discovered. To search for sequence determinants of murine (12R)-LOX, we carried out multiple amino acid sequence alignments and found that Phe(390), Gly(441), Ala(455), and Val(631) align with previously identified positional determinants of S-LOX isoforms. Multiple site-directed mutagenesis studies on Phe(390) and Ala(455) did not induce specific alterations in the reaction specificity, but yielded enzyme species with reduced specific activities and stereo random product patterns. Mutation of Gly(441) to Ala, which caused drastic alterations in the reaction specificity of other LOX isoforms, failed to induce major alterations in the positional specificity of mouse (12R)-LOX, but markedly modified the enantioselectivity of the enzyme. When Val(631), which aligns with the positional determinant Ile(593) of rabbit 15-LOX, was mutated to a less space-filling residue (Ala or Gly), we obtained an enzyme species with augmented catalytic activity and specifically altered reaction characteristics (major formation of chiral (11R)-hydroxyeicosatetraenoic acid methyl ester). The importance of Val(631) for the stereo control of murine (12R)-LOX was confirmed with other substrates such as methyl linoleate and 20-hydroxyeicosatetraenoic acid methyl ester. These data identify Val(631) as the major sequence determinant for the specificity of murine (12R)-LOX. Furthermore, we conclude that substrate fatty acids may adopt different catalytically productive arrangements at the active site of murine (12R)-LOX and that each of these arrangements may lead to the formation of chiral oxygenation products.

12-Hydroxy-5,8,10,14-eicosatetraenoic Acid↗