Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

Immune cascade of Spodoptera litura: cloning, expression, and characterization of inducible prophenol oxidase.

Haemolymph associated phenol oxidase is a critical component of invertebrate immune reaction and cuticle sclerotization. Phenol oxidase catalyses the conversion of mono-phenols to diphenols and quinones which finally leads to melanin formation. We have cloned the c-DNA encoding phenol oxidase from the haemocytes of Spodoptera litura and expressed it in Escherichia coli. The encoding gene is 2452bp with an open reading frame of 2091 bp translating into a 697 amino acid protein. Multiple alignment analysis of the predicted protein sequence shows close homology to other lepidopeteran PPOII type genes. The transcription of the gene is induced upon microbial challenge of 6th instar larvae with E. coli and is unresponsive to injury. Cloning of the ORF of SLPPO in-frame in the E. coli expression vector pQE30 resulted in its expression. Enzymatic analysis of the recombinant protein reveals that the recombinant protein is catalytically active on 4-methyl pyrocatechol upon activation by cetyl pyridinium chloride.

Amino Acid Sequence↗

Improving profile HMM discrimination by adapting transition probabilities.

Profile hidden Markov models (HMMs) are used to model protein families and for detecting evolutionary relationships between proteins. Such a profile HMM is typically constructed from a multiple alignment of a set of related sequences. Transition probability parameters in an HMM are used to model insertions and deletions in the alignment. We show here that taking into account unrelated sequences when estimating the transition probability parameters helps to construct more discriminative models for the global/local alignment mode. After normal HMM training, a simple heuristic is employed that adjusts the transition probabilities between match and delete states according to observed transitions in the training set relative to the unrelated (noise) set. The method is called adaptive transition probabilities (ATP) and is based on the HMMER package implementation. It was benchmarked in two remote homology tests based on the Pfam and the SCOP classifications. Compared to the HMMER default procedure, the rate of misclassification was reduced significantly in both tests and across all levels of error rate.

Algorithms↗

Characterization of a recombinant immunodiagnostic antigen (NIE) from Strongyloides stercoralis L3-stage larvae.

Due to the process of internal autoinfection, even chronic asymptomatic infections with Strongyloides stercoralis have the potential to become severe disseminated disease with fatal outcome. Intermittent and scanty larval excretion makes parasitologic diagnosis difficult. Serodiagnosis is helpful, but antigen preparation from infective larvae requires access to patients or immunosuppressed experimental animals. For these reasons, attention has turned to recombinant antigens for immunodiagnosis. A 31-kDa candidate antigen (NIE) derived from an L3 cDNA library is described in this report. Multiple alignment of the deduced amino acid sequence of NIE showed approximately 12-18% identity with various other organisms, including 17.9% of Asp1 of Ancylostoma caninum, 12.6% of Hemonchus contortus, and 17.6% of insect venom allergen 5 of yellow jacket. By ELISA, antibodies to the purified recombinant NIE antigen were demonstrated in 87.5% of 48 sera from strongyloides-infected patients and in only 6.5% of sera from presumed normal controls. Immunoreactivity of purified NIE antigen with parasite-specific IgE from sera of strongyloides-infected patients indicated its potential use as an immediate sensitivity skin test antigen. This application of the NIE antigen was supported by its capacity to trigger release of histamine upon in vitro exposure to blood from strongyloides-infected patients and its failure to produce histamine release from blood of normal controls.

Amino Acid Sequence↗

Involvement of Gln937 of Streptococcus downei GTF-I glucansucrase in transition-state stabilization.

Multiple alignment of deduced amino-acid sequences of glucansucrases (glucosyltransferases and dextransucrases) from oral streptococci and Leuconostoc mesenteroides has shown them to share a well-conserved catalytic domain. A portion of this domain displays homology to members of the alpha-amylase family (glycoside hydrolase family 13), which all have a (beta/alpha)8 barrel structure. In the glucansucrases, however, the alpha-helix and beta-strand elements are circularly permuted with respect to the order in family 13. Previous work has shown that amino-acid residues contributing to the active site of glucansucrases are situated in structural elements that align with those of family 13. In alpha-amylase and cyclodextrin glucanotransferase, a histidine residue has been identified that acts to stabilize the transition state, and a histidine is conserved at the corresponding position in all other members of family 13. In all the glucansucrases, however, the aligned position is occupied by glutamine. Mutants of glucosyltransferase I were constructed in which this glutamine, Gln937, was changed to histidine, glutamic acid, aspartic acid, asparagine or alanine. The effects on specific activity, ability to form glucan and ability to transfer glucose to a maltose acceptor were examined. Only histidine could substitute for glutamine and maintain Michaelis-Menten kinetics, albeit at a greatly reduced kcat, showing that Gln937 plays a functionally equivalent role to the histidine in family 13. This provides additional evidence in support of the proposed alignment of the (beta/alpha)8 barrel structures. Mutation at position 937 altered the acceptor reaction with maltose, and resulted in the synthesis of novel gluco-oligosaccharides in which alpha1,3-linked glucosyl units are joined sequentially to maltose.

Enzyme Stability↗

Changing a single amino acid residue switches processive and non-processive behavior of Aspergillus niger endopolygalacturonase I and II.

Processivity, also known as multiple attack on a single chain, is a feature commonly encountered only in enzymes in which the substrate binds in a tunnel. However, of the seven Aspergillus niger endopolygalacturonases, which have an open substrate binding cleft, four enzymes show processive behavior, whereas the other endopolygalacturonases are randomly acting enzymes. In a previous study (Benen, J.A.E., Kester, H.C.M., and Visser, J. (1999) Eur. J. Biochem. 259, 577-585) we proposed that the high affinity for the substrate of subsite -5 of processive endopolygalacturonase I constitutes the origin of the multiple attack behavior. Based on primary sequence alignments of A. niger endopolygalacturonases and three-dimensional structure analysis of endopolygalacturonase II, an arginine residue was identified in the processive enzymes at a position commensurate with subsite -5, whereas a serine residue was present at this position in the non-processive enzymes. In endopolygalacturonase I mutation R95S was introduced, and in endopolygalacturonase II mutation S91R was introduced. Product progression analysis on polymer substrate and bond cleavage frequency studies using oligogalacturonides of defined chain length for the mutant enzymes revealed that processive/non-processive behavior is indeed interchangeable by one single amino acid substitution at subsite -5, Arg-->Ser or Ser-->Arg.

Amino Acid Sequence↗

Detection of BCR-ABL mutations and resistance to imatinib mesylate.

The major mechanism of imatinib resistance for patients with chronic myeloid leukemia (CML) is clonal expansion of leukemic cells with mutations in the Bcr-Abl fusion tyrosine kinase that reduce the capacity of imatinib to inhibit kinase activity. The early detection of such mutations may allow timely treatment intervention to prevent or overcome resistance. Direct sequencing of the BCR-ABL kinase domain is relatively rapid and allows detection of emerging mutations at a sensitivity of approx 20%. Mutations have been detected over a range of 242 amino acids, which spans the entire kinase domain. For optimal sensitivity, the kinase domain of the abnormal gene should be isolated by reverse-transcription (RT) polymerase chain reaction (PCR) amplification using primers that hybridize to the BCR and ABL genes. The quality of the RNA is assessed by real-time quantitative PCR prior to analysis, and BCR-ABL levels are determined. Only RNA of adequate quality is used to ensure accurate and reproducible mutation analysis. Depending on the level of BCR-ABL transcripts, a one- or two-step PCR is required to amplify the kinase domain. Direct sequencing with dye terminator chemistry is performed using PCR-purified products. The sequence is compared to an ABL kinase domain reference sequence using sequencing analysis software, which aligns the sequences and highlights single or multiple mutations.

Amino Acid Substitution↗

The deferred path heuristic for the generalized tree alignment problem.

Many multiple alignment methods implicitly or explicitly try to minimize the amount of biological change implied by an alignment. At the level of sequences, biological change is measured along a phylogenetic tree, a structure frequently being predicted only after the multiple alignment instead of together with it. The Generalized Tree Alignment problem addresses both questions simultaneously. It can formally be viewed as a Steiner tree problem in sequence space and our approach merges a path heuristic for the construction of a Steiner tree with a clustering method as usually applied only to distance data. This combination is achieved using sequence graphs, a data structure for efficient representation of similar sequences. Although somewhat slower in practice than an earlier method by Hein (1989) the current approach achieves significantly better results in terms of the underlying scoring function. Furthermore, a variant of the algorithm is introduced that maintains a guaranteed error bound of (2 - 2/n) for n sequences.

Algorithms↗

Improved genetic algorithm-based protein structure comparisons: pairwise and multiple superpositions.

Three major improvements to a previously described method for automatic protein structure comparison are described. First, a limit to translations for the rigid-body superposition is now assigned according to the dimensions of the structures being compared. Second, examination of the effect of the gap penalty on the derivation of a sequence alignment corresponding to a given structure superposition has led to a method to evaluate alternative structure-based sequence alignments. Third, the pairwise procedure has been generalized to multiple structure alignment. This implementation of rigid-body superposition can recognize well documented distant relationships which hitherto have required consideration of additional features and properties as well as those relationships between proteins of different sizes. A much larger common scaffold or framework between six globins can be extracted than that obtained using a standard algorithm for multiple structure superposition.

Algorithms↗

Finding the most significant common sequence and structure motifs in a set of RNA sequences.

We present a computational scheme to locally align a collection of RNA sequences using sequence and structure constraints. In addition, the method searches for the resulting alignments with the most significant common motifs, among all possible collections. The first part utilizes a simplified version of the Sankoff algorithm for simultaneous folding and alignment of RNA sequences, but maintains tractability by constructing multi-sequence alignments from pairwise comparisons. The algorithm finds the multiple alignments using a greedy approach and has similarities to both CLUSTAL and CONSENSUS, but the core algorithm assures that the pairwise alignments are optimized for both sequence and structure conservation. The choice of scoring system and the method of progressively constructing the final solution are important considerations that are discussed. Example solutions, and comparisons with other approaches, are provided. The solutions include finding consensus structures identical to published ones.

Algorithms↗

BLAST Filter and GraphAlign: rule-based formation and analysis of sets of related DNA and protein sequences.

BLAST Filter and GraphAlign are web-based tools that offer novel methods for building and analyzing sets of related (i.e. similar) DNA and protein sequences. They can be used separately or together. BLAST Filter generates related sequence sets in an automated, objective and reproducible way based on an input query sequence. Sequences matched by BLAST are filtered through a set of 15 user-configurable rules based on full-length query/subject comparisons, high-scoring segment pair statistics and the level of redundancy in the sequence set. Such sets can be used for multiple alignments, profile hidden Markov models and other bioinformatics applications, including GraphAlign, which provides several novel methods for analyzing global query/subject alignments along with graphical representations of sequence similarities. These services are available at the following URLs: http://darwin.nmsu.edu/cgi-bin/blast_filter.cgi and http://darwin.nmsu.edu/cgi-bin/graph_align.cgi.

Computational Biology↗

The phylogenetic analysis of variable-length sequence data: elongation factor-1alpha introns in European populations of the parasitoid wasp genus Pauesia (Hymenoptera: Braconidae: Aphidiinae).

Elongation factor-1alpha (EF-1alpha) is a highly conserved nuclear coding gene that can be used to investigate recent divergences due to the presence of rapidly evolving introns. However, a universal feature of intron sequences is that even closely related species exhibit insertion and deletion events, which cause variation in the lengths of the sequences. Indels are frequently rich in evolutionary information, but most investigators ignore sites that fall within these variable regions, largely because the analytical tools and theory are not well developed. We examined this problem in the taxonomically problematic parasitoid wasp genus Pauesia (Hymenoptera: Braconidae: Aphidiinae) using congruence as a criterion for assessing a range of methods for aligning such variable-length EF-1alpha intron sequences. These methods included distance- and parsimony-based multiple-alignment programs (CLUSTAL W and MALIGN), direct optimization (POY), and two "by eye" alignment strategies. Furthermore, with one method (CLUSTAL W) we explored in detail the robustness of results to changes in the gap cost parameters. Phenetic-based alignments ("by eye" and CLUSTAL W) appeared, under our criterion, to perform as well as more readily defensible, but computationally more demanding, methods. In general, all of our alignment and tree-building strategies recovered the same basic topological structure, which means that an underlying phylogenetic signal remained regardless of the strategy chosen. However, several relationships between clades were sensitive both to alignment and to tree-building protocol. Further alignments, considering only sequences belonging to the same group, allowed us to infer a range of phylogenetic relationships that were highly robust to tree-building protocol. By comparing these topologies with those obtained by varying the CLUSTAL parameters, we generated the distribution area of congruence and taxonomic compatibility. Finally, we present the first robust estimate of the European Pauesia phylogeny by using two EF-1alpha introns and 38 taxa (plus 3 outgroups). This estimate conflicts markedly with the traditional subgeneric classification. We recommend that this classification be abandoned, and we propose a series of monophyletic species groups.

Animals↗

Multiple RNA structure alignment.

Ribonucleic Acid (RNA) structures can be viewed as a special kind of strings where characters in a string can bond with each other. The question of aligning two RNA structures has been studied for a while, and there are several successful algorithms that are based upon different models. In this paper, by adopting the model introduced in Wang and Zhang,(19) we propose two algorithms to attack the question of aligning multiple RNA structures. Our methods are to reduce the multiple RNA structure alignment problem to the problem of aligning two RNA structure alignments. Meanwhile, we will show that the framework of sequence center star alignment algorithm can be applied to the problem of multiple RNA structure alignment, and if the triangle inequality is met in the scoring matrix, the approximation ratio of the algorithm remains to be 2-2(over)n, where n is the total number of structures.

Algorithms↗

ProMSED: protein multiple sequence editor for Windows 3.11/95.

MOTIVATION: Most protein sequence alignment algorithms give similar results on closely related proteins, while manual intervention may be needed for distantly related molecules. To correct the alignment, it is often necessary to repeat calculations on selected parts of the alignments and edit the alignment manually. Software implementing such interactive alignment procedures is of significance. RESULTS: This paper presents a new MS Windows application called ProMSED for both automatic and manual protein sequence alignment. The program reads main sequence formats and has a user-friendly interface. ProMSED performs automatic (ClustalV algorithm) alignments, alignment visualization and editing, and it allows sequences to be aligned interactively leaving previously aligned regions unchanged. Manual alignment and sequence analysis are facilitated by colouring schemes reflecting amino acid similarity of mutational and physicochemical properties. The interactive alignment of a diverged set of reverse transcriptases has located four out of six known conserved motifs. AVAILABILITY: ProMSED is available on request from the authors. DEMO is available from ftp://ftp.ebi.ac.uk/pub/ software/dos/promsed/ or ftp://iubio.bio.indiana.edu/molbio/ ibmpc/.

Algorithms↗

Local multiple alignment by consensus matrix.

A new algorithm for aligning several sequences based on the calculation of a consensus matrix and the comparison of all the sequences using this consensus matrix is described. This consensus matrix contains the preference scores of each nucleotide/amino acid and gaps in every position of the alignment. Two modifications of the algorithm corresponding to the evolutionary and functional meanings of the alignment were developed. The first one solves the best-fitting problem without any penalty for end gaps and with an internal gap penalty function independent on the gap length. This algorithm should be used when comparing evolutionary-related proteins for identifying the most conservative residues. The other modification of the algorithm finds the most similar segments in the given sequences. It can be used for finding those parts of the sequences that are responsible for the same biological function. In this case the gap penalty function was chosen to be proportional to the gap length. The result of aligning amino acid sequences of neutral proteases and a compilation of 65 allosteric effectors and substrates of PEP carboxylase are presented.

Algorithms↗

GRIL: genome rearrangement and inversion locator.

UNLABELLED: GRIL is a tool to automatically identify collinear regions in a set of bacterial-size genome sequences. GRIL uses three basic steps. First, regions of high sequence identity are located. Second, some of these regions are filtered based on user-specified criteria. Finally, the remaining regions of sequence identity are used to define significant collinear regions among the sequences. By locating collinear regions of sequence, GRIL provides a basis for multiple genome alignment using current alignment systems. GRIL also provides a basis for using current inversion distance tools to infer phylogeny. AVAILABILITY: GRIL is implemented in C++ and runs on any x86-based Linux or Windows platform. It is available from http://asap.ahabs.wisc.edu/gril

Chromosome Inversion↗

Six-fold speed-up of Smith-Waterman sequence database searches using parallel processing on common microprocessors.

MOTIVATION: Sequence database searching is among the most important and challenging tasks in bioinformatics. The ultimate choice of sequence-search algorithm is that of Smith-Waterman. However, because of the computationally demanding nature of this method, heuristic programs or special-purpose hardware alternatives have been developed. Increased speed has been obtained at the cost of reduced sensitivity or very expensive hardware. RESULTS: A fast implementation of the Smith-Waterman sequence-alignment algorithm using Single-Instruction, Multiple-Data (SIMD) technology is presented. This implementation is based on the MultiMedia eXtensions (MMX) and Streaming SIMD Extensions (SSE) technology that is embedded in Intel's latest microprocessors. Similar technology exists also in other modern microprocessors. Six-fold speed-up relative to the fastest previously known Smith-Waterman implementation on the same hardware was achieved by an optimized 8-way parallel processing approach. A speed of more than 150 million cell updates per second was obtained on a single Intel Pentium III 500 MHz microprocessor. This is probably the fastest implementation of this algorithm on a single general-purpose microprocessor described to date.

Algorithms↗

Assessing strategies for improved superfamily recognition.

There are more than 200 completed genomes and over 1 million nonredundant sequences in public repositories. Although the structural data are more sparse (approximately 13,000 nonredundant structures solved to date), several powerful sequence-based methodologies now allow these structures to be mapped onto related regions in a significant proportion of genome sequences. We review a number of publicly available strategies for providing structural annotations for genome sequences, and we describe the protocol adopted to provide CATH structural annotations for completed genomes. In particular, we assess the performance of several sequence-based protocols employing Hidden Markov model (HMM) technologies for superfamily recognition, including a new approach (SAMOSA [sequence augmented models of structure alignments]) that exploits multiple structural alignments from the CATH domain structure database when building the models. Using a data set of remote homologs detected by structure comparison and manually validated in CATH, a single-seed HMM library was able to recognize 76% of the data set. Including the SAMOSA models in the HMM library showed little gain in homolog recognition, although a slight improvement in alignment quality was observed for very remote homologs. However, using an expanded 1D-HMM library, CATH-ISL increased the coverage to 86%. The single-seed HMM library has been used to annotate the protein sequences of 120 genomes from all three major kingdoms, allowing up to 70% of the genes or partial genes to be assigned to CATH superfamilies. It has also been used to recruit sequences from Swiss-Prot and TrEMBL into CATH domain superfamilies, expanding the CATH database eightfold.

Databases, Protein↗

SSAHA: a fast search method for large DNA databases.

We describe an algorithm, SSAHA (Sequence Search and Alignment by Hashing Algorithm), for performing fast searches on databases containing multiple gigabases of DNA. Sequences in the database are preprocessed by breaking them into consecutive k-tuples of k contiguous bases and then using a hash table to store the position of each occurrence of each k-tuple. Searching for a query sequence in the database is done by obtaining from the hash table the "hits" for each k-tuple in the query sequence and then performing a sort on the results. We discuss the effect of the tuple length k on the search speed, memory usage, and sensitivity of the algorithm and present the results of computational experiments which show that SSAHA can be three to four orders of magnitude faster than BLAST or FASTA, while requiring less memory than suffix tree methods. The SSAHA algorithm is used for high-throughput single nucleotide polymorphism (SNP) detection and very large scale sequence assembly. Also, it provides Web-based sequence search facilities for Ensembl projects.

Algorithms↗