Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Malate dehydrogenase: a model for structure, evolution, and catalysis.

Malate dehydrogenases are widely distributed and alignment of the amino acid sequences show that the enzyme has diverged into 2 main phylogenetic groups. Multiple amino acid sequence alignments of malate dehydrogenases also show that there is a low degree of primary structural similarity, apart from in several positions crucial for nucleotide binding, catalysis, and the subunit interface. The 3-dimensional structures of several malate dehydrogenases are similar, despite their low amino acid sequence identity. The coenzyme specificity of malate dehydrogenase may be modulated by substitution of a single residue, as can the substrate specificity. The mechanism of catalysis of malate dehydrogenase is similar to that of lactate dehydrogenase, an enzyme with which it shares a similar 3-dimensional structure. Substitution of a single amino acid residue of a lactate dehydrogenase changes the enzyme specificity to that of a malate dehydrogenase, but a similar substitution in a malate dehydrogenase resulted in relaxation of the high degree of specificity for oxaloacetate. Knowledge of the 3-dimensional structures of malate and lactate dehydrogenases allows the redesign of enzymes by rational rather than random mutation and may have important commercial implications.

Amino Acid Sequence↗

Optimal classification of protein sequences and selection of representative sets from multiple alignments: application to homologous families and lessons for structural genomics.

Hierarchical classification is probably the most popular approach to group related proteins. However, there are a number of problems associated with its use for this purpose. One is that the resulting tree showing a nested sequence of groups may not be the most suitable representation of the data. Another is that visual inspection is the most common method to decide the most appropriate number of subsets from a tree. In fact, classification of proteins in general is bedevilled with the need for subjective thresholds to define group membership (e.g., 'significant' sequence identity for homologous families). Such arbitrariness is not only intellectually unsatisfying but also has important practical consequences. For instance, it hinders meaningful identification of protein targets for structural genomics. I describe an alternative approach to cluster related proteins without the need for an a priori threshold: one, through its use of dynamic programming, which is guaranteed to produce globally optimal solutions at all levels of partition granularity. Grouping proteins according to weights assigned to their aligned sequences makes it possible to delineate dynamically a 'core-periphery' structure within families. The 'core' of a protein family comprises the most typical sequences while the 'periphery' consists of the atypical ones. Further, a new sequence weighting scheme that combines the information in all the multiply aligned positions of an alignment in a novel way is put forward. Instead of averaging over all positions, this procedure takes into account directly the distribution of sequence variability along an alignment. The relationships between sequence weights and sequence identity are investigated for 168 families taken from HOMSTRAD, a database of protein structure alignments for homologous families. An exact solution is presented for the problem of how to select the most representative pair of sequences for a protein family. Extension of this approach by a greedy algorithm allows automatic identification of a minimal set of aligned sequences. The results of this analysis are available on the Web at http://mathbio.nimr.mrc.ac.uk/~amay.

Algorithms↗

Vertebrate gene finding from multiple-species alignments using a two-level strategy.

BACKGROUND: One way in which the accuracy of gene structure prediction in vertebrate DNA sequences can be improved is by analyzing alignments with multiple related species, since functional regions of genes tend to be more conserved. RESULTS: We describe DOGFISH, a vertebrate gene finder consisting of a cleanly separated site classifier and structure predictor. The classifier scores potential splice sites and other features, using sequence alignments between multiple vertebrate species, while the structure predictor hypothesizes coding transcripts by combining these scores using a simple model of gene structure. This also identifies and assigns confidence scores to possible additional exons. Performance is assessed on the ENCODE regions. We predict transcripts and exons across the whole human genome, and identify over 10,000 high confidence new coding exons not in the Ensembl gene set. CONCLUSION: We present a practical multiple species gene prediction method. Accuracy improves as additional species, up to at least eight, are introduced. The novel predictions of the whole-genome scan should support efficient experimental verification.

Animals↗

Evolutionary HMMs: a Bayesian approach to multiple alignment.

MOTIVATION: We review proposed syntheses of probabilistic sequence alignment, profiling and phylogeny. We develop a multiple alignment algorithm for Bayesian inference in the links model proposed by Thorne et al. (1991, J. Mol. Evol., 33, 114-124). The algorithm, described in detail in Section 3, samples from and/or maximizes the posterior distribution over multiple alignments for any number of DNA or protein sequences, conditioned on a phylogenetic tree. The individual sampling and maximization steps of the algorithm require no more computational resources than pairwise alignment. METHODS: We present a software implementation (Handel) of our algorithm and report test results on (i) simulated data sets and (ii) the structurally informed protein alignments of BAliBASE (Thompson et al., 1999, Nucleic Acids Res., 27, 2682-2690). RESULTS: We find that the mean sum-of-pairs score (a measure of residue-pair correspondence) for the BAliBASE alignments is only 13% lower for Handelthan for CLUSTALW(Thompson et al., 1994, Nucleic Acids Res., 22, 4673-4680), despite the relative simplicity of the links model (CLUSTALW uses affine gap scores and increased penalties for indels in hydrophobic regions). With reference to these benchmarks, we discuss potential improvements to the links model and implications for Bayesian multiple alignment and phylogenetic profiling. AVAILABILITY: The source code to Handelis freely distributed on the Internet at http://www.biowiki.org/Handel under the terms of the GNU Public License (GPL, 2000, http://www.fsf.org./copyleft/gpl.html).

Algorithms↗

ProClass protein family database.

ProClass is a protein family database that organizes non-redundant sequence entries into families defined collectively by PIR superfamilies and PROSITE patterns. By combining global similarities and functional motifs into a single classification scheme, ProClass helps to reveal domain and family relationships and classify multi-domain proteins. The database currently consists of >155 000 sequence entries retrieved from both PIR-International and SWISS-PROT databases. Approximately 92 000 or 60% of the ProClass entries are classified into approximately 6000 families, including a large number of new members detected by our GeneFIND family identification system. The ProClass motif collection contains approximately 72 000 motif sequences and >1300 multiple alignments for all PROSITE patterns, including >21 000 matches not listed in PROSITE and mostly detected from unique PIR sequences. To maximize family information retrieval, the database provides links to various protein family, domain, alignment and structural class databases. With its high classification rate and comprehensive family relationships, ProClass can be used to support full-scale genomic annotation. The database, now being implemented in an object-relational database management system, is available for online sequence search and record retrieval from our WWW server at http://pir.georgetown.edu/gfserver/proclass.html

Databases, Factual↗

Identifying multiple alignment regions satisfying simple formulas and patterns.

MOTIVATION: When studying multiple alignments of genomic sequences one frequently aims to locate and count regions which satisfy a set of constraints. These regions may be putatively functional, but researchers may also be interested in quantifying the frequency of occurrences of certain patterns. RESULTS: We have developed a program that applies simple formulas and pattern specifications to multiple alignments, reporting the positions and counts of conforming regions. As an example, we have navigated a 15-species alignment of the CAV2-CAV1 region and outlined some findings regarding PPARgamma binding sites. AVAILABILITY: Our software and the accompanying documentation can be obtained at no charge by contacting the authors. It can also be accessed at http://ranger.uta.edu/~nick/compgen

Algorithms↗

Computing TaqMan probes for multiplex PCR detection of E. coli O157 serotypes in water.

Diarrheagenic E. coli strains contribute to water related diseases in urban and rural environment in developing and developed world. E. coli pathotype and pathogenicity varies due to complex multifactorial mechanism involving a large number of virulence factors. Rapid assessment of the virulence pattern of E. coli isolates is possible by Real-Time PCR probes like TaqMan. For designing TaqMan probes and primers for multiplex PCR selected E. coli gene sequences: stx1, stx2, hlyA, chuA, eae, lacZ, lamB and fimA were retrieved from NCBI's GenBank database. The alignment of the multiple sequences and analysis of conserved sequences was carried out using ClustalW and BLAST programs. The primers and Taqmen probes were designed using Beacon Designer software version 2.1 for two multiplexed PCR assays. In silico PCR simulation of these assays showed PCR products for stx2 (248bp) stx1 (102 bp), lacZ (228bp) and lamB (86 bp) in multiplex #1 and eae (200bp), chuA (147 bp), hlyA (141bp) and fimA (79 bp) in multiplex #2, respectively. These multiplexed PCR amplification products and probes can be used to identify and confirm presence of O157:H7/ H7-, O157:H43/45 and O26:H-/H11 serotypes. In conclusion, multiplex Real-Time Polymerase Chain Reaction oligomers and TaqMan probes designed and validated in silico will be helpful in management of water quality and outbreaks, by improving specificity and minimizing time needed for in vitro verification work.

Base Sequence↗

Sequence determinants for the reaction specificity of murine (12R)-lipoxygenase: targeted substrate modification and site-directed mutagenesis.

Mammalian lipoxygenases (LOXs) are categorized with respect to their positional specificity of arachidonic acid oxygenation. Site-directed mutagenesis identified sequence determinants for the positional specificity of these enzymes, and a critical amino acid for the stereoselectivity was recently discovered. To search for sequence determinants of murine (12R)-LOX, we carried out multiple amino acid sequence alignments and found that Phe(390), Gly(441), Ala(455), and Val(631) align with previously identified positional determinants of S-LOX isoforms. Multiple site-directed mutagenesis studies on Phe(390) and Ala(455) did not induce specific alterations in the reaction specificity, but yielded enzyme species with reduced specific activities and stereo random product patterns. Mutation of Gly(441) to Ala, which caused drastic alterations in the reaction specificity of other LOX isoforms, failed to induce major alterations in the positional specificity of mouse (12R)-LOX, but markedly modified the enantioselectivity of the enzyme. When Val(631), which aligns with the positional determinant Ile(593) of rabbit 15-LOX, was mutated to a less space-filling residue (Ala or Gly), we obtained an enzyme species with augmented catalytic activity and specifically altered reaction characteristics (major formation of chiral (11R)-hydroxyeicosatetraenoic acid methyl ester). The importance of Val(631) for the stereo control of murine (12R)-LOX was confirmed with other substrates such as methyl linoleate and 20-hydroxyeicosatetraenoic acid methyl ester. These data identify Val(631) as the major sequence determinant for the specificity of murine (12R)-LOX. Furthermore, we conclude that substrate fatty acids may adopt different catalytically productive arrangements at the active site of murine (12R)-LOX and that each of these arrangements may lead to the formation of chiral oxygenation products.

12-Hydroxy-5,8,10,14-eicosatetraenoic Acid↗

MAVID multiple alignment server.

MAVID is a multiple alignment program suitable for many large genomic regions. The MAVID web server allows biomedical researchers to quickly obtain multiple alignments for genomic sequences and to subsequently analyse the alignments for conserved regions. MAVID has been successfully used for the alignment of closely related species such as primates and also for the alignment of more distant organisms such as human and fugu. The server is fast, capable of aligning hundreds of kilobases in less than a minute. The multiple alignment is used to build a phylogenetic tree for the sequences, which is subsequently used as a basis for identifying conserved regions in the alignment. The server can be accessed at http://baboon.math.berkeley.edu/mavid/.

Animals↗

Prediction of unfolded segments in a protein sequence based on amino acid composition.

MOTIVATION: Partially and wholly unstructured proteins have now been identified in all kingdoms of life--more commonly in eukaryotic organisms. This intrinsic disorder is related to certain critical functions. Apart from their fundamental interest, unstructured regions in proteins may prevent crystallization. Therefore, the prediction of disordered regions is an important aspect for the understanding of protein function, but may also help to devise genetic constructs. RESULTS: In this paper we present a computational tool for the detection of unstructured regions in proteins based on two properties of unfolded fragments: (1) disordered regions have a biased composition and (2) they usually contain either small or no hydrophobic clusters. In order to quantify these two facts we first calculate the amino acid distributions in structured and unstructured regions. Using this distribution, we calculate for a given sequence fragment the probability to be part of either a structured or an unstructured region. For each amino acid, the distance to the nearest hydrophobic cluster is also computed. Using these three values along a protein sequence allows us to predict unstructured regions, with very simple rules. This method requires only the primary sequence, and no multiple alignment, which makes it an adequate method for orphan proteins. AVAILABILITY: http://genomics.eu.org/

Algorithms↗

Structure, specificity and function of cyclomaltodextrinase, a multispecific enzyme of the alpha-amylase family.

Cyclomaltodextrinase (CDase, EC 3.2.1.54), maltogenic amylase (EC 3. 2.1.133), and neopullulanase (EC 3.2.1.135) are reported to be capable of hydrolyzing all or two of the following three types of substrates: cyclomaltodextrins (CDs); pullulan; and starch. These enzymes hydrolyze CDs and starch to maltose and pullulan to panose by cleavage of alpha-1,4 glycosidic bonds whereas alpha-amylases essentially lack activity on CDs and pullulan. They also catalyze transglycosylation of oligosaccharides to the C3-, C4- or C6-hydroxyl groups of various acceptor sugar molecules. The present review surveys the biochemical, enzymatic, and structural properties of three types of such enzymes as defined based on the substrate specificity toward the CDs: type I, cyclomaltodextrinase and maltogenic amylase that hydrolyze CDs much faster than pullulan and starch; type II, Thermoactinomyces vulgaris amylase II (TVA II) that hydrolyzes CDs much less efficiently than pullulan; and type III, neopullulanase that hydrolyzes pullulan efficiently, but remains to be reported to hydrolyze CDs. These three types of enzymes exhibit 40-60% amino acid sequence identity. They occur in the cytoplasm of bacteria and have molecular masses from 62 to 90 kDa which are slightly larger than those of most alpha-amylases. Multiple amino acid sequence alignment and crystal structures of maltogenic amylase and TVA II reveal the presence of an N-terminal extension of approximately 130 residues not found in alpha-amylases. This unique N-terminal domain as seen in the crystal structures apparently contributes to the active site structure leading to the distinct substrate specificity through a dimer formation. In aqueous solution, most of these enzymes show a monomer-dimer equilibrium. The present review discusses the multiple specificity in the light of the oligomerization and the molecular structures arriving at a clarified enzyme classification. Finally, a physiological role of the enzymes is proposed.

Amino Acid Sequence↗

Eukaryotic DNA polymerase amino acid sequence required for 3'----5' exonuclease activity.

We have identified an amino-proximal sequence motif, Phe-Asp-Ile-Glu-Thr, in Saccharomyces cerevisiae DNA polymerase II that is almost identical to a sequence comprising part of the 3'----5' exonuclease active site of Escherichia coli DNA polymerase I. Similar motifs were identified by amino acid sequence alignment in related, aphidicolin-sensitive DNA polymerases possessing 3'----5' proofreading exonuclease activity. Substitution of Ala for the Asp and Glu residues in the motif reduced the exonuclease activity of partially purified DNA polymerase II at least 100-fold while preserving the polymerase activity. Yeast strains expressing the exonuclease-deficient DNA polymerase II had on average about a 22-fold increase in spontaneous mutation rate, consistent with a presumed proofreading role in vivo. In multiple amino acid sequence alignments of this and two other conserved motifs described previously, five residues of the 3'----5' exonuclease active site of E. coli DNA polymerase I appeared to be invariant in aphidicolin-sensitive DNA polymerases known to possess 3'----5' proofreading exonuclease activity. None of these residues, however, appeared to be identifiable in the catalytic subunits of human, yeast, or Drosophila alpha DNA polymerases.

Amino Acid Sequence↗

The HHpred interactive server for protein homology detection and structure prediction.

HHpred is a fast server for remote protein homology detection and structure prediction and is the first to implement pairwise comparison of profile hidden Markov models (HMMs). It allows to search a wide choice of databases, such as the PDB, SCOP, Pfam, SMART, COGs and CDD. It accepts a single query sequence or a multiple alignment as input. Within only a few minutes it returns the search results in a user-friendly format similar to that of PSI-BLAST. Search options include local or global alignment and scoring secondary structure similarity. HHpred can produce pairwise query-template alignments, multiple alignments of the query with a set of templates selected from the search results, as well as 3D structural models that are calculated by the MODELLER software from these alignments. A detailed help facility is available. As a demonstration, we analyze the sequence of SpoVT, a transcriptional regulator from Bacillus subtilis. HHpred can be accessed at http://protevo.eb.tuebingen.mpg.de/hhpred.

Bacterial Proteins↗

Selection of circularization sites in a group I IVS RNA requires multiple alignments of an internal template-like sequence.

Circularization and reverse circularization of the Tetrahymena thermophila rRNA intervening sequence resemble the first and second steps in splicing, respectively. However, site-specific base substitutions show that different nucleotides are involved in selection of the 5' splice site and the circularization sites. Furthermore, a substitution at the major circularization site that prevents circularization can be suppressed by second substitutions at two different nucleotide positions. A model is proposed in which adjacent and overlapping sequences can function as a binding site, forming a short duplex with the sequence at the circularization site and thus directing circularization and reverse circularization. Because the 5' exon-binding site and three potential circularization binding sites fall within a contiguous eight nucleotide region, this sequence may translocate relative to the catalytic core of the ribozyme in a template-like manner.

Animals↗

CLANS: a Java application for visualizing protein families based on pairwise similarity.

SUMMARY: The main source of hypotheses on the structure and function of new proteins is their homology to proteins with known properties. Homologous relationships are typically established through sequence similarity searches, multiple alignments and phylogenetic reconstruction. In cases where the number of potential relationships is large, for example in P-loop NTPases with many thousands of members, alignments and phylogenies become computationally demanding, accumulate errors and lose resolution. In search of a better way to analyze relationships in large sequence datasets we have developed a Java application, CLANS (CLuster ANalysis of Sequences), which uses a version of the Fruchterman-Reingold graph layout algorithm to visualize pairwise sequence similarities in either two-dimensional or three-dimensional space. AVAILABILITY: CLANS can be downloaded at http://protevo.eb.tuebingen.mpg.de/download.

Algorithms↗

OrthoGUI: graphical presentation of Orthostrapper results.

SUMMARY: Orthostrapper is a program that calculates orthology support values for pairs of sequences in a multiple alignment (Storm and Sonnhammer, Bioinformatics, 18, 92-99, 2002). Here we present OrthoGUI, a web interface and display tool for Orthostrapper analysis. OrthoGUI visualizes the Orthostrapper output in both tabular and tree representations, and can also apply a clustering algorithm to identify groups of multiple orthologs, which are indicated by colour coding. AVAILABILITY: http://www.cgb.ki.se/OrthoGUI CONTACT: erik.sonnhammer@cgb.ki.se

ATP-Binding Cassette Transporters↗

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software↗

Immunoglobulin-binding FcrA and Enn proteins and M proteins of group A streptococci evolved independently from a common ancestral protein.

Significant sequence homology between M proteins and immunoglobulin (Ig)-binding proteins of group A streptococci suggests that these proteins arose by gene duplication followed by the development of functional diversity due to mutations and intragenic recombinations. The deduced sequence of multiple Ig-binding proteins and M proteins were compared to distinguish between two evolutionary models. Did these functionally distinct genes originate in the distant past from duplication of a common ancestral gene and then functionally evolve independently or did they evolve more recently, one from the other by duplication of a fixed gene? Multiple alignments of conserved sequences of these proteins are consistent with the former hypothesis. Comparison of N termini of Ig-binding proteins revealed less diversity than that of the M proteins' N termini, suggesting that these proteins are under less selective pressure to change.

Amino Acid Sequence↗