Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Sequence of the canine herpesvirus thymidine kinase gene: taxon-preferred amino acid residues in the alphaherpesviral thymidine kinases.

Multiple sequence alignments of evolutionarily related proteins are finding increasing use as indicators of critical amino acid residues necessary for structural stability or involved in functional domains responsible for catalytic activities. In the past, a number of alignments have provided such information for the herpesviral thymidine kinases, for which three-dimensional structures are not yet available. We have sequenced the thymidine kinase gene of a canine herpesvirus, and with a multiple alignment have identified amino acids preferentially conserved in either of two taxons, the genera Varicellovirus and Simplexvirus, of the subfamily Alphaherpesvirinae. Since some regions of the thymidine kinases show otherwise elevated levels of substitutional tolerance, these conserved amino acids are candidates for critical residues which have become fixed through selection during the evolutionary divergence of these enzymes. Several pairs with distinctive patterns of distribution among the various viruses occur in or near highly conserved sequence motifs previously proposed to form the catalytic site, and we speculate that they may represent interacting, co-ordinately variable residues.

Alphaherpesvirinae↗

CINEMA--a novel colour INteractive editor for multiple alignments.

CINEMA is a new editor for manipulating and generating multiple sequence alignments. The program provides both an interface to existing databases of alignments on the Internet and a tool for constructing and modifying alignments locally. It is written in Java, so executable code will run on most major desktop platforms without modification. The implementation is highly flexible, so the applet can be easily customised with additional functions; and the object classes are reusable, promoting rapid development of program extensions. Formerly, such extended functionality might have been provided via browser plug-ins, which have to be downloaded and installed on every client before loading data. Now, for the first time, an applet is available that allows interactive client-side processing of an alignment, which can then be stored or processed automatically on the server. The program is embedded in a comprehensive help file and is accessible both as a stand-alone tool on UCL's Bioinformatics Server; http:/(/)www.biochem.ucl.ac.uk/bsm/dbbrowser+ ++/CINEMA2.02/, and as an integral part of the PRINTS protein fingerprint database. Exploitation of such novel technologies revolutionises the way users may interact with public databases in the future: bioinformatics centres need not simply provide data, but are now able to offer the means by which information is visualised and manipulated, without the requirement for users to install software.

Color Perception↗

An assessment of the phylogenetic relationship among sugarcane and related taxa based on the nucleotide sequence of 5S rRNA intergenic spacers.

5S rRNA intergenic spacers were amplified from two elite sugarcane (Saccharum hybrids) cultivars and their related taxa by polymerase chain reaction (PCR) with 5S rDNA consensus primers. Resulting PCR products were uniform in length from each accession but exhibited some degree of length variation among the sugarcane accessions and related taxa. These PCR products did not always cross hybridize in Southern blot hybridization experiments. These PCR products were cloned into a commercial plasmid vector PCR 2.1 and sequenced. Direct sequencing of cloned PCR products revealed spacer length of 231-237 bp for S. officinarum, 233-237 for sugarcane cultivars, 228-238 bp for S. spontaneum, 239-252 bp for S. giganteum, 385-410 bp for Erianthus spp., 226-230 bp for Miscanthus sinensis Zebra, 206-207 bp for M. sinensis IMP 3057, 207-209 bp for Sorghum bicolor, and 247-249 bp for Zea mays. Nucleotide sequence polymorphism were found at both the segment and single nucleotide level. A consensus sequence for each taxon was obtained by Align X. Multiple sequences were aligned and phylogenetic trees constructed using Align X. CLUSTAL and DNAMAN programs. In general, accessions of the following taxa tended to group together to form distinct clusters: S. giganteum, Erianthus spp., M. sinensis, S. bicolor, and Z. mays. However, the two S. officinarum clones and two sugarcane cultivars did not form distinct clusters but interrelated within the S. spontaneum cluster. The disclosure of these 5S rRNA intergenic spacer sequences will facilitate marker-assisted breeding in sugarcane.

Base Sequence↗

Protein family annotation in a multiple alignment viewer.

SUMMARY: The Pfaat protein family alignment annotation tool is a Java-based multiple sequence alignment editor and viewer designed for protein family analysis. The application merges display features such as dendrograms, secondary and tertiary protein structure with SRS retrieval, subgroup comparison, and extensive user-annotation capabilities. AVAILABILITY: The program and source code are freely available from the authors under the GNU General Public License at http://www.pfizerdtc.com

Amino Acid Sequence↗

Simultaneous sequence alignment and tree construction using hidden Markov models.

We present a new algorithm (SATCHMO) that simultaneously estimates a tree and generates a set of multiple sequence alignments given a set of protein sequences. Alignments are constructed for each node in the tree. These alignments predict the structurally conserved elements of the sequences in a subtree and are therefore of different lengths, and represent different amino acid preferences, at different nodes. Hidden Markov Models (HMMs) are also generated for each node and are used to determine branching order, to align sequences and to predict structurally alignable regions. In experiments on the BAliBASE benchmark alignment database, SATCHMO is shown to perform comparably to ClustalW and the UCSC SAM HMM software. Results using SATCHMO to identify protein domains are demonstrated on potassium channels, with implications for the mechanism by which tumor necrosis factor alpha affects potassium current.

Algorithms↗

A simple and fast approach to prediction of protein secondary structure from multiply aligned sequences with accuracy above 70%.

To improve secondary structure predictions in protein sequences, the information residing in multiple sequence alignments of substituted but structurally related proteins is exploited. A database comprised of 70 protein families and a total of 2,500 sequences, some of which were aligned by tertiary structural superpositions, was used to calculate residue exchange weight matrices within alpha-helical, beta-strand, and coil substructures, respectively. Secondary structure predictions were made based on the observed residue substitutions in local regions of the multiple alignments and the largest possible associated exchange weights in each of the three matrix types. Comparison of the observed and predicted secondary structure on a per-residue basis yielded a mean accuracy of 72.2%. Individual alpha-helix, beta-strand, and coil states were respectively predicted at 66.7, and 75.8% correctness, representing a well-balanced three-state prediction. The accuracy level, verified by cross-validation through jack-knife tests on all protein families, dropped, on average, to only 70.9%, indicating the rigor of the prediction procedure. On the basis of robustness, conceptual clarity, accuracy, and executable efficiency, the method has considerable advantage, especially with its sole reliance on amino acid substitutions within structurally related proteins.

Algorithms↗

CLOURE: Clustal Output Reformatter, a program for reformatting ClustalX/ClustalW outputs for SNP analysis and molecular systematics.

We describe a program (and a website) to reformat the ClustalX/ClustalW outputs to a format that is widely used in the presentation of sequence alignment data in SNP analysis and molecular systematic studies. This program, CLOURE, CLustal OUtput REformatter, takes the multiple sequence alignment file (nucleic acid or protein) generated from Clustal as input files. The CLOURE-D format presents the Clustal alignment in a format that highlights only the different nucleotides/residues relative to the first query sequence. The program has been written in Visual Basic and will run on a Windows platform. The downloadable program, as well as a web-based server which has also been developed, can be accessed at http://imtech.res.in/~anand/cloure.html.

Internet↗

PROTOGENE: turning amino acid alignments into bona fide CDS nucleotide alignments.

We describe Protogene, a server that can turn a protein multiple sequence alignment into the equivalent alignment of the original gene coding DNA. Protogene relies on a pipeline where every initial protein sequence is BLASTed against RefSeq or NR. The annotation associated with potential matches is used to identify the gene sequence. This gene sequence is then aligned with the query protein using Exonerate in order to extract a coding nucleotide sequence matching the original protein. Protogene can handle protein fragments and will return every CDS coding for a given protein, even if they occur in different genomes. Protogene is available from http://www.tcoffee.org/.

Base Sequence↗

THoR: a tool for domain discovery and curation of multiple alignments.

We describe a tool, THoR, that automatically creates and curates multiple sequence alignments representing protein domains. This exploits both PSI-BLAST and HMMER algorithms and provides an accurate and comprehensive alignment for any domain family. The entire process is designed for use via a web-browser, with simple links and cross-references to relevant information, to assist the assessment of biological significance. THoR has been benchmarked for accuracy using the SMART and pufferfish genome databases.

Algorithms↗

Increased detection of structural templates using alignments of designed sequences.

Protein structure prediction by comparative modeling benefits greatly from the use of multiple sequence alignment information to improve the accuracy of structural template identification and the alignment of target sequences to structural templates. Unfortunately, this benefit is limited to those protein sequences for which at least several natural sequence homologues exist. We show here that the use of large diverse alignments of computationally designed protein sequences confers many of the same benefits as natural sequences in identifying structural templates for comparative modeling targets. A large-scale massively parallelized application of an all-atom protein design algorithm, including a simple model of peptide backbone flexibility, has allowed us to generate 500 diverse, non-native, high-quality sequences for each of 264 protein structures in our test set. PSI-BLAST searches using the sequence profiles generated from the designed sequences ("reverse" BLAST searches) give near-perfect accuracy in identifying true structural homologues of the parent structure, with 54% coverage. In 41 of 49 genomes scanned using reverse BLAST searches, at least one novel structural template (not found by the standard method of PSI-BLAST against PDB) is identified. Further improvements in coverage, through optimizing the scoring function used to design sequences and continued application to new protein structures beyond the test set, will allow this method to mature into a useful strategy for identifying distantly related structural templates.

Algorithms↗

BAliBASE (Benchmark Alignment dataBASE): enhancements for repeats, transmembrane sequences and circular permutations.

BAliBASE is specifically designed to serve as an evaluation resource to address all the problems encountered when aligning complete sequences. The database contains high quality, manually constructed multiple sequence alignments together with detailed annotations. The alignments are all based on three-dimensional structural superpositions, with the exception of the transmembrane sequences. The first release provided sets of reference alignments dealing with the problems of high variability, unequal repartition and large N/C-terminal extensions and internal insertions. Here we describe version 2.0 of the database, which incorporates three new reference sets of alignments containing structural repeats, trans-membrane sequences and circular permutations to evaluate the accuracy of detection/prediction and alignment of these complex sequences. BAliBASE can be viewed at the web site http://www-igbmc.u-strasbg. fr/BioInfo/BAliBASE2/index.html or can be downloaded from ftp://ftp-igbmc.u-strasbg.fr/pub/BAliBASE2 /.

Algorithms↗

A method for finding single-nucleotide polymorphisms with allele frequencies in sequences of deep coverage.

BACKGROUND: The allele frequencies of single-nucleotide polymorphisms (SNPs) are needed to select an optimal subset of common SNPs for use in association studies. Sequence-based methods for finding SNPs with allele frequencies may need to handle thousands of sequences from the same genome location (sequences of deep coverage). RESULTS: We describe a computational method for finding common SNPs with allele frequencies in single-pass sequences of deep coverage. The method enhances a widely used program named PolyBayes in several aspects. We present results from our method and PolyBayes on eighteen data sets of human expressed sequence tags (ESTs) with deep coverage. The results indicate that our method used almost all single-pass sequences in computation of the allele frequencies of SNPs. CONCLUSION: The new method is able to handle single-pass sequences of deep coverage efficiently. Our work shows that it is possible to analyze sequences of deep coverage by using pairwise alignments of the sequences with the finished genome sequence, instead of multiple sequence alignments.

Computers, Molecular↗

Two Sample Logo: a graphical representation of the differences between two sets of sequence alignments.

SUMMARY: Two Sample Logo is a web-based tool that detects and displays statistically significant differences in position-specific symbol compositions between two sets of multiple sequence alignments. In a typical scenario, two groups of aligned sequences will share a common motif but will differ in their functional annotation. The inclusion of the background alignment provides an appropriate underlying amino acid or nucleotide distribution and addresses intersite symbol correlations. In addition, the difference detection process is sensitive to the sizes of the aligned groups. Two Sample Logo extends WebLogo, a widely-used sequence logo generator. The source code is distributed under the MIT Open Source license agreement and is available for download free of charge.

Algorithms↗

DINAMO: interactive protein alignment and model building.

MOTIVATION: To facilitate the process of structure prediction by both comparative modeling and fold recognition, we describe DINAMO, an interactive protein alignment building and model evaluation tool that dynamically couples a multiple sequence alignment editor to a molecular graphics display. DINAMO allows the user to optimize the alignment and model to satisfy the known heuristics of protein structure by means of a set of analysis tools. The analysis tools return information to both the alignment editor and graphics model in the form of visual cues (color, shape), allowing for rapid evaluation. Several analysis tools may be employed, including residue conservation, residue properties (charge, hydrophobicity, volume), residue environmental preference, and secondary structure propensity. RESULTS: We demonstrate DINAMO by building a model for submission in the 3rd annual Critical Assessment of Techniques for Protein Structure Prediction (CASP3) contest. AVAILABILITY: DINAMO is freely available as a local application or Web-based Java applet at http://tito.ucsc.edu/dinamo

Amino Acid Sequence↗

Correlating patterns in alignments of polymorphic sequences with experimental assays.

A general algorithm is presented for identifying sets of positions in multiple sequence alignments that best characterize an a priori partitioning such as those determined by inhibition studies or other experimental techniques. The algorithm explores combinations of polymorphic columns in the alignment and evaluates how well these sites reflect the original input partition. Partitions across the polymorphic columns are derived using a tree building procedure with conventional amino acid substitution matrices. Elucidation of those amino acids which govern the biochemical behaviour of a protein with a given substrate or inhibitor can provide insights towards an understanding of the tertiary conformation of the protein. Since it is likely that such positions will be spatially clustered in the protein fold, these positions may give rise to useful distance constraints for substantiating model protein structures. The method is exemplified using data for a set of human mu class glutathione S-transferases. A novel aspect for predicting the behaviour of new polymorphic sequences is also discussed.

Algorithms↗

Sequence-structure homology recognition by iterative alignment refinement and comparative modeling.

Our approach to fold recognition for the fourth critical assessment of techniques for protein structure prediction (CASP4) experiment involved the use of the FUGUE sequence-structure homology recognition program (http://www-cryst.bioc.cam.ac.uk/fugue), followed by model building. We treat models as hypotheses and examine these to determine whether they explain the available data. Our method depends heavily on environment-specific substitution tables derived from our database of structural alignments of homologous proteins (HOMSTRAD, http://www-cryst.bioc.cam.ac.uk/homstrad/). FUGUE uses these tables to incorporate structural information into profiles created from HOMSTRAD alignments that are matched against a profile created for the target from multiple sequence alignment. In addition, environment-specific substitution tables are used throughout the modeling procedure and as part of the model evaluation. Annotation of sequence alignments with JOY, to reflect local structural features, proved valuable, both for modifying hypotheses, and for rejecting predictions when the expected pattern of conservation is not observed. Our stringency in rejecting incorrect predictions led us to submit a relatively small number of models, including only a low number of false positives, resulting in a high average score.

Amino Acid Sequence↗

Indelign: a probabilistic framework for annotation of insertions and deletions in a multiple alignment.

MOTIVATION: A quantitative study of molecular evolutionary events such as substitutions, insertions and deletions from closely related genomes requires (1) an accurate multiple sequence alignment program and (2) a method to annotate the insertions and deletions that explain the 'gaps' in the alignment. Although the former requirement has been extensively addressed, the latter problem has received little attention, especially in a comprehensive probabilistic framework. RESULTS: Here, we present Indelign, a program that uses a probabilistic evolutionary model to compute the most likely scenario of insertions and deletions consistent with an input multiple alignment. It is also capable of modifying the given alignment so as to obtain a better agreement with the evolutionary model. We find close to optimal performance and substantial improvement over alternative methods, in tests of Indelign on synthetic data. We use Indelign to analyze regulatory sequences in Drosophila, and find an excess of insertions over deletions, which is different from what has been reported for neutral sequences. AVAILABILITY: The Indelign program may be downloaded from the website http://veda.cs.uiuc.edu/indelign/ SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.

Algorithms↗