Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

DINAMO: interactive protein alignment and model building.

MOTIVATION: To facilitate the process of structure prediction by both comparative modeling and fold recognition, we describe DINAMO, an interactive protein alignment building and model evaluation tool that dynamically couples a multiple sequence alignment editor to a molecular graphics display. DINAMO allows the user to optimize the alignment and model to satisfy the known heuristics of protein structure by means of a set of analysis tools. The analysis tools return information to both the alignment editor and graphics model in the form of visual cues (color, shape), allowing for rapid evaluation. Several analysis tools may be employed, including residue conservation, residue properties (charge, hydrophobicity, volume), residue environmental preference, and secondary structure propensity. RESULTS: We demonstrate DINAMO by building a model for submission in the 3rd annual Critical Assessment of Techniques for Protein Structure Prediction (CASP3) contest. AVAILABILITY: DINAMO is freely available as a local application or Web-based Java applet at http://tito.ucsc.edu/dinamo

Amino Acid Sequence↗

A secondary structure model of the integrin alpha subunit N-terminal domain based on analysis of multiple alignments.

The integrins are alpha/beta heterodimeric proteins which mediate cell-matrix and cell-cell interactions. Current data indicate that the N-terminal moiety of the alpha subunit is involved in ligand binding. This region of the receptor is made up of a seven-fold repeated sequence of unknown structure which contains EF-hand-like putative divalent cation-binding sites. Recent studies have shown that multiple sequence alignments can be analysed to yield secondary structure predictions. Therefore, to obtain a model structure for the integrin alpha subunit N-terminal domain repeat, a large alignment of the seven repeats from sixteen integrin sequences was generated. Two methods of analysis were used: First, Chou and Fasman and Garnier, Osguthorpe and Robson predictions were carried out for individual sequences and the consensus predictions derived. Consensus hydrophobicity and chain flexibility data were also used to provide additional data. Second, sites of conservation and variation were analysed by a computer program STAMA (STructure After Multiple Alignment) to yield a secondary structure prediction. The two analyses gave essentially the same predicted structure: undefined region, loop, alpha-helix, beta-strand, divalent cation-binding loop, beta-strand, putative turn, loop, beta-strand. This is the first model structure to be presented for an integrin domain. Its implications for integrin function are discussed.

Amino Acid Sequence↗

Correlating patterns in alignments of polymorphic sequences with experimental assays.

A general algorithm is presented for identifying sets of positions in multiple sequence alignments that best characterize an a priori partitioning such as those determined by inhibition studies or other experimental techniques. The algorithm explores combinations of polymorphic columns in the alignment and evaluates how well these sites reflect the original input partition. Partitions across the polymorphic columns are derived using a tree building procedure with conventional amino acid substitution matrices. Elucidation of those amino acids which govern the biochemical behaviour of a protein with a given substrate or inhibitor can provide insights towards an understanding of the tertiary conformation of the protein. Since it is likely that such positions will be spatially clustered in the protein fold, these positions may give rise to useful distance constraints for substantiating model protein structures. The method is exemplified using data for a set of human mu class glutathione S-transferases. A novel aspect for predicting the behaviour of new polymorphic sequences is also discussed.

Algorithms↗

Sequence-structure homology recognition by iterative alignment refinement and comparative modeling.

Our approach to fold recognition for the fourth critical assessment of techniques for protein structure prediction (CASP4) experiment involved the use of the FUGUE sequence-structure homology recognition program (http://www-cryst.bioc.cam.ac.uk/fugue), followed by model building. We treat models as hypotheses and examine these to determine whether they explain the available data. Our method depends heavily on environment-specific substitution tables derived from our database of structural alignments of homologous proteins (HOMSTRAD, http://www-cryst.bioc.cam.ac.uk/homstrad/). FUGUE uses these tables to incorporate structural information into profiles created from HOMSTRAD alignments that are matched against a profile created for the target from multiple sequence alignment. In addition, environment-specific substitution tables are used throughout the modeling procedure and as part of the model evaluation. Annotation of sequence alignments with JOY, to reflect local structural features, proved valuable, both for modifying hypotheses, and for rejecting predictions when the expected pattern of conservation is not observed. Our stringency in rejecting incorrect predictions led us to submit a relatively small number of models, including only a low number of false positives, resulting in a high average score.

Amino Acid Sequence↗

Indelign: a probabilistic framework for annotation of insertions and deletions in a multiple alignment.

MOTIVATION: A quantitative study of molecular evolutionary events such as substitutions, insertions and deletions from closely related genomes requires (1) an accurate multiple sequence alignment program and (2) a method to annotate the insertions and deletions that explain the 'gaps' in the alignment. Although the former requirement has been extensively addressed, the latter problem has received little attention, especially in a comprehensive probabilistic framework. RESULTS: Here, we present Indelign, a program that uses a probabilistic evolutionary model to compute the most likely scenario of insertions and deletions consistent with an input multiple alignment. It is also capable of modifying the given alignment so as to obtain a better agreement with the evolutionary model. We find close to optimal performance and substantial improvement over alternative methods, in tests of Indelign on synthetic data. We use Indelign to analyze regulatory sequences in Drosophila, and find an excess of insertions over deletions, which is different from what has been reported for neutral sequences. AVAILABILITY: The Indelign program may be downloaded from the website http://veda.cs.uiuc.edu/indelign/ SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.

Algorithms↗

Conservation of amino acids in multiple alignments: aspartic acid has unexpected conservation.

Analysis of the relationship between surface accessibility and amino acid conservation in multiple sequence alignments of homologous proteins confirms expected trends for hydrophobic amino acids, but reveals an unexpected difference between the conservation of Asp, Glu and Gln. Even when not in an active site, Asp is more highly conserved than Glu. There is a clear preference for conserved and buried Asp to be present in coil, but there is no tendency for Asp to conserve phi/psi in the ++ region of the Ramachandran map. Glu does not show any preference to be conserved in a particular secondary structure. Analysis of recently derived substitution matrices (e.g. BLOSUM) confirms that Glu tends to substitute more frequently with other amino acids than does Asp. Analysis of relative accessibility versus relative conservation for individual amino acid positions in alignments shows a negative correlation for all amino acid types. With the exception of Arg, Lys, Gly, Glu, Asp and Tyr, a relative conservation of > 2 suggests the amino acid will have a relative accessibility of < 50%. Observation of conserved Cys, Gly or Asp in a reliable multiple alignment suggests a position important for the structure of the protein. Furthermore, the Asp is likely to be involved in polar interactions through its side chain oxygen atoms. In contrast, Gln is the least conserved amino acid overall.

Amino Acid Sequence↗

Visual BLAST and visual FASTA: graphic workbenches for interactive analysis of full BLAST and FASTA outputs under MICROSOFT WINDOWS 95/NT.

MOTIVATION: When routinely analysing protein sequences, detailed analysis of database search results made with BLAST and FASTA becomes exceedingly time consuming and tedious work, as the resultant file may contain a list of hundreds of potential homologies. The interpretation of these results is usually carried out with a text editor which is not a convenient tool for this analysis. In addition, the format of data within BLAST and FASTA output files makes them difficult to read. RESULTS: To facilitate and accelerate this analysis, we present for the first time, two easy-to-use programs designed for interactive analysis of full BLAST and FASTA output files containing protein sequence alignments. The programs, Visual BLAST and Visual FASTA, run under Microsoft Windows 95 or NT systems. They are based on the same intuitive graphical user interface (GUI) with extensive viewing, searching, editing, printing and multithreading capabilities. These programs improve the browsing of BLAST/FASTA results by offering a more convenient presentation of these results. They also implement on a computer several analytical tools which automate a manual methodology used for detailed analysis of BLAST and FASTA outputs. These tools include a pairwise sequence alignment viewer, a Hydrophobic Cluster Analysis plot alignment viewer and a tool displaying a graphical map of all database sequences aligned with the query sequence. In addition. Visual Blast includes tools for multiple sequence alignment analysis (with an amino acid patterns search engine), and Visual FASTA provides a GUI to the FASTA program.

Amino Acid Sequence↗

Addressing the issue of sequence-to-structure alignments in comparative modeling of CASP3 target proteins.

During a blind protein structure prediction experiment (the third round of the Critical Assessment of Techniques for Protein Structure Prediction; URL http://PredictionCenter.llnl.gov/casp3/) , four target proteins, T0047, T0048, T0055, and T0070, were modeled by comparison. These proteins display 62%, 29%, 24%, and 19% sequence identity, respectively, to the structurally homologous proteins most similar in sequence. The issue of sequence-to-structure alignment in cases of low sequence homology was the main emphasis. Selection of alignments was made by constructing and evaluating three-dimensional models based on series of samples produced mainly by automatic multiple sequence alignments. Sequence-to-structure alignments were correct in all but two regions, in which significant changes in target structures compared with related proteins were the source of errors. Template choice is an important determinant of model quality, and a correct selection was made of a lower homology template for modeling of T0070; however, in the case of T0055, a template with 8% greater sequence homology proved deceptive. Loops and some ungapped template regions were assigned conformations taken from other proteins. Using fragments from homologous structures led to improvement over template backbone more often than cases in which nonhomologous structures were the source. The results also indicate that side-chain prediction accuracy depends not only on sequence similarity but also on accuracy of the backbone.

Algorithms↗

Improving protein secondary structure prediction with aligned homologous sequences.

Most recent protein secondary structure prediction methods use sequence alignments to improve the prediction quality. We investigate the relationship between the location of secondary structural elements, gaps, and variable residue positions in multiple sequence alignments. We further investigate how these relationships compare with those found in structurally aligned protein families. We show how such associations may be used to improve the quality of prediction of the secondary structure elements, using the Quadratic-Logistic method with profiles. Furthermore, we analyze the extent to which the number of homologous sequences influences the quality of prediction. The analysis of variable residue positions shows that surprisingly, helical regions exhibit greater variability than do coil regions, which are generally thought to be the most common secondary structure elements in loops. However, the correlation between variability and the presence of helices does not significantly improve prediction quality. Gaps are a distinct signal for coil regions. Increasing the coil propensity for those residues occurring in gap regions enhances the overall prediction quality. Prediction accuracy increases initially with the number of homologues, but changes negligibly as the number of homologues exceeds about 14. The alignment quality affects the prediction more than other factors, hence a careful selection and alignment of even a small number of homologues can lead to significant improvements in prediction accuracy.

Amino Acid Sequence↗

RNA secondary structure prediction based on free energy and phylogenetic analysis.

We describe a computational method for the prediction of RNA secondary structure that uses a combination of free energy and comparative sequence analysis strategies. Using a homology-based sequence alignment as a starting point, all favorable pairings with respect to the Turner energy function are identified. Each potentially paired region within a multiple sequence alignment is scored using a function that combines both predicted free energy and sequence covariation with optimized weightings. High scoring regions are ranked and sequentially incorporated to define a growing secondary structure. Using a single set of optimized parameters, it is possible to accurately predict the foldings of several test RNAs defined previously by extensive phylogenetic and experimental data (including tRNA, 5 S rRNA, SRP RNA, tmRNA, and 16 S rRNA). The algorithm correctly predicts approximately 80% of the secondary structure. A range of parameters have been tested to define the minimal sequence information content required to accurately predict secondary structure and to assess the importance of individual terms in the prediction scheme. This analysis indicates that prediction accuracy most strongly depends upon covariational information and only weakly on the energetic terms. However, relatively few sequences prove sufficient to provide the covariational information required for an accurate prediction. Secondary structures can be accurately defined by alignments with as few as five sequences and predictions improve only moderately with the inclusion of additional sequences.

Algorithms↗

Prediction of protein secondary structure content using amino acid composition and evolutionary information.

Knowing protein structure and inferring its function from the structure are one of the main issues of computational structural biology, and often the first step is studying protein secondary structure. There have been many attempts to predict protein secondary structure contents. Previous attempts assumed that the content of protein secondary structure can be predicted successfully using the information on the amino acid composition of a protein. Recent methods achieved remarkable prediction accuracy by using the expanded composition information. The overall average error of the most successful method is 3.4%. Here, we demonstrate that even if we only use the simple amino acid composition information alone, it is possible to improve the prediction accuracy significantly if the evolutionary information is included. The idea is motivated by the observation that evolutionarily related proteins share the similar structure. After calculating the homolog-averaged amino acid composition of a protein, which can be easily obtained from the multiple sequence alignment by running PSI-BLAST, those 20 numbers are learned by a multiple linear regression, an artificial neural network and a support vector regression. The overall average error of method by a support vector regression is 3.3%. It is remarkable that we obtain the comparable accuracy without utilizing the expanded composition information such as pair-coupled amino acid composition. This work again demonstrates that the amino acid composition is a fundamental characteristic of a protein. It is anticipated that our novel idea can be applied to many areas of protein bioinformatics where the amino acid composition information is utilized, such as subcellular localization prediction, enzyme subclass prediction, domain boundary prediction, signal sequence prediction, and prediction of unfolded segment in a protein sequence, to name a few.

Amino Acid Sequence↗

Consistency of optimal sequence alignments.

Pairwise optimal alignments between three or more sequences are not necessarily consistent as a whole, but consistent and inconsistent residues are usually distributed in clusters. An efficient method has been developed for locating consistent regions when each pairwise alignment is given in the form of a "skeletal representation" (Bull. math. Biol. 52, 359-373). This method is further extended so that the combination of pairwise alignments that gives the greatest consistency is found when possibly many alignments are equally optimal for each pairwise comparison. A method for acceleration of simultaneous multiple sequence alignment is proposed in which consistent regions serve as "anchor points" limiting application of direct multi-way alignment to the rest of "inconsistent" regions.

Algorithms↗

A tool for analyzing and annotating genomic sequences.

We describe a tool for analyzing and annotating large genomic sequences containing introns. The analysis and annotation tool (AAT) includes two sets of programs, one for comparing the query sequence with a protein database and the other for comparing the query with a cDNA database. Each set contains a fast database search program and a rigorous alignment program. The database search program quickly identifies regions of the query sequence that are similar to a database sequence. Then the alignment program constructs an optimal alignment for each region and the database sequence. The alignment program also reports the coordinates of exons in the query sequence. Pairwise alignments of the query sequence with protein and cDNA database sequences are combined into multiple sequence alignments, which provide a view of all protein and cDNA sequences matching a query region. On a data set of 570 DNA sequences, AAT identified 94% of coding nucleotides correctly and 74% of exons exactly. Results of analyzing a human BAC sequence with the AAT tool are also presented. The AAT tool reduces the labor-intensive work of locating the exons of the query sequence and improves the process of defining intron-exon boundaries by using the wealth of available protein and cDNA data.

Amino Acid Sequence↗

Nickel trafficking: insights into the fold and function of UreE, a urease metallochaperone.

UreE is a metallo-chaperone assisting the incorporation of two adjacent Ni(2+) ions in the active site of urease. This study describes an attempt to distill general information on this protein using a computational post-genomic approach for the understanding of the structural details of the molecular function of UreE in nickel trafficking. The two crystal structures recently determined for UreE from Bacillus pasteurii (BpUreE) and Klebsiella aerogenes (KaUreE) were comparatively analyzed. This analysis provided insights into the protein structural and conformational features. A structural database of UreE proteins from a large number of different genomes was built using homology modeling. All available sequences of UreE were retrieved from protein and cDNA databases, and their structures were modeled on the crystal structures of BpUreE and KaUreE. A self-consistent iterative protocol was devised for multiple sequence alignment optimization involving secondary structure prediction and evaluation of the energy features of the obtained modeled structures. The quality of all models was tested using standard assessment procedures. The final optimized structure-based multiple alignment and the derived model structures provided insightful information on the evolutionary conservation of key residues in the protein sequence and surface patches presumably involved in protein recognition during the urease active site assembly.

Amino Acid Sequence↗

Phylogenetic relationships of three porcine mycoplasmas, Mycoplasma hyopneumoniae, Mycoplasma flocculare, and Mycoplasma hyorhinis, and complete 16S rRNA sequence of M. flocculare.

The nucleotide sequence of the 16S rRNA gene of Mycoplasma flocculare was determined and was compared with the sequence of a related porcine mycoplasma, Mycoplasma hyopneumoniae. While the overall level of DNA-DNA homology was approximately 11%, sequence alignment of the two 16S rRNA genes yielded a homology value of more than 95%, emphasizing the highly conserved nature of the 16S rRNA gene. Multiple sequence alignments with other mollicutes indicated that M. flocculare, M. hyopneumoniae, and Mycoplasma hyorhinis form a subcluster within the fermentans phylogroup, and this subcluster is distinct from the Mycoplasma pneumoniae phylogroup. Thus, the three mycoplasmas isolated from porcine respiratory systems exhibit phylogenetic similarities.

Animals↗

MUSTANG: a multiple structural alignment algorithm.

Multiple structural alignment is a fundamental problem in structural genomics. In this article, we define a reliable and robust algorithm, MUSTANG (MUltiple STructural AligNment AlGorithm), for the alignment of multiple protein structures. Given a set of protein structures, the program constructs a multiple alignment using the spatial information of the C(alpha) atoms in the set. Broadly based on the progressive pairwise heuristic, this algorithm gains accuracy through novel and effective refinement phases. MUSTANG reports the multiple sequence alignment and the corresponding superposition of structures. Alignments generated by MUSTANG are compared with several handcurated alignments in the literature as well as with the benchmark alignments of 1033 alignment families from the HOMSTRAD database. The performance of MUSTANG was compared with DALI at a pairwise level, and with other multiple structural alignment tools such as POSA, CE-MC, MALECON, and MultiProt. MUSTANG performs comparably to popular pairwise and multiple structural alignment tools for closely related proteins, and performs more reliably than other multiple structural alignment methods on hard data sets containing distantly related proteins or proteins that show conformational changes.

Algorithms↗

Human-specific insertions and deletions inferred from mammalian genome sequences.

It has been suggested that insertions and deletions (indels) have contributed to the sequence divergence between the human and chimpanzee genomes more than do nucleotide changes (3% vs. 1.2%). However, although there have been studies of large indels between the two genomes, no systematic analysis of small indels (i.e., indels </= 100 bp) has been published. In this study, we first estimated that the false-positive rate of small indels inferred from human-chimpanzee pairwise sequence alignments is quite high, suggesting that the chimpanzee genome draft is not sufficiently accurate for our purpose. We have therefore inferred only human-specific indels using multiple sequence alignments of mammalian genomes. We identified >840,000 "small" indels, which affect >7000 UCSC-annotated human genes (>11,000 transcripts). These indels, however, amount to only approximately 0.21% sequence change in the human lineage for the regions compared, whereas in pseudogenes indels contribute to a sequence divergence of 1.40%, suggesting that most of the indels that occurred in genic regions have been eliminated. Functional analysis reveals that the genes whose coding exons have been affected by human-specific indels are enriched in transcription and translation regulatory activities but are underrepresented in catalytic and transporter activities, cellular and physiological processes, and extracellular region/matrix. This functional bias suggests that human-specific indels might have contributed to human unique traits by causing changes at the RNA and protein level.

Animals↗