Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

A configuration space of homologous proteins conserving mutual information and allowing a phylogeny inference based on pair-wise Z-score probabilities.

BACKGROUND: Popular methods to reconstruct molecular phylogenies are based on multiple sequence alignments, in which addition or removal of data may change the resulting tree topology. We have sought a representation of homologous proteins that would conserve the information of pair-wise sequence alignments, respect probabilistic properties of Z-scores (Monte Carlo methods applied to pair-wise comparisons) and be the basis for a novel method of consistent and stable phylogenetic reconstruction. RESULTS: We have built up a spatial representation of protein sequences using concepts from particle physics (configuration space) and respecting a frame of constraints deduced from pair-wise alignment score properties in information theory. The obtained configuration space of homologous proteins (CSHP) allows the representation of real and shuffled sequences, and thereupon an expression of the TULIP theorem for Z-score probabilities. Based on the CSHP, we propose a phylogeny reconstruction using Z-scores. Deduced trees, called TULIP trees, are consistent with multiple-alignment based trees. Furthermore, the TULIP tree reconstruction method provides a solution for some previously reported incongruent results, such as the apicomplexan enolase phylogeny. CONCLUSION: The CSHP is a unified model that conserves mutual information between proteins in the way physical models conserve energy. Applications include the reconstruction of evolutionary consistent and robust trees, the topology of which is based on a spatial representation that is not reordered after addition or removal of sequences. The CSHP and its assigned phylogenetic topology, provide a powerful and easily updated representation for massive pair-wise genome comparisons based on Z-score computations.

Algorithms↗

Consensus folding of aligned sequences as a new measure for the detection of functional RNAs by comparative genomics.

Facing the ever-growing list of newly discovered classes of functional RNAs, it can be expected that further types of functional RNAs are still hidden in recently completed genomes. The computational identification of such RNA genes is, therefore, of major importance. While most known functional RNAs have characteristic secondary structures, their free energies are generally not statistically significant enough to distinguish RNA genes from the genomic background. Additional information is required. Considering the wide availability of new genomic data of closely related species, comparative studies seem to be the most promising approach. Here, we show that prediction of consensus structures of aligned sequences can be a significant measure to detect functional RNAs. We report a new method to test multiple sequence alignments for the existence of an unusually structured and conserved fold. We show for alignments of six types of well-known functional RNA that an energy score consisting of free energy and a covariation term significantly improves sensitivity compared to single sequence predictions. We further test our method on a number of non-coding RNAs from Caenorhabditis elegans/Caenorhabditis briggsae and seven Saccharomyces species. Most RNAs can be detected with high significance. We provide a Perl implementation that can be used readily to score single alignments and discuss how the methods described here can be extended to allow for efficient genome-wide screens.

Algorithms↗

A simple method for aligning many protein sequences.

A simple extension of the Needleman and Wunsch algorithm for aligning pairs of protein sequences allows it to be used for the efficient generation of very large multiple-sequence alignments whose members are similar. This technique could have applications in a broad range of high-volume genomics projects.

Algorithms↗

Membrane-associated proteins in eicosanoid and glutathione metabolism (MAPEG). A widespread protein superfamily.

The members of the MAPEG superfamily have been aligned and found to be distantly related, with a common pattern of hydropathy. Figure 2A shows the multiple sequence alignments of the human members and Figure 2B the corresponding superimposed hydropathy profiles. The alignment in Figure 2A demonstrates a total of six strictly conserved residues. The Arg-51 in LTC4 synthase has been suggested to function as proton donor for the opening of the LTA4 epoxide. This arginine is found in all but the FLAP sequences in accordance with the observation that FLAP has no known enzyme activity. Also the Tyr-93 in LTC4 synthase has been suggested to function as a base for the formation of the thiolate anion of glutathione. This tyrosine is not conserved in MGST1 or MGST1-L1. Table 1 summarizes some other properties of the individual human proteins. They are all of the same size, ranging from 147 to 161 amino acids. Only FLAP differs in that its isoelectric point is more neutral than that of the other, more basic proteins. The genes encoding these proteins all reside on different chromosomes (when known) (Table 1). In addition to the human proteins, MAPEG members have been identified in plants, fungi, and bacteria. It is clearly a challenge to elucidate their role in these different phyla in relation to their defined physiological functions in humans.

Amino Acid Sequence↗

Improvements to GALA and dbERGE II: databases featuring genomic sequence alignment, annotation and experimental results.

We describe improvements to two databases that give access to information on genomic sequence similarities, functional elements in DNA and experimental results that demonstrate those functions. GALA, the database of Genome ALignments and Annotations, is now a set of interlinked relational databases for five vertebrate species, human, chimpanzee, mouse, rat and chicken. For each species, GALA records pairwise and multiple sequence alignments, scores derived from those alignments that reflect the likelihood of being under purifying selection or being a regulatory element, and extensive annotations such as genes, gene expression patterns and transcription factor binding sites. The user interface supports simple and complex queries, including operations such as subtraction and intersections as well as clustering and finding elements in proximity to features. dbERGE II, the database of Experimental Results on Gene Expression, contains experimental data from a variety of functional assays. Both databases are now run on the DB2 database management system. Improved hardware and tuning has reduced response times and increased querying capacity, while simplified query interfaces will help direct new users through the querying process. Links are available at http://www.bx.psu.edu/.

Animals↗

Panta rhei (QAlign2): an open graphical environment for sequence analysis.

MOTIVATION: The first version of the graphical multiple sequence alignment environment QAlign was published in 2003. Heavy response from the molecular-biological user community clearly demonstrated the need for such a platform. RESULTS: Panta rhei extends QAlign by several features. Major redesigns on the user interface, for instance, allow users to flexibily compose views for multiple projects. The new sequence viewer handles datasets with arbitrarily many and arbitrarily large sequences that may still be edited by guided block moving. More distance-based algorithms are available to interactively reconstruct phylogenetic trees which can now also be zoomed and navigated graphicaly. AVAILABILITY: Executables and the JAVA source code are available under the Apache license at http://gi.cebitec.uni-bielefeld.de/qalign CONTACT: qalign@cebitec.uni-bielefeld.de.

Algorithms↗

Detecting recombination with MCMC.

MOTIVATION: We present a statistical method for detecting recombination, whose objective is to accurately locate the recombinant breakpoints in DNA sequence alignments of small numbers of taxa (4 or 5). Our approach explicitly models the sequence of phylogenetic tree topologies along a multiple sequence alignment. Inference under this model is done in a Bayesian way, using Markov chain Monte Carlo (MCMC). The algorithm returns the site-dependent posterior probability of each tree topology, which is used for detecting recombinant regions and locating their breakpoints. RESULTS: The method was tested on a synthetic and three real DNA sequence alignments, where it was found to outperform the established detection methods PLATO, RECPARS, and TOPAL.

Algorithms↗

Computational methods for the analysis of differential conservation in groups of similar DNA sequences.

Multiple sequence alignments are a powerful tool for identifying the regions of DNA which have been constrained in evolutionary divergence, presumably due to their functional role. However, such constraints rarely manifest themselves as perfect conservation of a site clearly standing out in its broader environment, as they reflect the species-specific differences in proteins, as well as the ability of some proteins to interact with multiple variants of their binding sequence. In this paper we explore the use of alignment column uncertainty as an aid in locating differential phylogenetic footprints, which refer to the sites in DNA where groups of related species exhibit sequence conservation, but where the pattern may vary between the groups. We use efficient, linear-time algorithms to locate such sites. We have performed a study of the mammalian CAV2-CAV1 gene region using our software, and we conclude with several observations concerning the differential conservation and the use of computational methods for its detection. The software developed for this project is available, free of charge, by contacting the author.

Animals↗

Suboptimal sequence alignment in molecular biology. Alignment with error analysis.

A molecular sequence alignment algorithm based on dynamic programming has been extended to allow the computation of all pairs of residues that can be part of optimal and suboptimal sequence alignments. The uncertainties inherent in sequence alignment can be displayed using a new form of dot plot. The method allows the qualitative assessment of whether or not two sequences are related, and can reveal what parts of the alignment are better determined than others. It also permits the computation of representative optimal and suboptimal alignments. The relation between alignment reliability and alignment parameters is discussed. Other applications are to cyclical permutations of sequences and the detection of self-similarity. An application to multiple sequence alignment is noted.

Algorithms↗

topoSNP: a topographic database of non-synonymous single nucleotide polymorphisms with and without known disease association.

The database of topographic mapping of Single Nucleotide Polymorphism (topoSNP) provides an online resource for analyzing non-synonymous SNPs (nsSNPs) that can be mapped onto known 3D structures of proteins. These include disease- associated nsSNPs derived from the Online Mendelian Inheritance in Man (OMIM) database and other nsSNPs derived from dbSNP, a resource at the National Center for Biotechnology Information that catalogs SNPs. TopoSNP further classifies each nsSNP site into three categories based on their geometric location: those located in a surface pocket or an interior void of the protein, those on a convex region or a shallow depressed region, and those that are completely buried in the interior of the protein structure. These unique geometric descriptions provide more detailed mapping of nsSNPs to protein structures. The current release also includes relative entropy of SNPs calculated from multiple sequence alignment as obtained from the Pfam database (a database of protein families and conserved protein motifs) as well as manually adjusted multiple alignments obtained from ClustalW. These structural and conservational data can be useful for studying whether nsSNPs in coding regions are likely to lead to phenotypic changes. TopoSNP includes an interactive structural visualization web interface, as well as downloadable batch data. The database will be updated at regular intervals and can be accessed at: http://gila.bioengr.uic.edu/snp/toposnp.

Computational Biology↗

Pollen allergens are restricted to few protein families and show distinct patterns of species distribution.

BACKGROUND: Inhalative allergies are elicited predominantly by pollen of various plant species. However, a classification of the large number of identified pollen allergens is still missing. OBJECTIVE: To analyze pollen allergen sequences with respect to protein family membership, taxonomic distribution of protein families, and interspecies variability. METHODS: Protein family memberships of all plant allergen sequences from the Allergome database were determined by using the Protein Families Database of Alignments and Hidden Markov Models. The taxonomic distribution of pollen allergens was established from the Integrated Taxonomic Information System. Members of abundant pollen allergen families were compared with allergenic and nonallergenic homologues by database similarity searches and multiple sequence alignments. RESULTS: Pollen allergens were classified into 29 of 7868 protein families. Expansins, profilins, and calcium-binding proteins constitute the major pollen allergen families, whereas most plant food allergens belong to the prolamin, cupin, or profilin families. Pollen allergens were revealed to be ubiquitous (eg, profilins), present in certain plant families (eg, pectate lyases), or limited to a single taxon (eg, thaumatin-like proteins). Allergenic plant profilins constitute a highly conserved family with sequence identities of 70% to 85% among each other but low identities of 30% to 40% with nonallergenic profilins from other eukaryotes, including human beings. Similarly, allergenic polcalcins possess sequence identities of 64% to 92% but show low identities of 39% to 42% to related nonallergenic calmodulins and calmodulin-like proteins from vegetative plant tissues and man. CONCLUSION: This classification of pollen allergens into protein families will aid in predicting cross-reactivity, designing comprehensive diagnostic devices, and assessing the allergenic potential of novel proteins.

Allergens↗

Conservation analysis and structure prediction of the protein serine/threonine phosphatases. Sequence similarity with diadenosine tetraphosphatase from Escherichia coli suggests homology to the protein phosphatases.

A multiple sequence alignment of 44 serine/threonine-specific protein phosphatases has been performed. This reveals the position of a common conserved catalytic core, the location of invariant residues, insertions and deletions. The multiple alignment has been used to guide and improve a consensus secondary-structure prediction for the common catalytic core. The location of insertions and deletions has aided in defining the positions of surface loops and turns. The prediction suggests that the core protein phosphatase structure comprises two domains: the first has a single, beta sheet flanked by alpha helices, while the second is predominantly alpha helical. Knowledge of the core secondary structures provides a guide for the design of site-directed-mutagenesis experiments that will not disrupt the native phosphatase fold. A sequence similarity between eukaryotic serine/threonine protein phosphatases and the Escherichia coli diadenosine tetraphosphatase has been identified. This extends over the N-terminal 100 residues of bacteriophage phosphatases and E. coli diadenosine tetraphosphatase. Residues which are invariant amongst these classes are likely to be important in catalysis and protein folding. These include Arg92, Asn138, Asp59, Asp88, Gly58, Gly62, Gly87, Gly93, Gly137, His61, His139 and Val90 and fall into three clusters with the consensus sequences GD(IVTL)HG, GD(LYF)V(DA)RG and GNH, where brackets surround alternative amino acids. The first two consensus sequences are predicted to fall in the beta-alpha and beta-beta loops of a beta-alpha-beta-beta secondary-structure motif. This places the predicted phosphate-binding site at the N-terminus of the alpha helix, where phosphate binding may be stabilised by the alpha-helix dipole.

Acid Anhydride Hydrolases↗

Structural interpretation of mutations and SNPs using STRAP-NT.

Visualization of residue positions in protein alignments and mapping onto suitable structural models is an important first step in the interpretation of mutations or polymorphisms in terms of protein function, interaction, and thermodynamic stability. Selecting and highlighting large numbers of residue positions in a protein structure can be time-consuming and tedious with currently available software. Previously, a series of tasks and analyses had to be performed one-by-one to map mutations onto 3D protein structures; STRAP-NT is an extension of STRAP that automates these tasks so that users can quickly and conveniently map mutations onto 3D protein structures. When the structure of the protein of interest is not yet available, a related protein can frequently be found in the structure databases. In this case the alignment of both proteins becomes the crucial part of the analysis. Therefore we embedded these program modules into the Java-based multiple sequence alignment program STRAP-NT. STRAP-NT can simultaneously map an arbitrary number of mutations denoted using either the nucleotide or amino acid sequence. When the designations of the mutations refer to genomic sites, STRAP-NT translates them into the corresponding amino acid positions, taking intron-exon boundaries into account. STRAP-NT tightly integrates a number of current protein structure viewers (currently PYMOL, RASMOL, JMOL, and VMD) with which mutations and polymorphisms can be directly displayed on the 3D protein structure model. STRAP-NT is available at the PDB site and at http://www.charite.de/bioinf/strap/ or http://strapjava.de.

DNA Mutational Analysis↗

OrthoMCL-DB: querying a comprehensive multi-species collection of ortholog groups.

The OrthoMCL database (http://orthomcl.cbil.upenn.edu) houses ortholog group predictions for 55 species, including 16 bacterial and 4 archaeal genomes representing phylogenetically diverse lineages, and most currently available complete eukaryotic genomes: 24 unikonts (12 animals, 9 fungi, microsporidium, Dictyostelium, Entamoeba), 4 plants/algae and 7 apicomplexan parasites. OrthoMCL software was used to cluster proteins based on sequence similarity, using an all-against-all BLAST search of each species' proteome, followed by normalization of inter-species differences, and Markov clustering. A total of 511,797 proteins (81.6% of the total dataset) were clustered into 70,388 ortholog groups. The ortholog database may be queried based on protein or group accession numbers, keyword descriptions or BLAST similarity. Ortholog groups exhibiting specific phyletic patterns may also be identified, using either a graphical interface or a text-based Phyletic Pattern Expression grammar. Information for ortholog groups includes the phyletic profile, the list of member proteins and a multiple sequence alignment, a statistical summary and graphical view of similarities, and a graphical representation of domain architecture. OrthoMCL software, the entire FASTA dataset employed and clustering results are available for download. OrthoMCL-DB provides a centralized warehouse for orthology prediction among multiple species, and will be updated and expanded as additional genome sequence data become available.

Animals↗

New features of the Blocks Database servers.

Blocks are ungapped multiple sequence alignments representing conserved protein regions, and the Blocks Database consists of blocks from documented protein families. World Wide Web (http://www. blocks.fhcrc.org) and Email (blocks@blocks.fhcrc.org) servers provide tools for homology searching and for analyzing protein family relationships. New enhancements include a multiple alignment processor that extends the use of these tools to imported multiple alignments of families not present in the database and a PCR primer designer that implements a new strategy for gene isolation.

DNA Primers↗

A novel approach to phylogeny reconstruction from protein sequences.

The reliable reconstruction of tree topology from a set of homologous sequences is one of the main goals in the study of molecular evolution. If consistent estimators of distances from a multiple sequence alignment are known, the distance method is attractive because the tree reconstruction is consistent. To obtain a distance estimate d, the observed proportion of differences p (p-distance) is usually "corrected" for multiple and back substitutions by means of a functional relationship d = f(p). In this paper the conditions under which this correction of p-distances will not alter the selection of the tree topology are specified. When these conditions are not fulfilled the selection of the tree topology may depend on the correction function applied. A novel method which includes estimates of distances not only between sequence pairs, but between triplets, quadruplets, etc., is proposed to strengthen the proper selection of correction function and tree topology. A "super" tree that includes all tree topologies as special cases is introduced.

Animals↗

Multiple protein structure alignment.

A method was developed to compare protein structures and to combine them into a multiple structure consensus. Previous methods of multiple structure comparison have only concatenated pairwise alignments or produced a consensus structure by averaging coordinate sets. The current method is a fusion of the fast structure comparison program SSAP and the multiple sequence alignment program MULTAL. As in MULTAL, structures are progressively combined, producing intermediate consensus structures that are compared directly to each other and all remaining single structures. This leads to a hierarchic "condensation," continually evaluated in the light of the emerging conserved core regions. Following the SSAP approach, all interatomic vectors were retained with well-conserved regions distinguished by coherent vector bundles (the structural equivalent of a conserved sequence position). Each bundle of vectors is summarized by a resultant, whereas vector coherence is captured in an error term, which is the only distinction between conserved and variable positions. Resultant vectors are used directly in the comparison, which is weighted by their error values, giving greater importance to the matching of conserved positions. The resultant vectors and their errors can also be used directly in molecular modeling. Applications of the method were assessed by the quality of the resulting sequence alignments, phylogenetic tree construction, and databank scanning with the consensus. Visual assessment of the structural superpositions and consensus structure for various well-characterized families confirmed that the consensus had identified a reasonable core.

Amino Acid Sequence↗

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural (3-D) and sequence (1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in SwissProt using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 29% of all SwissProt-stored sequences.

Amino Acid Sequence↗