Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Profile hidden Markov models.

The recent literature on profile hidden Markov model (profile HMM) methods and software is reviewed. Profile HMMs turn a multiple sequence alignment into a position-specific scoring system suitable for searching databases for remotely homologous sequences. Profile HMM analyses complement standard pairwise comparison methods for large-scale sequence analysis. Several software implementations and two large libraries of profile HMMs of common protein domains are available. HMM methods performed comparably to threading methods in the CASP2 structure prediction exercise.

Humans↗

Application of distance geometry to 3D visualization of sequence relationships.

SUMMARY: We describe the application of distance geometry methods to the three-dimensional visualization of sequence relationships, with examples for mumps virus SH gene cDNA and prion protein sequences. Sequence-sequence distance measures may be obtained from either a multiple sequence alignment or from sets of pairwise alignments. AVAILABILITY: C/Perl code and HTML/VRML files from http://www.nibsc.ac.uk/dg3dseq/

Algorithms↗

Markovian domain fingerprinting: statistical segmentation of protein sequences.

MOTIVATION: Characterization of a protein family by its distinct sequence domains is crucial for functional annotation and correct classification of newly discovered proteins. Conventional Multiple Sequence Alignment (MSA) based methods find difficulties when faced with heterogeneous groups of proteins. However, even many families of proteins that do share a common domain contain instances of several other domains, without any common underlying linear ordering. Ignoring this modularity may lead to poor or even false classification results. An automated method that can analyze a group of proteins into the sequence domains it contains is therefore highly desirable. RESULTS: We apply a novel method to the problem of protein domain detection. The method takes as input an unaligned group of protein sequences. It segments them and clusters the segments into groups sharing the same underlying statistics. A Variable Memory Markov (VMM) model is built using a Prediction Suffix Tree (PST) data structure for each group of segments. Refinement is achieved by letting the PSTs compete over the segments, and a deterministic annealing framework infers the number of underlying PST models while avoiding many inferior solutions. We show that regions of similar statistics correlate well with protein sequence domains, by matching a unique signature to each domain. This is done in a fully automated manner, and does not require or attempt an MSA. Several representative cases are analyzed. We identify a protein fusion event, refine an HMM superfamily classification into the underlying families the HMM cannot separate, and detect all 12 instances of a short domain in a group of 396 sequences. CONTACT: jill@cs.huji.ac.il; tishby@cs.huji.ac.il.

Algorithms↗

PhyloBLAST: facilitating phylogenetic analysis of BLAST results.

PhyloBLAST is an internet-accessed application based on CGI/Perl programming that compares a users protein sequence to a SwissProt/TREMBL database using BLAST2 and then allows phylogenetic analyses to be performed on selected sequences from the BLAST output. Flexible features such as ability to input your own multiple sequence alignment and use PHYLIP program options provide additional web-based phylogenetic analysis functionality beyond the analysis of a BLAST result.

Internet↗

Euclidian space and grouping of biological objects.

MOTIVATION: Biological objects tend to cluster into discrete groups. Objects within a group typically possess similar properties. It is important to have fast and efficient tools for grouping objects that result in biologically meaningful clusters. Protein sequences reflect biological diversity and offer an extraordinary variety of objects for polishing clustering strategies. Grouping of sequences should reflect their evolutionary history and their functional properties. Visualization of relationships between sequences is of no less importance. Tree-building methods are typically used for such visualization. An alternative concept to visualization is a multidimensional sequence space. In this space, proteins are defined as points and distances between the points reflect the relationships between the proteins. Such a space can also be a basis for model-based clustering strategies that typically produce results correlating better with biological properties of proteins. RESULTS: We developed an approach to classification of biological objects that combines evolutionary measures of their similarity with a model-based clustering procedure. We apply the methodology to amino acid sequences. On the first step, given a multiple sequence alignment, we estimate evolutionary distances between proteins measured in expected numbers of amino acid substitutions per site. These distances are additive and are suitable for evolutionary tree reconstruction. On the second step, we find the best fit approximation of the evolutionary distances by Euclidian distances and thus represent each protein by a point in a multidimensional space. The Euclidian space may be projected in two or three dimensions and the projections can be used to visualize relationships between proteins. On the third step, we find a non-parametric estimate of the probability density of the points and cluster the points that belong to the same local maximum of this density in a group. The number of groups is controlled by a sigma-parameter that determines the shape of the density estimate and the number of maxima in it. The grouping procedure outperforms commonly used methods such as UPGMA and single linkage clustering.

Algorithms↗

A sequence-profile-based HMM for predicting and discriminating beta barrel membrane proteins.

MOTIVATION: Membrane proteins are an abundant and functionally relevant subset of proteins that putatively include from about 15 up to 30% of the proteome of organisms fully sequenced. These estimates are mainly computed on the basis of sequence comparison and membrane protein prediction. It is therefore urgent to develop methods capable of selecting membrane proteins especially in the case of outer membrane proteins, barely taken into consideration when proteome wide analysis is performed. This will also help protein annotation when no homologous sequence is found in the database. Outer membrane proteins solved so far at atomic resolution interact with the external membrane of bacteria with a characteristic beta barrel structure comprising different even numbers of beta strands (beta barrel membrane proteins). In this they differ from the membrane proteins of the cytoplasmic membrane endowed with alpha helix bundles (all alpha membrane proteins) and need specialised predictors. RESULTS: We develop a HMM model, which can predict the topology of beta barrel membrane proteins using, as input, evolutionary information. The model is cyclic with 6 types of states: two for the beta strand transmembrane core, one for the beta strand cap on either side of the membrane, one for the inner loop, one for the outer loop and one for the globular domain state in the middle of each loop. The development of a specific input for HMM based on multiple sequence alignment is novel. The accuracy per residue of the model is 83% when a jack knife procedure is adopted. With a model optimisation method using a dynamic programming algorithm seven topological models out of the twelve proteins included in the testing set are also correctly predicted. When used as a discriminator, the model is rather selective. At a fixed probability value, it retains 84% of a non-redundant set comprising 145 sequences of well-annotated outer membrane proteins. Concomitantly, it correctly rejects 90% of a set of globular proteins including about 1200 chains with low sequence identity (<30%) and 90% of a set of all alpha membrane proteins, including 188 chains.

Algorithms↗

Comparative genomics of microbial pathogens and symbionts.

We are interested in quantifying the contribution of gene acquisition, loss, expansion and rearrangements to the evolution of microbial genomes. Here, we discuss factors influencing microbial genome divergence based on pair-wise genome comparisons of closely related strains and species with different lifestyles. A particular focus is on intracellular pathogens and symbionts of the genera Rickettsia, Bartonella and BUCHNERA: Extensive gene loss and restricted access to phage and plasmid pools may provide an explanation for why single host pathogens are normally less successful than multihost pathogens. We note that species-specific genes tend to be shorter than orthologous genes, suggesting that a fraction of these may represent fossil-orfs, as also supported by multiple sequence alignments among species. The results of our genome comparisons are placed in the context of phylogenomic analyses of alpha and gamma proteobacteria. We highlight artefacts caused by different rates and patterns of mutations, suggesting that atypical phylogenetic placements can not a priori be taken as evidence for horizontal gene transfer events. The flexibility in genome structure among free-living microbes contrasts with the extreme stability observed for the small genomes of aphid endosymbionts, in which no rearrangements or inflow of genetic material have occurred during the past 50 millions years (1). Taken together, the results suggest that genomic stability correlate with the content of repeated sequences and mobile genetic elements, and thereby indirectly with bacterial lifestyles.

Alphaproteobacteria↗

MAVG: locating non-overlapping maximum average segments in a given sequence.

SUMMARY: MAVG is a software tool for finding k non-overlapping maximum-average segments that are sufficiently long in a given sequence of real numbers, for any k > 0. It has applications in several areas of biomolecular sequence analysis including locating GC-rich regions and CpG islands in a genomic sequence, and annotating multiple sequence alignments. AVAILABILITY: http://iubio.bio.indiana.edu/soft/molbio/pattern/cpg_islands/.

Algorithms↗

ModView, visualization of multiple protein sequences and structures.

SUMMARY: We describe ModView, a web application for visualization of multiple protein sequences and structures. ModView integrates a multiple structure viewer, a multiple sequence alignment editor, and a database querying engine. It is possible to interactively manipulate hundreds of proteins, to visualize conservative and variable residues, active and binding sites, fragments, and domains in protein families, as well as to display large macromolecular complexes such as ribosomes or viruses. As a Netscape plug-in, ModView can be included in HTML pages along with text and figures, which makes it useful for teaching and presentations. ModView is also suitable as a graphical interface to various databases because it can be controlled through JavaScript commands and called from CGI scripts. AVAILABILITY: ModView is available at http://guitar.rockefeller.edu/modview.

Database Management Systems↗

ENVIRON: a software package to compare protein three-dimensional structures with homologous sequences using local structural motifs.

This work presents a method to compare local clusters of interacting residues as observed in a known three-dimensional protein structure with corresponding clusters inferred from homologous protein sequences, assuming conserved protein folding. For this purpose the local environment of a selected residue in a known protein structure is defined as the ensemble of amino acids in contact with it in the folded state. Using a multiple sequence alignment to identify corresponding residues in homologous proteins, a detailed comparison can be performed between the local environment of a selected amino acid in the template protein structure and the expected local environments at the sets of equivalent residues, derived from the aligned protein sequences. The comparison makes it possible to detect conserved local features such as hydrogen bonding or complementarity in residue substitution. A global measure of environmental similarity is also defined, to search for conserved amino acid clusters subject to functional or structural constraints. The proposed approach is useful for investigating protein function as well as for site-directed mutagenesis experiments, where appropriate amino acid substitutions can be suggested by observing naturally occurring protein variants.

Algorithms↗

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny↗

SNP and mutation discovery using base-specific cleavage and MALDI-TOF mass spectrometry.

MOTIVATION: Single Nucleotide Polymorphisms (SNPs) are believed to contribute strongly to the genetic variability in living beings, in particular their disease or drug side effect predispositions. Mutation-induced sequence variations are playing an important role in the development of cancer, among others. From this, it is clear that SNP and mutation discovery is of great interest in today's Life Sciences. Currently, such discovery is often performed utilizing electrophoresis-based Sanger Sequencing. Discovery of SNPs can also be performed by multiple sequence alignment of publicly available sequence data, but recent studies indicate that only a small percentage of SNPs can be discovered using this approach and, in particular, that SNPs with low frequency are often missed. Other SNP discovery methods only indicate the presence of a SNP in a sample region, but fail to resolve its characterization and localization. RESULTS: We present a method to discover mutations and SNPs using base-specific cleavage and mass spectrometry. An amplicon of known reference sequence with length usually between 100 and 1000 nt is amplified, transcribed, and cleaved using base-specific endonucleases such as RNAse A or T1. The resulting cleavage products (or fragments) are analyzed by MALDI-TOF mass spectrometry and, comparing the measured spectra with those predicted in-silico, the goal is to discover and pinpoint sequence variations of the sample sequence compared to the reference sequence. A time-efficient algorithm for discovering sequence variations is presented that enables fast analysis of such variations even if the sample sequence differs significantly from the reference sequence.

Algorithms↗

Pyranose oxidase identified as a member of the GMC oxidoreductase family.

Fungal pyranose oxidase is a flavoenzyme whose preferred substrate among several monosaccharides is D-glucose. After a comprehensive analysis of conserved features in a structure-based multiple sequence alignment of homologous proteins, we could classify this enzyme into the GMC oxidoreductase family. The identified homology also suggests a three-dimensional protein structure similar to the functionally related glucose oxidase.

Amino Acid Sequence↗

MutDB: annotating human variation with functionally relevant data.

SUMMARY: We have developed a resource, MutDB (http://mutdb.org/), to aid in determining which single nucleotide polymorphisms (SNPs) are likely to alter the function of their associated protein product. MutDB contains protein structure annotations and comparative genomic annotations for 8000 disease-associated mutations and SNPs found in the UCSC Annotated Genome and the human RefSeq gene set. MutDB provides interactive mutation maps at the gene and protein levels, and allows for ranking of their predicted functional consequences based on conservation in multiple sequence alignments. AVAILABILITY: http://mutdb.org/ SUPPLEMENTARY INFORMATION: http://mutdb.org/about/about.html

Database Management Systems↗

Disease-associated variants in PYPAF1 and NOD2 result in similar alterations of conserved sequence.

Sequence variations in the gene products PYPAF1/CIAS1 and NOD2/CARD15 have been associated with several autoinflammatory diseases that, although clinically different, share a similar inflammatory pathophysiology. A multiple sequence alignment of homologous proteins demonstrates that some of the missense variants are located in highly conserved regions of the NTPase domain and possibly impair NTP-hydrolysis. Intriguingly, one of the variations, which is found identically in PYPAF1 and NOD2, is located at the same alignment position. Our findings suggest that evolutionary gene duplication can give rise to disease families because variants affect conserved sequence in a similar fashion.

Animals↗

Predicting allergenic proteins using wavelet transform.

MOTIVATION: With many transgenic proteins introduced today, the ability to predict their potential allergenicity has become an important issue. Previous studies were based on either sequence similarity or the protein motifs identified from known allergen databases. The similarity-based approaches, although being able to produce high recalls, usually have low prediction precisions. Previous motif-based approaches have been shown to be able to improve the precisions on cross-validation experiments. In this study, a system that combines the advantages of similarity-based and motif-based prediction is described. RESULTS: The new prediction system uses a clustering algorithm that groups the known allergenic proteins into clusters. Proteins within each cluster are assumed to carry one or more common motifs. After a multiple sequence alignment, proteins in each cluster go through a wavelet analysis program whereby conserved motifs will be identified. A hidden Markov model (HMM) profile will then be prepared for each identified motif. The allergens that do not appear to carry detectable allergen motifs will be saved in a small database. The allergenicity of an unknown protein may be predicted by comparing it against the HMM profiles, and, if no matching profiles are found, against the small allergen database by BLASTP. Over 70% of recall and over 90% of precision were observed using cross-validation experiments. Using the entire Swiss-Prot as the query, we predicted about 2000 potential allergens. AVAILABILITY: The software is available upon request from the authors.

Algorithms↗

GLAD: a system for developing and deploying large-scale bioinformatics grid.

MOTIVATION: Grid computing is used to solve large-scale bioinformatics problems with gigabytes database by distributing the computation across multiple platforms. Until now in developing bioinformatics grid applications, it is extremely tedious to design and implement the component algorithms and parallelization techniques for different classes of problems, and to access remotely located sequence database files of varying formats across the grid. In this study, we propose a grid programming toolkit, GLAD (Grid Life sciences Applications Developer), which facilitates the development and deployment of bioinformatics applications on a grid. RESULTS: GLAD has been developed using ALiCE (Adaptive scaLable Internet-based Computing Engine), a Java-based grid middleware, which exploits the task-based parallelism. Two bioinformatics benchmark applications, such as distributed sequence comparison and distributed progressive multiple sequence alignment, have been developed using GLAD.

Computational Biology↗

An alternative model of amino acid replacement.

MOTIVATION: The observed correlations between pairs of homologous protein sequences are typically explained in terms of a Markovian dynamic of amino acid substitution. This model assumes that every location on the protein sequence has the same background distribution of amino acids, an assumption that is incompatible with the observed heterogeneity of protein amino acid profiles and with the success of profile multiple sequence alignment. RESULTS: We propose an alternative model of amino acid replacement during protein evolution based upon the assumption that the variation of the amino acid background distribution from one residue to the next is sufficient to explain the observed sequence correlations of homologs. The resulting dynamical model of independent replacements drawn from heterogeneous backgrounds is simple and consistent, and provides a unified homology match score for sequence-sequence, sequence-profile and profile-profile alignment.

Algorithms↗