Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Mauve: multiple alignment of conserved genomic sequence with rearrangements.

As genomes evolve, they undergo large-scale evolutionary processes that present a challenge to sequence comparison not posed by short sequences. Recombination causes frequent genome rearrangements, horizontal transfer introduces new sequences into bacterial chromosomes, and deletions remove segments of the genome. Consequently, each genome is a mosaic of unique lineage-specific segments, regions shared with a subset of other genomes and segments conserved among all the genomes under consideration. Furthermore, the linear order of these segments may be shuffled among genomes. We present methods for identification and alignment of conserved genomic DNA in the presence of rearrangements and horizontal transfer. Our methods have been implemented in a software package called Mauve. Mauve has been applied to align nine enterobacterial genomes and to determine global rearrangement structure in three mammalian genomes. We have evaluated the quality of Mauve alignments and drawn comparison to other methods through extensive simulations of genome evolution.

Chromosomes, Bacterial↗

Evolutionary divergence plots of homologous proteins.

A simple and efficient method is described for analyzing quantitatively multiple protein sequence alignments and finding the most conserved blocks as well as the maxima of divergence within the set of aligned sequences. It consists of calculating the mean distance and the root-mean-square distance in each column of the multiple alignment, averaging the values in a window of defined length and plotting the results as a function of the position of the window. Due attention is paid to the presence of gaps in the columns. Several examples are provided, using the sequences of several cytochromes c, serine proteases, lysozymes and globins. Two distance matrices are compared, namely the matrix derived by Gribskov and Burgess from the Dayhoff matrix, and the Risler Structural Superposition Matrix. In each case, the divergence plots effectively point to the specific residues which are known to be essential for the catalytic activity of the proteins. In addition, the regions of maximum divergence are clearly delineated. Interestingly, they are generally observed in positions immediately flanking the most conserved blocks. The method should therefore be useful for delineating the peptide segments which will be good candidates for site-directed mutagenesis and for visualizing the evolutionary constraints along homologous polypeptide chains.

Amino Acid Sequence↗

Ancestral sequence alignment under optimal conditions.

BACKGROUND: Multiple genome alignment is an important problem in bioinformatics. An important subproblem used by many multiple alignment approaches is that of aligning two multiple alignments. Many popular alignment algorithms for DNA use the sum-of-pairs heuristic, where the score of a multiple alignment is the sum of its induced pairwise alignment scores. However, the biological meaning of the sum-of-pairs of pairs heuristic is not obvious. Additionally, many algorithms based on the sum-of-pairs heuristic are complicated and slow, compared to pairwise alignment algorithms. An alternative approach to aligning alignments is to first infer ancestral sequences for each alignment, and then align the two ancestral sequences. In addition to being fast, this method has a clear biological basis that takes into account the evolution implied by an underlying phylogenetic tree. In this study we explore the accuracy of aligning alignments by ancestral sequence alignment. We examine the use of both maximum likelihood and parsimony to infer ancestral sequences. Additionally, we investigate the effect on accuracy of allowing ambiguity in our ancestral sequences. RESULTS: We use synthetic sequence data that we generate by simulating evolution on a phylogenetic tree. We use two different types of phylogenetic trees: trees with a period of rapid growth followed by a period of slow growth, and trees with a period of slow growth followed by a period of rapid growth. We examine the alignment accuracy of four ancestral sequence reconstruction and alignment methods: parsimony, maximum likelihood, ambiguous parsimony, and ambiguous maximum likelihood. Additionally, we compare against the alignment accuracy of two sum-of-pairs algorithms: ClustalW and the heuristic of Ma, Zhang, and Wang. CONCLUSION: We find that allowing ambiguity in ancestral sequences does not lead to better multiple alignments. Regardless of whether we use parsimony or maximum likelihood, the success of aligning ancestral sequences containing ambiguity is very sensitive to the choice of gap open cost. Surprisingly, we find that using maximum likelihood to infer ancestral sequences results in less accurate alignments than when using parsimony to infer ancestral sequences. Finally, we find that the sum-of-pairs methods produce better alignments than all of the ancestral alignment methods.

Data Interpretation, Statistical↗

A heuristic Bayesian method for segmenting DNA sequence alignments and detecting evidence for recombination and gene conversion.

We propose a heuristic approach to the detection of evidence for recombination and gene conversion in multiple DNA sequence alignments. The proposed method consists of two stages. In the first stage, a sliding window is moved along the DNA sequence alignment, and phylogenetic trees are sampled from the conditional posterior distribution with MCMC. To reduce the noise intrinsic to inference from the limited amount of data available in the typically short sliding window, a clustering algorithm based on the Robinson-Foulds distance is applied to the trees thus sampled, and the posterior distribution over tree clusters is obtained for each window position. While changes in this posterior distribution are indicative of recombination or gene conversion events, it is difficult to decide when such a change is statistically significant. This problem is addressed in the second stage of the proposed algorithm, where the distributions obtained in the first stage are post-processed with a Bayesian hidden Markov model (HMM). The emission states of the HMM are associated with posterior distributions over phylogenetic tree topology clusters. The hidden states of the HMM indicate putative recombinant segments. Inference is done in a Bayesian sense, sampling parameters from the posterior distribution with MCMC. Of particular interest is the determination of the number of hidden states as an indication of the number of putative recombinant regions. To this end, we apply reversible jump MCMC, and sample the number of hidden states from the respective posterior distribution.

Actins↗

Rapid detection of conserved regions in protein sequences using wavelets.

We present an algorithm to detect protein sub-structural motifs from primary sequence. The input to the algorithm is a set of aligned multiple protein sequences. It uses wavelet transforms to decompose protein sequences represented numerically by different indices (such as polarity, accessible surface area or electron-ion integration potentials of the amino acids). The numerical representation of a protein sequence has significant correlation with its biological activity, thus common motifs are expected to be observable from the wavelet spectrum. The decomposed signals are then up-sampled and similarity search techniques are used to identify similar regions across all the proteins at multiple scales. Results indicate that wavelet transform techniques are a promising approach for rapid motif detection.

Algorithms↗

Joint Bayesian estimation of alignment and phylogeny.

We describe a novel model and algorithm for simultaneously estimating multiple molecular sequence alignments and the phylogenetic trees that relate the sequences. Unlike current techniques that base phylogeny estimates on a single estimate of the alignment, we take alignment uncertainty into account by considering all possible alignments. Furthermore, because the alignment and phylogeny are constructed simultaneously, a guide tree is not needed. This sidesteps the problem in which alignments created by progressive alignment are biased toward the guide tree used to generate them. Joint estimation also allows us to model rate variation between sites when estimating the alignment and to use the evidence in shared insertion/deletions (indels) to group sister taxa in the phylogeny. Our indel model makes use of affine gap penalties and considers indels of multiple letters. We make the simplifying assumption that the indel process is identical on all branches. As a result, the probability of a gap is independent of branch length. We use a Markov chain Monte Carlo (MCMC) method to sample from the posterior of the joint model, estimating the most probable alignment and tree and their support simultaneously. We describe a new MCMC transition kernel that improves our algorithm's mixing efficiency, allowing the MCMC chains to converge even when started from arbitrary alignments. Our software implementation can estimate alignment uncertainty and we describe a method for summarizing this uncertainty in a single plot.

Algorithms↗

RAGA: RNA sequence alignment by genetic algorithm.

We describe a new approach for accurately aligning two homologous RNA sequences when the secondary structure of one of them is known. To do so we developed two software packages, called RAGA and PRAGA, which use a genetic algorithm approach to optimize the alignments. RAGA is mainly an extension of SAGA, an earlier package for multiple protein sequence alignment. In PRAGA several genetic algorithms run in parallel and exchange individual solutions. This method allows us to optimize an objective function that describes the quality of a RNA pairwise alignment, taking into account both primary and secondary structure, including pseudoknots. We report results obtained using PRAGA on nine test cases of pairs of eukaryotic small subunit rRNA sequence (nuclear and mitochondrial).

Algorithms↗

The human genome browser at UCSC.

As vertebrate genome sequences near completion and research refocuses to their analysis, the issue of effective genome annotation display becomes critical. A mature web tool for rapid and reliable display of any requested portion of the genome at any scale, together with several dozen aligned annotation tracks, is provided at http://genome.ucsc.edu. This browser displays assembly contigs and gaps, mRNA and expressed sequence tag alignments, multiple gene predictions, cross-species homologies, single nucleotide polymorphisms, sequence-tagged sites, radiation hybrid data, transposon repeats, and more as a stack of coregistered tracks. Text and sequence-based searches provide quick and precise access to any region of specific interest. Secondary links from individual features lead to sequence details and supplementary off-site databases. One-half of the annotation tracks are computed at the University of California, Santa Cruz from publicly available sequence data; collaborators worldwide provide the rest. Users can stably add their own custom tracks to the browser for educational or research purposes. The conceptual and technical framework of the browser, its underlying MYSQL database, and overall use are described. The web site currently serves over 50,000 pages per day to over 3000 different users.

California↗

ddbRNA: detection of conserved secondary structures in multiple alignments.

MOTIVATION: Structured non-coding RNAs (ncRNAs) have a very important functional role in the cell. No distinctive general features common to all ncRNA have yet been discovered. This makes it difficult to design computational tools able to detect novel ncRNAs in the genomic sequence. RESULTS: We devised an algorithm able to detect conserved secondary structures in both pairwise and multiple DNA sequence alignments with computational time proportional to the square of the sequence length. We implemented the algorithm for the case of pairwise and three-way alignments and tested it on ncRNAs obtained from public databases. On the test sets, the pairwise algorithm has a specificity greater than 97% with a sensitivity varying from 22.26% for Blast alignments to 56.35% for structural alignments. The three-way algorithm behaves similarly. Our algorithm is able to efficiently detect a conserved secondary structure in multiple alignments.

Algorithms↗

Phylogenetic relationships among megabats, microbats, and primates.

We present 744 nucleotide base positions from the mitochondrial 12S rRNA gene and 236 base positions from the mitochondrial cytochrome oxidase subunit I gene for a microbat, Brachyphylla cavernarum, and a megabat, Pteropus capestratus, in phylogenetic analyses with homologous DNA sequences from Homo sapiens, Mus musculus (house mouse), and Gallus gallus (chicken). We use information on evolutionary rate differences for different types of sequence change to establish phylogenetic character weights, and we consider alternative rRNA alignment strategies in finding that this mtDNA data set clearly supports bat monophyly. This result is found despite variations in outgroup used, gap coding scheme, and order of input for DNA sequences in multiple alignment bouts. These findings are congruent with morphological characters including details of wing structure as well as cladistic analyses of amino acid sequences for three globin genes and indicate that neurological similarities between megabats and primates are due to either retention of primitive characters or to convergent evolution rather than to inheritance from a common ancestor. This finding also indicates a single origin for flight among mammals.

Amino Acid Sequence↗

Flexible algorithm for direct multiple alignment of protein structures and sequences.

The recently described equivalence between the alignment of two proteins and a conformation of a lattice chain on a two-dimensional square lattice is extended to multiple alignments. The search for the optimal multiple alignment between several proteins, which is equivalent to finding the energy minimum in the conformational space of a multi-dimensional lattice chain, is studied by the Monte Carlo approach. This method, while not deterministic, and for two-dimensional problems slower than dynamic programming, can accept arbitrary scoring functions, including non-local ones, and its speed decreases slowly with increasing number of dimensions. For the local scoring functions, the MC algorithm can also reproduce known exact solutions for the direct multiple alignments. As illustrated by examples, both for structure- and sequence-based alignments, direct multi-dimensional alignments are able to capture weak similarities between divergent families much better than ones built from pairwise alignments by a hierarchical approach.

Algorithms↗

Intraspecific phylogeny of the New Zealand short-tailed bat Mystacina tuberculata inferred from multiple mitochondrial gene sequences.

An intraspecific phylogeny was established for the New Zealand short-tailed bat Mystacina tuberculata using a 2,878-bp sequence alignment from multiple mitochondrial genes (control region, ND2, 12S ribosomal RNA [rRNA], 16S rRNA, and tRNA). The inferred phylogeny comprises six lineages, with estimated divergences extending back between 0.93 and 0.68 million years to the middle Pleistocene. The lineages do not correspond to the existing subspecific taxonomy. Although multiple lineages occur sympatrically in many populations, the lineages are geographically structured. This structure has persisted despite repeated cycles of range expansion and contraction in response to climatic oscillations and catastrophic volcanic eruptions. The distribution of lineages among populations in central North Island indicates that a hybrid zone was formed by simultaneous colonization from single-lineage source populations inhabiting remote forest refugia. The observed pattern is not typical of microbats, which because of their high mobility generally exhibit low levels of genetic differentiation and geographic structure over continental ranges. Although lineages of M. tuberculata occur sympatrically in many populations, genetic distances between them are sufficiently large to suggest that they may be considered evolutionary significant units or taxonomic subspecies.

Animals↗

Sequencing of 42kb of the APO E-C2 gene cluster reveals a new gene: PEREC1.

Through the sequencing of a 42kb cosmid clone we describe a new gene, designated PEREC1, located approximately 1.5kb centromeric of the human apolipoprotein (APO) E-C2 cluster. The combination of dotplot analysis, predicted coding potential and interrogation of the Expressed Sequence Tag (EST) database determined the genomic organisation of PEREC1. Sequence alignment with multiple overlapping ESTs confirmed the predicted splice sites. The predicted cDNA and amino acid sequences of PEREC1 have extensive similarity to the Caenorhabditis elegans protein, C18E9.6. Conserved structural and functional motifs have been defined by combining nucleotide and amino acid analyses to identify third base degeneracy and therefore selection at the protein level. The Poliovirus Receptor Related Protein2 gene (PRR2), previously mapped to chromosome 19q13.2 by Fluorescent In-Situ Hybridisation, has also been located approximately 17kb centromeric of APO E.

Alzheimer Disease↗

Calculating percent identity between protein or DNA sequences with a word processor.

Two macros, to calculate percentage identity between protein or DNA sequences using the Microsoft Word word processor, are described. The user prepares an alignment file of multiple sequences which is used by the macros to calculate number of matches, number of mismatches, total number of compared positions, and the percent identity. The macros are especially useful when alignment of multiple sequences is possible only by eye.

Algorithms↗

CD-Search: protein domain annotations on the fly.

We describe the Conserved Domain Search service (CD-Search), a web-based tool for the detection of structural and functional domains in protein sequences. CD-Search uses BLAST(R) heuristics to provide a fast, interactive service, and searches a comprehensive collection of domain models. Search results are displayed as domain architecture cartoons and pairwise alignments between the query and domain-model consensus sequences. Search results may be visualized in further detail by embedding the query sequence into multiple alignment displays and by mapping onto three-dimensional molecular graphic displays of known structures within the domain family. CD-Search can be accessed at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi.

Amino Acid Sequence↗

JPred: a consensus secondary structure prediction server.

UNLABELLED: An interactive protein secondary structure prediction Internet server is presented. The server allows a single sequence or multiple alignment to be submitted, and returns predictions from six secondary structure prediction algorithms that exploit evolutionary information from multiple sequences. A consensus prediction is also returned which improves the average Q3 accuracy of prediction by 1% to 72.9%. The server simplifies the use of current prediction algorithms and allows conservation patterns important to structure and function to be identified. AVAILABILITY: http://barton.ebi.ac.uk/servers/jpred.h tml CONTACT: geoff@ebi.ac.uk

Algorithms↗

Molecular modelling studies on G protein-coupled receptors: from sequence to structure?

In pharmaceutical research G protein-coupled receptors (GPCR) emerged as a superfamily of prominent drug targets. Extensive protein sequence analyses on GPCRs revealed a common protein topology consisting of a membrane-spanning seven-helix bundle, which is believed to accommodate the binding site for low-molecular weight ligands. Enormous efforts are undertaken to generate GPCR structural models by means of molecular modelling, since these have already been shown to aid the process of lead structure finding and optimisation in that they provide atomistic models for structure-based drug design approaches. One of the most critical steps in modelling the transmembrane domains of GPCRs is the assignment of the putative transmembrane sequence stretches from multiple protein sequence alignment analyses. This study focuses on the comparative evaluation of protein sequence analysis tools, such as periodicity analyses, multiple sequence analyses or directional helix descriptors, especially developed for modelling the 7TM domains of GPCRs. In this context we will demonstrate that from application of different methods contradictory results can be obtained for the identification of the putative transmembrane sequence stretches of peptide-binding GPCRs, as exemplified with a comprehensive protein sequence analysis study based on the most prominent members of that receptor class (angiotensin II, CCK/gastrin, interleukin 8, endothelin, etc.).

Animals↗

4SALE--a tool for synchronous RNA sequence and secondary structure alignment and editing.

BACKGROUND: In sequence analysis the multiple alignment builds the fundament of all proceeding analyses. Errors in an alignment could strongly influence all succeeding analyses and therefore could lead to wrong predictions. Hand-crafted and hand-improved alignments are necessary and meanwhile good common practice. For RNA sequences often the primary sequence as well as a secondary structure consensus is well known, e.g., the cloverleaf structure of the t-RNA. Recently, some alignment editors are proposed that are able to include and model both kinds of information. However, with the advent of a large amount of reliable RNA sequences together with their solved secondary structures (available from e.g. the ITS2 Database), we are faced with the problem to handle sequences and their associated secondary structures synchronously. RESULTS: 4SALE fills this gap. The application allows a fast sequence and synchronous secondary structure alignment for large data sets and for the first time synchronous manual editing of aligned sequences and their secondary structures. This study describes an algorithm for the synchronous alignment of sequences and their associated secondary structures as well as the main features of 4SALE used for further analyses and editing. 4SALE builds an optimal and unique starting point for every RNA sequence and structure analysis. CONCLUSION: 4SALE, which provides an user-friendly and intuitive interface, is a comprehensive toolbox for RNA analysis based on sequence and secondary structure information. The program connects sequence and structure databases like the ITS2 Database to phylogeny programs as for example the CBCAnalyzer. 4SALE is written in JAVA and therefore platform independent. The software is freely available and distributed from the website at http://4sale.bioapps.biozentrum.uni-wuerzburg.de.

Algorithms↗