Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Using multiple alignments and phylogenetic trees to detect RNA secondary structure.

We describe a statistical method to determine if a pair of columns in a multiple alignment of a homologous family of RNA sequences shows evidence of being base paired. The method makes explicit use of a given phylogenetic tree for the sequences in the alignment. It is tested on a multiple alignment of 16S rRNA sequences with good results.

Base Composition↗

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software↗

SSMAL: similarity searching with alignment graphs.

MOTIVATION: We want to provide biologists with a fast and sensitive scanning tool for searching local alignments of a protein query sequence against databases of protein multiple alignments, such as ProDom. Conversely, we want to provide a tool for locally aligning a protein multiple alignment query against a protein database such as SWISSPROT. RESULTS: We developed the program SSMAL (Shuffling Similarities with Multiple Alignments) which utilizes features of the Blast (Altschul et al., J. Mol. Biol., 215, 403-410, 1990) algorithm and part of the Blast code. Our software allows both scanning of multiple alignments and searching with a multiple alignment. Deletions in the multiple alignment only are handled and a SSMAL search may miss some similarities found by a profile search. However, an SSMAL scan of a database such as ProDom would be 20-30 times faster that a profile scan. In the worst case, a SSMAL search is approximately 9 times faster than a profile search. AVAILABILITY: http://www.dkfz-heidelberg.de/tbi/ people/nicodeme and follow the hyperlink SSMAL. CONTACT: p.nicodeme@DKFZ-Heidelberg.de

Amino Acid Sequence↗

Automated whole-genome multiple alignment of rat, mouse, and human.

We have built a whole-genome multiple alignment of the three currently available mammalian genomes using a fully automated pipeline that combines the local/global approach of the Berkeley Genome Pipeline and the LAGAN program. The strategy is based on progressive alignment and consists of two main steps: (1) alignment of the mouse and rat genomes, and (2) alignment of human to either the mouse-rat alignments from step 1, or the remaining unaligned mouse and rat sequences. The resulting alignments demonstrate high sensitivity, with 87% of all human gene-coding areas aligned in both mouse and rat. The specificity is also high: <7% of the rat contigs are aligned to multiple places in human, and 97% of all alignments with human sequence >100 kb agree with a three-way synteny map built independently, using predicted exons in the three genomes. At the nucleotide level <1% of the rat nucleotides are mapped to multiple places in the human sequence in the alignment, and 96.5% of human nucleotides within all alignments agree with the synteny map. The alignments are publicly available online, with visualization through the novel Multi-VISTA browser that we also present.

Animals↗

The cDNA sequence of proton-pumping nicotinamide nucleotide transhydrogenase from man and mouse.

cDNA clones for the human and mouse nicotinamide nucleotide transhydrogenases have been isolated and their sequences have been determined. Multiple alignments show that the functional proteins are encoded by single mRNAs. The deduced amino acid sequences are approximately 95% identical for the previously known bovine, and the human and mouse proteins. The major variable region is located in the presequence. It is proposed that all mammalian transhydrogenases have a similar structure.

Amino Acid Sequence↗

PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments.

PAL2NAL is a web server that constructs a multiple codon alignment from the corresponding aligned protein sequences. Such codon alignments can be used to evaluate the type and rate of nucleotide substitutions in coding DNA for a wide range of evolutionary analyses, such as the identification of levels of selective constraint acting on genes, or to perform DNA-based phylogenetic studies. The server takes a protein sequence alignment and the corresponding DNA sequences as input. In contrast to other existing applications, this server is able to construct codon alignments even if the input DNA sequence has mismatches with the input protein sequence, or contains untranslated regions and polyA tails. The server can also deal with frame shifts and inframe stop codons in the input models, and is thus suitable for the analysis of pseudogenes. Another distinct feature is that the user can specify a subregion of the input alignment in order to specifically analyze functional domains or exons of interest. The PAL2NAL server is available at http://www.bork.embl.de/pal2nal.

Codon↗

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence↗

Using evolutionary Expectation Maximization to estimate indel rates.

MOTIVATION: The Expectation Maximization (EM) algorithm, in the form of the Baum-Welch algorithm (for hidden Markov models) or the Inside-Outside algorithm (for stochastic context-free grammars), is a powerful way to estimate the parameters of stochastic grammars for biological sequence analysis. To use this algorithm for multiple-sequence evolutionary modelling, it would be useful to apply the EM algorithm to estimate not only the probability parameters of the stochastic grammar, but also the instantaneous mutation rates of the underlying evolutionary model (to facilitate the development of stochastic grammars based on phylogenetic trees, also known as Statistical Alignment). Recently, we showed how to do this for the point substitution component of the evolutionary process; here, we extend these results to the indel process. RESULTS: We present an algorithm for maximum-likelihood estimation of insertion and deletion rates from multiple sequence alignments, using EM, under the single-residue indel model owing to Thorne, Kishino and Felsenstein (the 'TKF91' model). The algorithm converges extremely rapidly, gives accurate results on simulated data that are an improvement over parsimonious estimates (which are shown to underestimate the true indel rate), and gives plausible results on experimental data (coronavirus envelope domains). Owing to the algorithm's close similarity to the Baum-Welch algorithm for training hidden Markov models, it can be used in an 'unsupervised' fashion to estimate rates for unaligned sequences, or estimate several sets of rates for sequences with heterogenous rates. AVAILABILITY: Software implementing the algorithm and the benchmark is available under GPL from http://www.biowiki.org/

Algorithms↗

Blocks-based methods for detecting protein homology.

The most highly conserved regions of proteins can be represented as blocks of aligned sequence segments, typically with multiple blocks for a given protein family. The Blocks Database World Wide Web (http://blocks.fhcrc.org) and e-mail (blocks@blocks. fhcrc.org) servers provide tools to search DNA and protein queries against the Blocks+ Database of multiple alignments. We describe features for detection of distant relationships using blocks. Blocks+ includes protein families from the PROSITE, Prints, Pfam-A, ProDom and Domo databases. Other features include searching Blocks+ with the BLIMPS and NCBI's IMPALA programs, sequence logos, phylogenetic trees, three-dimensional display of blocks on PDB structures, and a polymerase chain reaction (PCR) primer design strategy based on blocks.

Amino Acid Sequence↗

Using guide trees to construct multiple-sequence evolutionary HMMs.

MOTIVATION: Score-based progressive alignment algorithms do dynamic programming on successive branches of a guide tree. The analogous probabilistic construct is an Evolutionary HMM. This is a multiple-sequence hidden Markov model (HMM) made by combining transducers (conditionally normalised Pair HMMs) on the branches of a phylogenetic tree. METHODS: We present general algorithms for constructing an Evolutionary HMM from any Pair HMM and for doing dynamic programming to any Multiple-sequence HMM. RESULTS: Our prototype implementation, Handel, is based on the Thorne-Kishino-Felsenstein evolutionary model and is benchmarked using structural reference alignments.

Algorithms↗

PowerBLAST: a new network BLAST application for interactive or automated sequence analysis and annotation.

As the rate of DNA sequencing increases, analysis by sequence similarity search will need to become much more efficient in terms of sensitivity, specificity, automation potential, and consistency in annotation. PowerBLAST was developed, in part, to address these problems. PowerBLAST includes a number of options for masking repetitive elements and low complexity subsequences. It also has the capacity to restrict the search to any level of NCBI's taxonomy index, thus supporting "comparative genomics" applications. Postprocessing of the BLAST output using the SIM series of algorithms produces optimal, gapped alignments, and multiple alignments when a region of the query sequence matches multiple database sequences. PowerBLAST is capable of processing sequences of any length because it divides long query sequences into overlapping fragments and then merges the results after searching. The results may be viewed graphically, as a textual representation, or as an HTML page with links to GenBank and Entrez. For matching database sequences, annotated features are superimposed on the aligned query sequence in the output, thus greatly increasing the ease of interpretation. Such features may be used for automated annotation of new sequence because PowerBLAST output in ASN.1 form may be "dragged and dropped" into NCBI's Sequin program for sequence annotation and submission. PowerBLAST is capable of analyzing and annotating a 100-kb query in 60 min on NCBI's BLAST server.

Amino Acid Sequence↗

Enhanced genome annotation using structural profiles in the program 3D-PSSM.

A method (three-dimensional position-specific scoring matrix, 3D-PSSM) to recognise remote protein sequence homologues is described. The method combines the power of multiple sequence profiles with knowledge of protein structure to provide enhanced recognition and thus functional assignment of newly sequenced genomes. The method uses structural alignments of homologous proteins of similar three-dimensional structure in the structural classification of proteins (SCOP) database to obtain a structural equivalence of residues. These equivalences are used to extend multiply aligned sequences obtained by standard sequence searches. The resulting large superfamily-based multiple alignment is converted into a PSSM. Combined with secondary structure matching and solvation potentials, 3D-PSSM can recognise structural and functional relationships beyond state-of-the-art sequence methods. In a cross-validated benchmark on 136 homologous relationships unambiguously undetectable by position-specific iterated basic local alignment search tool (PSI-Blast), 3D-PSSM can confidently assign 18 %. The method was applied to the remaining unassigned regions of the Mycoplasma genitalium genome and an additional 13 regions were assigned with 95 % confidence. 3D-PSSM is available to the community as a web server: http://www.bmm.icnet.uk/servers/3dpssm

Algorithms↗

A sequence alignment-independent method for protein classification.

Annotation of the rapidly accumulating body of sequence data relies heavily on the detection of remote homologues and functional motifs in protein families. The most popular methods rely on sequence alignment. These include programs that use a scoring matrix to compare the probability of a potential alignment with random chance and programs that use curated multiple alignments to train profile hidden Markov models (HMMs). Related approaches depend on bootstrapping multiple alignments from a single sequence. However, alignment-based programs have limitations. They make the assumption that contiguity is conserved between homologous segments, which may not be true in genetic recombination or horizontal transfer. Alignments also become ambiguous when sequence similarity drops below 40%. This has kindled interest in classification methods that do not rely on alignment. An approach to classification without alignment based on the distribution of contiguous sequences of four amino acids (4-grams) was developed. Interest in 4-grams stemmed from the observation that almost all theoretically possible 4-grams (20(4)) occur in natural sequences and the majority of 4-grams are uniformly distributed. This implies that the probability of finding identical 4-grams by random chance in unrelated sequences is low. A Bayesian probabilistic model was developed to test this hypothesis. For each protein family in Pfam-A and PIR-PSD, a feature vector called a probe was constructed from the set of 4-grams that best characterised the family. In rigorous jackknife tests, unknown sequences from Pfam-A and PIR-PSD were compared with the probes for each family. A classification result was deemed a true positive if the probe match with the highest probability was in first place in a rank-ordered list. This was achieved in 70% of cases. Analysis of false positives suggested that the precision might approach 85% if selected families were clustered into subsets. Case studies indicated that the 4-grams in common between an unknown and the best matching probe correlated with functional motifs from PRINTS. The results showed that remote homologues and functional motifs could be identified from an analysis of 4-gram patterns.

Algorithms↗

Flexible protein sequence patterns. A sensitive method to detect weak structural similarities.

The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Amino Acid Sequence↗

Use of long sequence alignments to study the evolution and regulation of mammalian globin gene clusters.

The determination of long segments of DNA sequences encompassing the beta- and alpha-globin gene clusters has provided an unprecedented data base for analysis of genome evolution and regulation of gene clusters. A newly developed computer tool kit generates local alignments between such long sequences in a space-efficient manner, helps the user analyze the alignments effectively, and finds consistently aligning blocks of sequences in multiple pairwise comparisons. Such sequence analyses among the beta-like globin gene clusters of human, galago, rabbit, and mouse have revealed the general patterns of evolution of this gene cluster. Alignments in the flanking regions are very useful in assigning orthologous relationships. Investigation of such matches between the mouse and human beta-like globin gene clusters has led to a reassessment of some orthologous assignments in mouse and to a revision of the proposed pathway for evolution of this gene cluster. In general, the interspersed repetitive elements have inserted independently, presumably via a retrotransposition mechanism, in the different mammalian lineages. However, some examples of ancient L1 repeats are found, including one between the epsilon- and gamma-globin genes that appears to have been in the ancestral eutherian gene cluster. Prominent matching sequences are found in a long region 5' to the epsilon-globin gene, the locus control region (LCR) that is a positive regulator of the entire gene cluster. Three-way alignments among the human, goat, and rabbit sequences can extend for > or = 3 kb in part of the LCR (DNase hypersensitive site 3), indicating that the cis-acting components of this complex regulatory region cover a long segment of DNA. In contrast to the beta-like globin gene clusters, the alpha-like globin gene clusters of many mammals occur in very G+C-rich isochores and contain prominent CpG islands. The regions between the alpha-like globin genes are evolving faster than the intergenic regions of the beta-like globin gene clusters. The contrasts between the two gene clusters can be attributed to differences in DNA metabolism in the isochore. The proximal control elements of the rabbit alpha-globin gene are located both 5' to and within the gene. All of this region is part of a prominent CpG island that may be acting as an extended, enhancer-independent promoter. One can hypothesize that the analogue to the LCR in the alpha-globin gene cluster may interface with the distinctive alpha-globin promoter in ways different from the interaction between the beta LCR and the promoters of beta-like globin genes.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Characterization of the active site and thermostability regions of endoxylanase from Thermoanaerobacterium saccharolyticum B6A-RI.

Deletion mutants were constructed from pZEP12, which contained the intact Thermoanaerobacterium saccharolyticum endoxylanase gene (xynA). Deletion of 1.75 kb from the N-terminal end of xynA resulted in a mutant enzyme that retained activity but lost thermostability. Deletion of 1.05 kb from the C terminus did not alter thermostability or activity. The deduced amino acid sequence of T. saccharolyticum B6A-RI endoxylanase XynA was aligned with five other family F beta-glycanases by using the PILEUP program of the Genetics Computer Group package. This multiple alignment of amino acid sequences revealed six highly conserved motifs which included the consensus sequence consisting of a hydrophobic amino acid, Ser or Thr, Glu, a hydrophobic amino acid, Asp, and a hydrophobic amino acid in the catalytic domain. Endoxylanase was inhibited by EDAC [1-(3-dimethylamino propenyl)-3-ethylcarbodiimide hydrochloride], suggesting that Asp and/or Glu was involved in catalysis. Three aspartic acids, two glutamic acids, and one histidine were conserved in all six enzymes aligned. Hydrophobic cluster analysis revealed that two Asp and one Glu occur in the same hydrophobic clusters in T. saccharolyticum B6A-RI endoxylanase and two other enzymes belonging to family F beta-glycanases and suggests their involvement in a catalytic triad. These two Asp and one Glu in XynA from T. saccharolyticum were targeted for analysis by site-specific mutagenesis. Substitution of Asp-537 and Asp-602 by Asn and Glu-600 by Gln completely destroyed endoxylanase activity. These results suggest that these three amino acids form a catalytic triad that functions in a general acid catalysis mechanism.

Amino Acid Sequence↗

Building multiple alignments from pairwise alignments.

Given a family of related sequences, one can first determine alignments between various pairs of those sequences, then construct a simultaneous alignment of all the sequences that is determined in a natural manner by the set of pairwise alignments. This approach is sometimes effective for exposing the existence and locations of conserved regions, which can then be aligned by more sensitive multiple-alignment methods. This paper presents an efficient algorithm for constructing a multiple alignment from a set of pairwise alignments.

Algorithms↗

Predicting DNA-binding sites of proteins from amino acid sequence.

BACKGROUND: Understanding the molecular details of protein-DNA interactions is critical for deciphering the mechanisms of gene regulation. We present a machine learning approach for the identification of amino acid residues involved in protein-DNA interactions. RESULTS: We start with a Naïve Bayes classifier trained to predict whether a given amino acid residue is a DNA-binding residue based on its identity and the identities of its sequence neighbors. The input to the classifier consists of the identities of the target residue and 4 sequence neighbors on each side of the target residue. The classifier is trained and evaluated (using leave-one-out cross-validation) on a non-redundant set of 171 proteins. Our results indicate the feasibility of identifying interface residues based on local sequence information. The classifier achieves 71% overall accuracy with a correlation coefficient of 0.24, 35% specificity and 53% sensitivity in identifying interface residues as evaluated by leave-one-out cross-validation. We show that the performance of the classifier is improved by using sequence entropy of the target residue (the entropy of the corresponding column in multiple alignment obtained by aligning the target sequence with its sequence homologs) as additional input. The classifier achieves 78% overall accuracy with a correlation coefficient of 0.28, 44% specificity and 41% sensitivity in identifying interface residues. Examination of the predictions in the context of 3-dimensional structures of proteins demonstrates the effectiveness of this method in identifying DNA-binding sites from sequence information. In 33% (56 out of 171) of the proteins, the classifier identifies the interaction sites by correctly recognizing at least half of the interface residues. In 87% (149 out of 171) of the proteins, the classifier correctly identifies at least 20% of the interface residues. This suggests the possibility of using such classifiers to identify potential DNA-binding motifs and to gain potentially useful insights into sequence correlates of protein-DNA interactions. CONCLUSION: Naïve Bayes classifiers trained to identify DNA-binding residues using sequence information offer a computationally efficient approach to identifying putative DNA-binding sites in DNA-binding proteins and recognizing potential DNA-binding motifs.

Algorithms↗