Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

ProMSED: protein multiple sequence editor for Windows 3.11/95.

MOTIVATION: Most protein sequence alignment algorithms give similar results on closely related proteins, while manual intervention may be needed for distantly related molecules. To correct the alignment, it is often necessary to repeat calculations on selected parts of the alignments and edit the alignment manually. Software implementing such interactive alignment procedures is of significance. RESULTS: This paper presents a new MS Windows application called ProMSED for both automatic and manual protein sequence alignment. The program reads main sequence formats and has a user-friendly interface. ProMSED performs automatic (ClustalV algorithm) alignments, alignment visualization and editing, and it allows sequences to be aligned interactively leaving previously aligned regions unchanged. Manual alignment and sequence analysis are facilitated by colouring schemes reflecting amino acid similarity of mutational and physicochemical properties. The interactive alignment of a diverged set of reverse transcriptases has located four out of six known conserved motifs. AVAILABILITY: ProMSED is available on request from the authors. DEMO is available from ftp://ftp.ebi.ac.uk/pub/ software/dos/promsed/ or ftp://iubio.bio.indiana.edu/molbio/ ibmpc/.

Algorithms↗

Local multiple alignment by consensus matrix.

A new algorithm for aligning several sequences based on the calculation of a consensus matrix and the comparison of all the sequences using this consensus matrix is described. This consensus matrix contains the preference scores of each nucleotide/amino acid and gaps in every position of the alignment. Two modifications of the algorithm corresponding to the evolutionary and functional meanings of the alignment were developed. The first one solves the best-fitting problem without any penalty for end gaps and with an internal gap penalty function independent on the gap length. This algorithm should be used when comparing evolutionary-related proteins for identifying the most conservative residues. The other modification of the algorithm finds the most similar segments in the given sequences. It can be used for finding those parts of the sequences that are responsible for the same biological function. In this case the gap penalty function was chosen to be proportional to the gap length. The result of aligning amino acid sequences of neutral proteases and a compilation of 65 allosteric effectors and substrates of PEP carboxylase are presented.

Algorithms↗

GRIL: genome rearrangement and inversion locator.

UNLABELLED: GRIL is a tool to automatically identify collinear regions in a set of bacterial-size genome sequences. GRIL uses three basic steps. First, regions of high sequence identity are located. Second, some of these regions are filtered based on user-specified criteria. Finally, the remaining regions of sequence identity are used to define significant collinear regions among the sequences. By locating collinear regions of sequence, GRIL provides a basis for multiple genome alignment using current alignment systems. GRIL also provides a basis for using current inversion distance tools to infer phylogeny. AVAILABILITY: GRIL is implemented in C++ and runs on any x86-based Linux or Windows platform. It is available from http://asap.ahabs.wisc.edu/gril

Chromosome Inversion↗

Six-fold speed-up of Smith-Waterman sequence database searches using parallel processing on common microprocessors.

MOTIVATION: Sequence database searching is among the most important and challenging tasks in bioinformatics. The ultimate choice of sequence-search algorithm is that of Smith-Waterman. However, because of the computationally demanding nature of this method, heuristic programs or special-purpose hardware alternatives have been developed. Increased speed has been obtained at the cost of reduced sensitivity or very expensive hardware. RESULTS: A fast implementation of the Smith-Waterman sequence-alignment algorithm using Single-Instruction, Multiple-Data (SIMD) technology is presented. This implementation is based on the MultiMedia eXtensions (MMX) and Streaming SIMD Extensions (SSE) technology that is embedded in Intel's latest microprocessors. Similar technology exists also in other modern microprocessors. Six-fold speed-up relative to the fastest previously known Smith-Waterman implementation on the same hardware was achieved by an optimized 8-way parallel processing approach. A speed of more than 150 million cell updates per second was obtained on a single Intel Pentium III 500 MHz microprocessor. This is probably the fastest implementation of this algorithm on a single general-purpose microprocessor described to date.

Algorithms↗

Assessing strategies for improved superfamily recognition.

There are more than 200 completed genomes and over 1 million nonredundant sequences in public repositories. Although the structural data are more sparse (approximately 13,000 nonredundant structures solved to date), several powerful sequence-based methodologies now allow these structures to be mapped onto related regions in a significant proportion of genome sequences. We review a number of publicly available strategies for providing structural annotations for genome sequences, and we describe the protocol adopted to provide CATH structural annotations for completed genomes. In particular, we assess the performance of several sequence-based protocols employing Hidden Markov model (HMM) technologies for superfamily recognition, including a new approach (SAMOSA [sequence augmented models of structure alignments]) that exploits multiple structural alignments from the CATH domain structure database when building the models. Using a data set of remote homologs detected by structure comparison and manually validated in CATH, a single-seed HMM library was able to recognize 76% of the data set. Including the SAMOSA models in the HMM library showed little gain in homolog recognition, although a slight improvement in alignment quality was observed for very remote homologs. However, using an expanded 1D-HMM library, CATH-ISL increased the coverage to 86%. The single-seed HMM library has been used to annotate the protein sequences of 120 genomes from all three major kingdoms, allowing up to 70% of the genes or partial genes to be assigned to CATH superfamilies. It has also been used to recruit sequences from Swiss-Prot and TrEMBL into CATH domain superfamilies, expanding the CATH database eightfold.

Databases, Protein↗

SSAHA: a fast search method for large DNA databases.

We describe an algorithm, SSAHA (Sequence Search and Alignment by Hashing Algorithm), for performing fast searches on databases containing multiple gigabases of DNA. Sequences in the database are preprocessed by breaking them into consecutive k-tuples of k contiguous bases and then using a hash table to store the position of each occurrence of each k-tuple. Searching for a query sequence in the database is done by obtaining from the hash table the "hits" for each k-tuple in the query sequence and then performing a sort on the results. We discuss the effect of the tuple length k on the search speed, memory usage, and sensitivity of the algorithm and present the results of computational experiments which show that SSAHA can be three to four orders of magnitude faster than BLAST or FASTA, while requiring less memory than suffix tree methods. The SSAHA algorithm is used for high-throughput single nucleotide polymorphism (SNP) detection and very large scale sequence assembly. Also, it provides Web-based sequence search facilities for Ensembl projects.

Algorithms↗

Cloning the human and mouse MMS19 genes and functional complementation of a yeast mms19 deletion mutant.

The MMS19 gene of the yeast Saccharomyces cerevisiae encodes a polypeptide of unknown function which is required for both nucleotide excision repair (NER) and RNA polymerase II (RNAP II) transcription. Here we report the molecular cloning of human and mouse orthologs of the yeast MMS19 gene. Both human and Drosophila MMS19 cDNAs correct thermosensitive growth and sensitivity to killing by UV radiation in a yeast mutant deleted for the MMS19 gene, indicating functional conservation between the yeast and mammalian gene products. Alignment of the translated sequences of MMS19 from multiple eukaryotes, including mouse and human, revealed the presence of several conserved regions, including a HEAT repeat domain near the C-terminus. The presence of HEAT repeats, coupled with functional complementation of yeast mutant phenotypes by the orthologous protein from higher eukaryotes, suggests a role of Mms19 protein in the assembly of a multiprotein complex(es) required for NER and RNAP II transcription. Both the mouse and human genes are ubiquitously expressed as multiple transcripts, some of which appear to derive from alternative splicing. The ratio of different transcripts varies in several different tissue types.

Alternative Splicing↗

Statistical modeling and analysis of the LAGLIDADG family of site-specific endonucleases and identification of an intein that encodes a site-specific endonuclease of the HNH family.

The LAGLIDADG and HNH families of site-specific DNA endonucleases encoded by viruses, bacteriophages as well as archaeal, eucaryotic nuclear and organellar genomes are characterized by the sequence motifs 'LAGLIDADG' and 'HNH', respectively. These endonucleases have been shown to occur in different environments: LAGLIDADG endonucleases are found in inteins, archaeal and group I introns and as free standing open reading frames (ORFs); HNH endonucleases occur in group I and group II introns and as ORFs. Here, statistical models (hidden Markov models, HMMs) that encompass both the conserved motifs and more variable regions of these families have been created and employed to characterize known and potential new family members. A number of new, putative LAGLIDADG and HNH endonucleases have been identified including an intein-encoded HNH sequence. Analysis of an HMM-generated multiple alignment of 130 LAGLIDADG family members and the three-dimensional structure of the I- Cre I endonuclease has enabled definition of the core elements of the repeated domain (approximately 90 residues) that is present in this family of proteins. A conserved negatively charged residue is proposed to be involved in catalysis. Phylogenetic analysis of the two families indicates a lack of exchange of endonucleases between different mobile elements (environments) and between hosts from different phylogenetic kingdoms. However, there does appear to have been considerable exchange of endonuclease domains amongst elements of the same type. Such events are suggested to be important for the formation of elements of new specficity.

Amino Acid Sequence↗

Genome-wide detection and analysis of homologous recombination among sequenced strains of Escherichia coli.

BACKGROUND: Comparisons of complete bacterial genomes reveal evidence of lateral transfer of DNA across otherwise clonally diverging lineages. Some lateral transfer events result in acquisition of novel genomic segments and are easily detected through genome comparison. Other more subtle lateral transfers involve homologous recombination events that result in substitution of alleles within conserved genomic regions. This type of event is observed infrequently among distantly related organisms. It is reported to be more common within species, but the frequency has been difficult to quantify since the sequences under comparison tend to have relatively few polymorphic sites. RESULTS: Here we report a genome-wide assessment of homologous recombination among a collection of six complete Escherichia coli and Shigella flexneri genome sequences. We construct a whole-genome multiple alignment and identify clusters of polymorphic sites that exhibit atypical patterns of nucleotide substitution using a random walk-based method. The analysis reveals one large segment (approximately 100 kb) and 186 smaller clusters of single base pair differences that suggest lateral exchange between lineages. These clusters include portions of 10% of the 3,100 genes conserved in six genomes. Statistical analysis of the functional roles of these genes reveals that several classes of genes are over-represented, including those involved in recombination, transport and motility. CONCLUSION: We demonstrate that intraspecific recombination in E. coli is much more common than previously appreciated and may show a bias for certain types of genes. The described method provides high-specificity, conservative inference of past recombination events.

Alleles↗

Redundancy based detection of sequence polymorphisms in expressed sequence tag data using autoSNP.

UNLABELLED: AutoSNP is a program to detect single nucleotide polymorphisms (SNPs) and insertion/deletion polymorphisms (indels) in expressed sequence tag (EST) data. The program uses d2cluster and cap3 to cluster and align EST sequences, and uses redundancy to differentiate between candidate SNPs and sequence errors. Candidate polymorphisms are identified as occurring in multiple reads within an alignment. For each candidate SNP, two measures of confidence are calculated, the redundancy of the polymorphism at a SNP locus and the co segregation of the candidate SNP with other SNPs in the alignment. AVAILABILITY: The program was written in PERL and is freely available to non-commercial users by request from the authors.

Cluster Analysis↗

Overexpression of the receptor for hyaluronan-mediated motility (RHAMM) characterizes the malignant clone in multiple myeloma: identification of three distinct RHAMM variants.

The receptor for hyaluronan (HA)-mediated motility (RHAMM) controls motility by malignant cells in myeloma and is abnormally expressed on the surface of most malignant B and plasma cells in blood or bone marrow (BM) of patients with multiple myeloma (MM). RHAMM cDNA was cloned and sequenced from the malignant B and plasma cells comprising the myeloma B lineage hierarchy. Three distinct RHAMM gene products, RHAMMFL, RHAMM-48, and RHAMM-147, were cloned from MM B and plasma cells. RHAMMFL was 99% homologous to the published sequence of RHAMM. RHAMM-48 and RHAMM-147 variants align with RHAMMFL, but are characterized by sequence deletions of 48 bp (16 amino acids [aa]) and 147 bp (49 aa), respectively. The relative frequency of these RHAMM transcripts in MM plasma cells was determined by cloning of reverse-transcriptase polymerase chain reaction (RT-PCR) products amplified from MM plasma cells. Of 115 randomly picked clones, 49% were RHAMMFL, 47% were RHAMM-48, and 4% were RHAMM-147. All of the detected RHAMM variants contain exon 4, which is alternatively spliced in murine RHAMM, and had only a single copy of the exon 8 repeat sequence detected in murine RHAMM. RT-PCR analysis of sorted blood or BM cells from 22 MM patients showed that overexpression of RHAMM variants is characteristic of MM B cells and BM plasma cells in all patients tested. RHAMM also appeared to be overexpressed in B lymphoma and B-chronic lymphocytic leukemia (CLL) cells. In B cells from normal donors, RHAMMFL was only weakly detectable in resting B cells from five of eight normal donors or in chronically activated B cells from three patients with Crohn's disease. RHAMM-48 was detectable in B cells from one of eight normal donors, but was undetectable in B cells of three donors with Crohn's disease. RHAMM-147 was undetectable in normal and Crohn's disease B cells. In situ RT-PCR was used to determine the number of individual cells with aggregate RHAMM transcripts. For six patients, 29% of BM plasma cells and 12% of MM B cells had detectable RHAMM transcripts, while for five normal donors, only 1. 2% of B cells expressed RHAMM transcripts. This work suggests that RHAMMFL, RHAMM-48, and RHAMM-147 splice variants are overexpressed in MM and other B lymphocyte malignancies relative to resting or in vivo-activated B cells, raising the possibility that RHAMM and its variants may contribute to the malignant process in B-cell malignancies such as lymphoma, CLL, and MM.

B-Lymphocytes↗

Gene structure prediction from consensus spliced alignment of multiple ESTs matching the same genomic locus.

MOTIVATION: Accurate gene structure annotation is a challenging computational problem in genomics. The best results are achieved with spliced alignment of full-length cDNAs or multiple expressed sequence tags (ESTs) with sufficient overlap to cover the entire gene. For most species, cDNA and EST collections are far from comprehensive. We sought to overcome this bottleneck by exploring the possibility of using combined EST resources from fairly diverged species that still share a common gene space. Previous spliced alignment tools were found inadequate for this task because they rely on very high sequence similarity between the ESTs and the genomic DNA. RESULTS: We have developed a computer program, GeneSeqer, which is capable of aligning thousands of ESTs with a long genomic sequence in a reasonable amount of time. The algorithm is uniquely designed to tolerate a high percentage of mismatches and insertions or deletions in the EST relative to the genomic template. This feature allows use of non-cognate ESTs for gene structure prediction, including ESTs derived from duplicated genes and homologous genes from related species. The increased gene prediction sensitivity results in part from novel splice site prediction models that are also available as a stand-alone splice site prediction tool. We assessed GeneSeqer performance relative to a standard Arabidopsis thaliana gene set and demonstrate its utility for plant genome annotation. In particular, we propose that this method provides a timely tool for the annotation of the rice genome, using abundant ESTs from other cereals and plants. AVAILABILITY: The source code is available for download at http://bioinformatics.iastate.edu/bioinformatics2go/gs/download.html. Web servers for Arabidopsis and other plant species are accessible at http://www.plantgdb.org/cgi-bin/AtGeneSeqer.cgi and http://www.plantgdb.org/cgi-bin/GeneSeqer.cgi, respectively. For non-plant species, use http://bioinformatics.iastate.edu/cgi-bin/gs.cgi. The splice site prediction tool (SplicePredictor) is distributed with the GeneSeqer code. A SplicePredictor web server is available at http://bioinformatics.iastate.edu/cgi-bin/sp.cgi

Algorithms↗

Evolution and phylogenetic utility of alignment gaps within intron sequences of three nuclear genes in bumble bees (Bombus).

To test whether gaps resulting from sequence alignment contain phylogenetic signal concordant with those of base substitutions, we analyzed the occurrence of indel mutations upon a well-resolved, substitution-based tree for three nuclear genes in bumble bees (Bombus, Apidae: Bombini). The regions analyzed were exon and intron sequences of long-wavelength rhodopsin (LW Rh), arginine kinase (ArgK), and elongation factor-1alpha (EF-1alpha) F2 copy genes. LW Rh intron had only a few uninformative gaps, ArgK intron had relatively long gaps that were easily aligned, and EF-1alpha intron had many short gaps, resulting in multiple optimal alignments. The unambiguously aligned gaps within ArgK intron sequences showed no homoplasy upon the substitution-based tree, and phylogenetic signals within ambiguously aligned regions of EF-1alpha intron were highly congruent with those of base substitutions. We further analyzed the contribution of gap characters to phylogenetic reconstruction by incorporating them in parsimony analysis. Inclusion of gap characters consistently improved support for nodes recovered by substitutions, and inclusion of ambiguously aligned regions of EF-1alpha intron resolved several additional nodes, most of which were apical on the phylogeny. We conclude that gaps are an exceptionally reliable source of phylogenetic information that can be used to corroborate and refine phylogenies hypothesized by base substitutions, at least at lower taxonomic levels. At present, full use of gaps in phylogenetic reconstruction is best achieved in parsimony analysis, pending development of well-justified and generally applicable methods for incorporating indels in explicitly model-based methods.

Animals↗

The three-dimensional structure of human transaldolase.

The crystal structure of human transaldolase has been determined to 2.45 A resolution. The enzyme folds into an alpha/beta barrel structure and is thus similar in structure to other class I aldolases. Structure-based sequence alignment of available sequences of the transaldolase subfamily reveals that eight active site residues are invariant in the whole subfamily. Other invariant residues are mainly involved in the formation of the hydrophobic core of the enzyme. Noteworthy is a hydrophobic cluster consisting of five invariant residues. Human transaldolase has been implicated as an autoantigen in multiple sclerosis and four immunodominant peptide segments are located at the surface of the enzyme, accessible to autoantibodies.

Amino Acid Sequence↗

Identification of mycobacterial species by comparative sequence analysis of the RNA polymerase gene (rpoB).

For the differentiation and identification of mycobacterial species, the rpoB gene, encoding the beta subunit of RNA polymerase, was investigated. rpoB DNAs (342 bp) were amplified from 44 reference strains of mycobacteria and clinical isolates (107 strains) by PCR. The nucleotide sequences were directly determined (306 bp) and aligned by using the multiple alignment algorithm in the MegAlign package (DNASTAR) and the MEGA program. A phylogenetic tree was constructed by the neighbor-joining method. Comparative sequence analysis of rpoB DNAs provided the basis for species differentiation within the genus Mycobacterium. Slowly and rapidly growing groups of mycobacteria were clearly separated, and each mycobacterial species was differentiated as a distinct entity in the phylogenetic tree. Pathogenic Mycobacterium kansasii was easily differentiated from nonpathogenic M. gastri; this differentiation cannot be achieved by using 16S rRNA gene (rDNA) sequences. By being grouped into species-specific clusters with low-level sequence divergence among strains of the same species, all of the clinical isolates could be easily identified. These results suggest that comparative sequence analysis of amplified rpoB DNAs can be used efficiently to identify clinical isolates of mycobacteria in parallel with traditional culture methods and as a supplement to 16S rDNA gene analysis. Furthermore, in the case of M. tuberculosis, rifampin resistance can be simultaneously determined.

Amino Acid Sequence↗

Gibbs motif sampling: detection of bacterial outer membrane protein repeats.

The detection and alignment of locally conserved regions (motifs) in multiple sequences can provide insight into protein structure, function, and evolution. A new Gibbs sampling algorithm is described that detects motif-encoding regions in sequences and optimally partitions them into distinct motif models; this is illustrated using a set of immunoglobulin fold proteins. When applied to sequences sharing a single motif, the sampler can be used to classify motif regions into related submodels, as is illustrated using helix-turn-helix DNA-binding proteins. Other statistically based procedures are described for searching a database for sequences matching motifs found by the sampler. When applied to a set of 32 very distantly related bacterial integral outer membrane proteins, the sampler revealed that they share a subtle, repetitive motif. Although BLAST (Altschul SF et al., 1990, J Mol Biol 215:403-410) fails to detect significant pairwise similarity between any of the sequences, the repeats present in these outer membrane proteins, taken as a whole, are highly significant (based on a generally applicable statistical test for motifs described here). Analysis of bacterial porins with known trimeric beta-barrel structure and related proteins reveals a similar repetitive motif corresponding to alternating membrane-spanning beta-strands. These beta-strands occur on the membrane interface (as opposed to the trimeric interface) of the beta-barrel. The broad conservation and structural location of these repeats suggests that they play important functional roles.

Algorithms↗

Molecular evolution of the genes encoding receptor tyrosine kinase with immunoglobulinlike domains.

Receptor tyrosine kinases (RTK) with five, three, or seven immunoglobulinlike domains in their extracellular regions are classified as subclasses III, IV, and V, respectively. Conservation of the exon/intron structure of the downstream part of the human KIT, FMS, and FLT3 genes that encode RTK of subclass III together with the particular chromosomal localization of these genes suggests that RTKIII genes have evolved from a common ancestor by cis and trans duplications. To strengthen this model of evolution and to determine if it can be extended to RTKIV and V genes, we constructed a phylogenetic tree of RTKIII, IV, and V on the basis of a multiple alignment of their catalytic tyrosine kinase domain sequences and determined the exon/intron structure of PDGFRA (subclass III), FGFR4 (subclass IV), and FLT4 (subclass V) genes in their downstream part. Phylogenetic analyses with amino acid or nucleotide sequences both resulted in one most parsimonious tree. The phylogenetic trees obtained indicate that all three subclasses are well individuated and that RTKIII and RTKV are closer to each other than RTKIV. Furthermore, RTKIII and FLT4 (subclass V) genes possess the same exon/intron structure in their downstream part while the structure of the RTKIV genes is very similar to that of RTKIII and FLT4. Both approaches are in complete agreement and indicate that RTKIII, IV, and V genes most probably evolved from a common ancestor already "in pieces" by successive duplications involving entire genes.

Alternative Splicing↗

Strategies for comparing gene expression profiles from different microarray platforms: application to a case-control experiment.

Meta-analysis of microarray data is increasingly important, considering both the availability of multiple platforms using disparate technologies and the accumulation in public repositories of data sets from different laboratories. We addressed the issue of comparing gene expression profiles from two microarray platforms by devising a standardized investigative strategy. We tested this procedure by studying MDA-MB-231 cells, which undergo apoptosis on treatment with resveratrol. Gene expression profiles were obtained using high-density, short-oligonucleotide, single-color microarray platforms: GeneChip (Affymetrix) and CodeLink (Amersham). Interplatform analyses were carried out on 8414 common transcripts represented on both platforms, as identified by LocusLink ID, representing 70.8% and 88.6% of annotated GeneChip and CodeLink features, respectively. We identified 105 differentially expressed genes (DEGs) on CodeLink and 42 DEGs on GeneChip. Among them, only 9 DEGs were commonly identified by both platforms. Multiple analyses (BLAST alignment of probes with target sequences, gene ontology, literature mining, and quantitative real-time PCR) permitted us to investigate the factors contributing to the generation of platform-dependent results in single-color microarray experiments. An effective approach to cross-platform comparison involves microarrays of similar technologies, samples prepared by identical methods, and a standardized battery of bioinformatic and statistical analyses.

Breast Neoplasms↗