Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

SSAHA: a fast search method for large DNA databases.

We describe an algorithm, SSAHA (Sequence Search and Alignment by Hashing Algorithm), for performing fast searches on databases containing multiple gigabases of DNA. Sequences in the database are preprocessed by breaking them into consecutive k-tuples of k contiguous bases and then using a hash table to store the position of each occurrence of each k-tuple. Searching for a query sequence in the database is done by obtaining from the hash table the "hits" for each k-tuple in the query sequence and then performing a sort on the results. We discuss the effect of the tuple length k on the search speed, memory usage, and sensitivity of the algorithm and present the results of computational experiments which show that SSAHA can be three to four orders of magnitude faster than BLAST or FASTA, while requiring less memory than suffix tree methods. The SSAHA algorithm is used for high-throughput single nucleotide polymorphism (SNP) detection and very large scale sequence assembly. Also, it provides Web-based sequence search facilities for Ensembl projects.

Algorithms↗

Cloning the human and mouse MMS19 genes and functional complementation of a yeast mms19 deletion mutant.

The MMS19 gene of the yeast Saccharomyces cerevisiae encodes a polypeptide of unknown function which is required for both nucleotide excision repair (NER) and RNA polymerase II (RNAP II) transcription. Here we report the molecular cloning of human and mouse orthologs of the yeast MMS19 gene. Both human and Drosophila MMS19 cDNAs correct thermosensitive growth and sensitivity to killing by UV radiation in a yeast mutant deleted for the MMS19 gene, indicating functional conservation between the yeast and mammalian gene products. Alignment of the translated sequences of MMS19 from multiple eukaryotes, including mouse and human, revealed the presence of several conserved regions, including a HEAT repeat domain near the C-terminus. The presence of HEAT repeats, coupled with functional complementation of yeast mutant phenotypes by the orthologous protein from higher eukaryotes, suggests a role of Mms19 protein in the assembly of a multiprotein complex(es) required for NER and RNAP II transcription. Both the mouse and human genes are ubiquitously expressed as multiple transcripts, some of which appear to derive from alternative splicing. The ratio of different transcripts varies in several different tissue types.

Alternative Splicing↗

Statistical modeling and analysis of the LAGLIDADG family of site-specific endonucleases and identification of an intein that encodes a site-specific endonuclease of the HNH family.

The LAGLIDADG and HNH families of site-specific DNA endonucleases encoded by viruses, bacteriophages as well as archaeal, eucaryotic nuclear and organellar genomes are characterized by the sequence motifs 'LAGLIDADG' and 'HNH', respectively. These endonucleases have been shown to occur in different environments: LAGLIDADG endonucleases are found in inteins, archaeal and group I introns and as free standing open reading frames (ORFs); HNH endonucleases occur in group I and group II introns and as ORFs. Here, statistical models (hidden Markov models, HMMs) that encompass both the conserved motifs and more variable regions of these families have been created and employed to characterize known and potential new family members. A number of new, putative LAGLIDADG and HNH endonucleases have been identified including an intein-encoded HNH sequence. Analysis of an HMM-generated multiple alignment of 130 LAGLIDADG family members and the three-dimensional structure of the I- Cre I endonuclease has enabled definition of the core elements of the repeated domain (approximately 90 residues) that is present in this family of proteins. A conserved negatively charged residue is proposed to be involved in catalysis. Phylogenetic analysis of the two families indicates a lack of exchange of endonucleases between different mobile elements (environments) and between hosts from different phylogenetic kingdoms. However, there does appear to have been considerable exchange of endonuclease domains amongst elements of the same type. Such events are suggested to be important for the formation of elements of new specficity.

Amino Acid Sequence↗

Genome-wide detection and analysis of homologous recombination among sequenced strains of Escherichia coli.

BACKGROUND: Comparisons of complete bacterial genomes reveal evidence of lateral transfer of DNA across otherwise clonally diverging lineages. Some lateral transfer events result in acquisition of novel genomic segments and are easily detected through genome comparison. Other more subtle lateral transfers involve homologous recombination events that result in substitution of alleles within conserved genomic regions. This type of event is observed infrequently among distantly related organisms. It is reported to be more common within species, but the frequency has been difficult to quantify since the sequences under comparison tend to have relatively few polymorphic sites. RESULTS: Here we report a genome-wide assessment of homologous recombination among a collection of six complete Escherichia coli and Shigella flexneri genome sequences. We construct a whole-genome multiple alignment and identify clusters of polymorphic sites that exhibit atypical patterns of nucleotide substitution using a random walk-based method. The analysis reveals one large segment (approximately 100 kb) and 186 smaller clusters of single base pair differences that suggest lateral exchange between lineages. These clusters include portions of 10% of the 3,100 genes conserved in six genomes. Statistical analysis of the functional roles of these genes reveals that several classes of genes are over-represented, including those involved in recombination, transport and motility. CONCLUSION: We demonstrate that intraspecific recombination in E. coli is much more common than previously appreciated and may show a bias for certain types of genes. The described method provides high-specificity, conservative inference of past recombination events.

Alleles↗

Redundancy based detection of sequence polymorphisms in expressed sequence tag data using autoSNP.

UNLABELLED: AutoSNP is a program to detect single nucleotide polymorphisms (SNPs) and insertion/deletion polymorphisms (indels) in expressed sequence tag (EST) data. The program uses d2cluster and cap3 to cluster and align EST sequences, and uses redundancy to differentiate between candidate SNPs and sequence errors. Candidate polymorphisms are identified as occurring in multiple reads within an alignment. For each candidate SNP, two measures of confidence are calculated, the redundancy of the polymorphism at a SNP locus and the co segregation of the candidate SNP with other SNPs in the alignment. AVAILABILITY: The program was written in PERL and is freely available to non-commercial users by request from the authors.

Cluster Analysis↗

Overexpression of the receptor for hyaluronan-mediated motility (RHAMM) characterizes the malignant clone in multiple myeloma: identification of three distinct RHAMM variants.

The receptor for hyaluronan (HA)-mediated motility (RHAMM) controls motility by malignant cells in myeloma and is abnormally expressed on the surface of most malignant B and plasma cells in blood or bone marrow (BM) of patients with multiple myeloma (MM). RHAMM cDNA was cloned and sequenced from the malignant B and plasma cells comprising the myeloma B lineage hierarchy. Three distinct RHAMM gene products, RHAMMFL, RHAMM-48, and RHAMM-147, were cloned from MM B and plasma cells. RHAMMFL was 99% homologous to the published sequence of RHAMM. RHAMM-48 and RHAMM-147 variants align with RHAMMFL, but are characterized by sequence deletions of 48 bp (16 amino acids [aa]) and 147 bp (49 aa), respectively. The relative frequency of these RHAMM transcripts in MM plasma cells was determined by cloning of reverse-transcriptase polymerase chain reaction (RT-PCR) products amplified from MM plasma cells. Of 115 randomly picked clones, 49% were RHAMMFL, 47% were RHAMM-48, and 4% were RHAMM-147. All of the detected RHAMM variants contain exon 4, which is alternatively spliced in murine RHAMM, and had only a single copy of the exon 8 repeat sequence detected in murine RHAMM. RT-PCR analysis of sorted blood or BM cells from 22 MM patients showed that overexpression of RHAMM variants is characteristic of MM B cells and BM plasma cells in all patients tested. RHAMM also appeared to be overexpressed in B lymphoma and B-chronic lymphocytic leukemia (CLL) cells. In B cells from normal donors, RHAMMFL was only weakly detectable in resting B cells from five of eight normal donors or in chronically activated B cells from three patients with Crohn's disease. RHAMM-48 was detectable in B cells from one of eight normal donors, but was undetectable in B cells of three donors with Crohn's disease. RHAMM-147 was undetectable in normal and Crohn's disease B cells. In situ RT-PCR was used to determine the number of individual cells with aggregate RHAMM transcripts. For six patients, 29% of BM plasma cells and 12% of MM B cells had detectable RHAMM transcripts, while for five normal donors, only 1. 2% of B cells expressed RHAMM transcripts. This work suggests that RHAMMFL, RHAMM-48, and RHAMM-147 splice variants are overexpressed in MM and other B lymphocyte malignancies relative to resting or in vivo-activated B cells, raising the possibility that RHAMM and its variants may contribute to the malignant process in B-cell malignancies such as lymphoma, CLL, and MM.

B-Lymphocytes↗

Gene structure prediction from consensus spliced alignment of multiple ESTs matching the same genomic locus.

MOTIVATION: Accurate gene structure annotation is a challenging computational problem in genomics. The best results are achieved with spliced alignment of full-length cDNAs or multiple expressed sequence tags (ESTs) with sufficient overlap to cover the entire gene. For most species, cDNA and EST collections are far from comprehensive. We sought to overcome this bottleneck by exploring the possibility of using combined EST resources from fairly diverged species that still share a common gene space. Previous spliced alignment tools were found inadequate for this task because they rely on very high sequence similarity between the ESTs and the genomic DNA. RESULTS: We have developed a computer program, GeneSeqer, which is capable of aligning thousands of ESTs with a long genomic sequence in a reasonable amount of time. The algorithm is uniquely designed to tolerate a high percentage of mismatches and insertions or deletions in the EST relative to the genomic template. This feature allows use of non-cognate ESTs for gene structure prediction, including ESTs derived from duplicated genes and homologous genes from related species. The increased gene prediction sensitivity results in part from novel splice site prediction models that are also available as a stand-alone splice site prediction tool. We assessed GeneSeqer performance relative to a standard Arabidopsis thaliana gene set and demonstrate its utility for plant genome annotation. In particular, we propose that this method provides a timely tool for the annotation of the rice genome, using abundant ESTs from other cereals and plants. AVAILABILITY: The source code is available for download at http://bioinformatics.iastate.edu/bioinformatics2go/gs/download.html. Web servers for Arabidopsis and other plant species are accessible at http://www.plantgdb.org/cgi-bin/AtGeneSeqer.cgi and http://www.plantgdb.org/cgi-bin/GeneSeqer.cgi, respectively. For non-plant species, use http://bioinformatics.iastate.edu/cgi-bin/gs.cgi. The splice site prediction tool (SplicePredictor) is distributed with the GeneSeqer code. A SplicePredictor web server is available at http://bioinformatics.iastate.edu/cgi-bin/sp.cgi

Algorithms↗

Evolution and phylogenetic utility of alignment gaps within intron sequences of three nuclear genes in bumble bees (Bombus).

To test whether gaps resulting from sequence alignment contain phylogenetic signal concordant with those of base substitutions, we analyzed the occurrence of indel mutations upon a well-resolved, substitution-based tree for three nuclear genes in bumble bees (Bombus, Apidae: Bombini). The regions analyzed were exon and intron sequences of long-wavelength rhodopsin (LW Rh), arginine kinase (ArgK), and elongation factor-1alpha (EF-1alpha) F2 copy genes. LW Rh intron had only a few uninformative gaps, ArgK intron had relatively long gaps that were easily aligned, and EF-1alpha intron had many short gaps, resulting in multiple optimal alignments. The unambiguously aligned gaps within ArgK intron sequences showed no homoplasy upon the substitution-based tree, and phylogenetic signals within ambiguously aligned regions of EF-1alpha intron were highly congruent with those of base substitutions. We further analyzed the contribution of gap characters to phylogenetic reconstruction by incorporating them in parsimony analysis. Inclusion of gap characters consistently improved support for nodes recovered by substitutions, and inclusion of ambiguously aligned regions of EF-1alpha intron resolved several additional nodes, most of which were apical on the phylogeny. We conclude that gaps are an exceptionally reliable source of phylogenetic information that can be used to corroborate and refine phylogenies hypothesized by base substitutions, at least at lower taxonomic levels. At present, full use of gaps in phylogenetic reconstruction is best achieved in parsimony analysis, pending development of well-justified and generally applicable methods for incorporating indels in explicitly model-based methods.

Animals↗

The three-dimensional structure of human transaldolase.

The crystal structure of human transaldolase has been determined to 2.45 A resolution. The enzyme folds into an alpha/beta barrel structure and is thus similar in structure to other class I aldolases. Structure-based sequence alignment of available sequences of the transaldolase subfamily reveals that eight active site residues are invariant in the whole subfamily. Other invariant residues are mainly involved in the formation of the hydrophobic core of the enzyme. Noteworthy is a hydrophobic cluster consisting of five invariant residues. Human transaldolase has been implicated as an autoantigen in multiple sclerosis and four immunodominant peptide segments are located at the surface of the enzyme, accessible to autoantibodies.

Amino Acid Sequence↗

Identification of mycobacterial species by comparative sequence analysis of the RNA polymerase gene (rpoB).

For the differentiation and identification of mycobacterial species, the rpoB gene, encoding the beta subunit of RNA polymerase, was investigated. rpoB DNAs (342 bp) were amplified from 44 reference strains of mycobacteria and clinical isolates (107 strains) by PCR. The nucleotide sequences were directly determined (306 bp) and aligned by using the multiple alignment algorithm in the MegAlign package (DNASTAR) and the MEGA program. A phylogenetic tree was constructed by the neighbor-joining method. Comparative sequence analysis of rpoB DNAs provided the basis for species differentiation within the genus Mycobacterium. Slowly and rapidly growing groups of mycobacteria were clearly separated, and each mycobacterial species was differentiated as a distinct entity in the phylogenetic tree. Pathogenic Mycobacterium kansasii was easily differentiated from nonpathogenic M. gastri; this differentiation cannot be achieved by using 16S rRNA gene (rDNA) sequences. By being grouped into species-specific clusters with low-level sequence divergence among strains of the same species, all of the clinical isolates could be easily identified. These results suggest that comparative sequence analysis of amplified rpoB DNAs can be used efficiently to identify clinical isolates of mycobacteria in parallel with traditional culture methods and as a supplement to 16S rDNA gene analysis. Furthermore, in the case of M. tuberculosis, rifampin resistance can be simultaneously determined.

Amino Acid Sequence↗

Gibbs motif sampling: detection of bacterial outer membrane protein repeats.

The detection and alignment of locally conserved regions (motifs) in multiple sequences can provide insight into protein structure, function, and evolution. A new Gibbs sampling algorithm is described that detects motif-encoding regions in sequences and optimally partitions them into distinct motif models; this is illustrated using a set of immunoglobulin fold proteins. When applied to sequences sharing a single motif, the sampler can be used to classify motif regions into related submodels, as is illustrated using helix-turn-helix DNA-binding proteins. Other statistically based procedures are described for searching a database for sequences matching motifs found by the sampler. When applied to a set of 32 very distantly related bacterial integral outer membrane proteins, the sampler revealed that they share a subtle, repetitive motif. Although BLAST (Altschul SF et al., 1990, J Mol Biol 215:403-410) fails to detect significant pairwise similarity between any of the sequences, the repeats present in these outer membrane proteins, taken as a whole, are highly significant (based on a generally applicable statistical test for motifs described here). Analysis of bacterial porins with known trimeric beta-barrel structure and related proteins reveals a similar repetitive motif corresponding to alternating membrane-spanning beta-strands. These beta-strands occur on the membrane interface (as opposed to the trimeric interface) of the beta-barrel. The broad conservation and structural location of these repeats suggests that they play important functional roles.

Algorithms↗

Molecular evolution of the genes encoding receptor tyrosine kinase with immunoglobulinlike domains.

Receptor tyrosine kinases (RTK) with five, three, or seven immunoglobulinlike domains in their extracellular regions are classified as subclasses III, IV, and V, respectively. Conservation of the exon/intron structure of the downstream part of the human KIT, FMS, and FLT3 genes that encode RTK of subclass III together with the particular chromosomal localization of these genes suggests that RTKIII genes have evolved from a common ancestor by cis and trans duplications. To strengthen this model of evolution and to determine if it can be extended to RTKIV and V genes, we constructed a phylogenetic tree of RTKIII, IV, and V on the basis of a multiple alignment of their catalytic tyrosine kinase domain sequences and determined the exon/intron structure of PDGFRA (subclass III), FGFR4 (subclass IV), and FLT4 (subclass V) genes in their downstream part. Phylogenetic analyses with amino acid or nucleotide sequences both resulted in one most parsimonious tree. The phylogenetic trees obtained indicate that all three subclasses are well individuated and that RTKIII and RTKV are closer to each other than RTKIV. Furthermore, RTKIII and FLT4 (subclass V) genes possess the same exon/intron structure in their downstream part while the structure of the RTKIV genes is very similar to that of RTKIII and FLT4. Both approaches are in complete agreement and indicate that RTKIII, IV, and V genes most probably evolved from a common ancestor already "in pieces" by successive duplications involving entire genes.

Alternative Splicing↗

Strategies for comparing gene expression profiles from different microarray platforms: application to a case-control experiment.

Meta-analysis of microarray data is increasingly important, considering both the availability of multiple platforms using disparate technologies and the accumulation in public repositories of data sets from different laboratories. We addressed the issue of comparing gene expression profiles from two microarray platforms by devising a standardized investigative strategy. We tested this procedure by studying MDA-MB-231 cells, which undergo apoptosis on treatment with resveratrol. Gene expression profiles were obtained using high-density, short-oligonucleotide, single-color microarray platforms: GeneChip (Affymetrix) and CodeLink (Amersham). Interplatform analyses were carried out on 8414 common transcripts represented on both platforms, as identified by LocusLink ID, representing 70.8% and 88.6% of annotated GeneChip and CodeLink features, respectively. We identified 105 differentially expressed genes (DEGs) on CodeLink and 42 DEGs on GeneChip. Among them, only 9 DEGs were commonly identified by both platforms. Multiple analyses (BLAST alignment of probes with target sequences, gene ontology, literature mining, and quantitative real-time PCR) permitted us to investigate the factors contributing to the generation of platform-dependent results in single-color microarray experiments. An effective approach to cross-platform comparison involves microarrays of similar technologies, samples prepared by identical methods, and a standardized battery of bioinformatic and statistical analyses.

Breast Neoplasms↗

Structural and functional differences between 3-repeat and 4-repeat tau isoforms. Implications for normal tau function and the onset of neurodegenetative disease.

Tau, MAP2, and MAP4 are members of a microtubule-associated protein (MAP) family that are each expressed as "3-repeat" and "4-repeat" isoforms. These isoforms arise from tightly controlled tissue-specific and/or developmentally regulated alternative splicing of a 31-amino acid long "inter-repeat:repeat module," raising the possibility that different MAP isoforms may possess some distinct functional capabilities. Consistent with this hypothesis, regulatory mutations in the human tau gene that disrupt the normal balance between 3-repeat and 4-repeat tau isoform expression lead to a collection of neurodegenerative diseases known as FTDP-17 (fronto-temporal dementias and Parkinsonism linked to chromosome 17), which are characterized by the formation of pathological tau filaments and neuronal cell death. Unfortunately, very little is known regarding structural and functional differences between the isoforms. In our previous analyses, we focused on 4-repeat tau structure and function. Here, we investigate 3-repeat tau, generating a series of truncations, amino acid substitutions, and internal deletions and examining the functional consequences. 3-Repeat tau possesses a "core microtubule binding domain" composed of its first two repeats and the intervening inter-repeat. This observation is in marked contrast to the widely held notion that tau possesses multiple independent tubulin-binding sites aligned in sequence along the length of the protein. In addition, we observed that the carboxyl-terminal sequences downstream of the repeat region make a strong but indirect contribution to microtubule binding activity in 3-repeat tau, which is in contrast to the negligible effect of these same sequences in 4-repeat tau. Taken together with previous work, these data suggest that 3-repeat and 4-repeat tau assume complex and distinct structures that are regulated differentially, which in turn suggests that they may possess isoform-specific functional capabilities. The relevance of isoform-specific structure and function to normal tau action and the onset of neurodegenerative disease are discussed.

Alternative Splicing↗

PairWise and SearchWise: finding the optimal alignment in a simultaneous comparison of a protein profile against all DNA translation frames.

DNA translation frames can be disrupted for several reasons, including: (i) errors in sequence determination; (ii) RNA processing, such as intron removal and guide RNA editing; (iii) less commonly, polymerase frameshifting during transcription or ribosomal frameshifting during translation. Frameshifts frequently confound computational activities involving homologous sequences, such as database searches and inferences on structure, function or phylogeny made from multiple alignments. A dynamic alignment algorithm is reported here which compares a protein profile (a residue scoring matrix for one or more aligned sequences) against the three translation frames of a DNA strand, allowing frameshifting. The algorithm has been incorporated into a new package, WiseTools, for comparison of biological sequences. A protein profile can be compared against either a DNA sequence or a protein sequence. The program PairWise may be used interactively for alignment of any two sequence inputs. SearchWise can perform combinations of searches through DNA or protein databases by a protein profile or DNA sequence. Routine application of the programs has revealed a set of database entries with frameshifts caused by errors in sequence determination.

Algorithms↗

Phylogenetic internal control for HIV-1 genotypic antiretroviral testing.

Genotypic testing includes several steps (RNA purification, RT-PCR amplification, DNA sequencing, sequence editing and analysis) that should be individually controlled. In our laboratory, we have added to this step-by-step internal control a final phylogenetic quality control: this is performed every time a sequence is obtained from a patient previously subjected to the same test. Each sequence with this characteristic is routinely compared with sequences from previous samples of the same patient by multiple alignment and a neighbor-joining tree by using Kimura two-parameter method is constructed. To validate the quality control procedure, we have aligned and calculated the mean similarity of the reverse transcriptase (first 984 nucleotides) and protease (whole gene) sequences from 30 patients whose virus was completely wild-type for both reverse transcriptase and protease. In the same tree, we have added the sequences obtained from 5 out of the 30 patients, tested at a second time point. The wild type sequences have shown a mean inter-sample divergence of 2.9%, and all the sequence pairs from individual patients clustered together in the tree constructed with the nucleotide sequences, while the tree constructed with the inferred aminoacid sequences did not always permit to cluster the sequences from the same patients. This indicates that: 1) the phylogenetic analysis of nucleic acid sequences can be useful to rule out sample mix-up; 2) the belonging of a sequence to each individual patient can efficiently be assessed also in the cases of extreme divergence in terms of drug resistance mutations.

Amino Acid Sequence↗

Scoredist: a simple and robust protein sequence distance estimator.

BACKGROUND: Distance-based methods are popular for reconstructing evolutionary trees thanks to their speed and generality. A number of methods exist for estimating distances from sequence alignments, which often involves some sort of correction for multiple substitutions. The problem is to accurately estimate the number of true substitutions given an observed alignment. So far, the most accurate protein distance estimators have looked for the optimal matrix in a series of transition probability matrices, e.g. the Dayhoff series. The evolutionary distance between two aligned sequences is here estimated as the evolutionary distance of the optimal matrix. The optimal matrix can be found either by an iterative search for the Maximum Likelihood matrix, or by integration to find the Expected Distance. As a consequence, these methods are more complex to implement and computationally heavier than correction-based methods. Another problem is that the result may vary substantially depending on the evolutionary model used for the matrices. An ideal distance estimator should produce consistent and accurate distances independent of the evolutionary model used. RESULTS: We propose a correction-based protein sequence estimator called Scoredist. It uses a logarithmic correction of observed divergence based on the alignment score according to the BLOSUM62 score matrix. We evaluated Scoredist and a number of optimal matrix methods using three evolutionary models for both training and testing Dayhoff, Jones-Taylor-Thornton, and Muller-Vingron, as well as Whelan and Goldman solely for testing. Test alignments with known distances between 0.01 and 2 substitutions per position (1-200 PAM) were simulated using ROSE. Scoredist proved as accurate as the optimal matrix methods, yet substantially more robust. When trained on one model but tested on another one, Scoredist was nearly always more accurate. The Jukes-Cantor and Kimura correction methods were also tested, but were substantially less accurate. CONCLUSION: The Scoredist distance estimator is fast to implement and run, and combines robustness with accuracy. Scoredist has been incorporated into the Belvu alignment viewer, which is available at ftp://ftp.cgb.ki.se/pub/prog/belvu/.

Algorithms↗