Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

The VPH1 gene encodes a 95-kDa integral membrane polypeptide required for in vivo assembly and activity of the yeast vacuolar H(+)-ATPase.

Yeast vacuolar acidification-defective (vph) mutants were identified using the pH-sensitive fluorescence of 6-carboxyfluorescein diacetate (Preston, R. A., Murphy, R. F., and Jones, E. W. (1989) Proc. Natl. Acad. Sci. U.S.A. 86, 7027-7031). Vacuoles purified from yeast bearing the vph1-1 mutation had no detectable bafilomycin-sensitive ATPase activity or ATP-dependent proton pumping. The peripherally bound nucleotide-binding subunits of the vacuolar H(+)-ATPase (60 and 69 kDa) were no longer associated with vacuolar membranes yet were present in wild type levels in yeast whole cell extracts. The VPH1 gene was cloned by complementation of the vph1-1 mutation and independently cloned by screening a lambda gt11 expression library with antibodies directed against a 95-kDa vacuolar integral membrane protein. Deletion disruption of the VPH1 gene revealed that the VPH1 gene is not essential for viability but is required for vacuolar H(+)-ATPase assembly and vacuolar acidification. VPH1 encodes a predicted polypeptide of 840 amino acid residues (molecular mass 95.6 kDa) and contains six putative membrane-spanning regions. Cell fractionation and immunodetection demonstrate that Vph1p is a vacuolar integral membrane protein that co-purifies with vacuolar H(+)-ATPase activity. Multiple sequence alignments show extensive homology over the entire lengths of the following four polypeptides: Vph1p, the 116-kDa polypeptide of the rat clathrin-coated vesicles/synaptic vesicle proton pump, the predicted polypeptide encoded by the yeast gene STV1 (Similar To VPH1, identified as an open reading frame next to the BUB2 gene), and the TJ6 mouse immune suppressor factor.

Amino Acid Sequence

Computational sequence analysis revisited: new databases, software tools, and the research opportunities they engender.

The increasing quantity and complexity of sequences and structural data for proteins and nucleic acids create both problems and opportunities for biomedical researchers. Fortunately, a new generation of practical computer tools for data analysis and integrated information retrieval is emerging. Recent developments in fast database searching, multiple sequence alignment, and molecular modeling are discussed and windows-based, mouse-driven software for CD-ROM and network information retrieval are described. Each method is illustrated with a practical example pertinent to lipid research. In particular, the connection among cholesteryl ester transfer protein, bactericidal permeability-increasing protein, and lipopolysaccharide-binding proteins is determined; novel repetitive sequence motifs in mammalian farnesyltransferase subunits and related yeast prenyltransferases are derived; biochemical insights from a three-dimensional model of human apolipoprotein D based on two insect lipocalins are discussed; the relationship between apolipoprotein D and gross cystic disease fluid protein from human breast is reviewed; and prospects for modeling apolipoprotein E-related proteins are described. In addition, information on a number of general and special-purpose sequence, motif, and structural databases is included.

Amino Acid Sequence

The megaprior heuristic for discovering protein sequence patterns.

Several computer algorithms for discovering patterns in groups of protein sequences are in use that are based on fitting the parameters of a statistical model to a group of related sequences. These include hidden Markov model (HMM) algorithms for multiple sequence alignment, and the MEME and Gibbs sampler algorithms for discovering motifs. These algorithms are sometimes prone to producing models that are incorrect because two or more patients have been combined. The statistical model produced in this situation is a convex combination (weighted average) of two or more different models. This paper presents a solution to the problem of convex combinations in the form of a heuristic based on using extremely low variance Dirichlet mixture priors as part of the statistical model. This heuristic, which we call the megaprior heuristic, increase the strength (i.e., decreases the variance) of the prior in proportion to the size of the sequence dataset. This causes each column in the final model to strongly resemble the mean of a single component of the prior, regardless of the size of the dataset. We describe the cause of the convex combination problem, analyze it mathematically, motivate and describe the implementation of the megaprior heuristic, and show how it can effectively eliminate the problem of convex combinations in protein sequence pattern discovery.

Algorithms

Multiple DNA and protein sequence alignment on a workstation and a supercomputer.

This paper describes a multiple alignment method using a workstation and supercomputer. The method is based on the alignment of a set of aligned sequences with the new sequence, and uses a recursive procedure of such alignment. The alignment is executed in a reasonable computation time on diverse levels from a workstation to a supercomputer, from the viewpoint of alignment results and computational speed by parallel processing. The application of the algorithm is illustrated by several examples of multiple alignment of 12 amino acid and DNA sequences of HIV (human immunodeficiency virus) env genes. Colour graphic programs on a workstation and parallel processing on a supercomputer are discussed.

Algorithms

Protein sequence alignments: a strategy for the hierarchical analysis of residue conservation.

An algorithm is described for the systematic characterization of the physico-chemical properties seen at each position in a multiple protein sequence alignment. The new algorithm allows questions important in the design of mutagenesis experiments to be quickly answered since positions in the alignment that show unusual or interesting residue substitution patterns may be rapidly identified. The strategy is based on a flexible set-based description of amino acid properties, which is used to define the conservation between any group of amino acids. Sequences in the alignment are gathered into subgroups on the basis of sequence similarity, functional, evolutionary or other criteria. All pairs of subgroups are then compared to highlight positions that confer the unique features of each subgroup. The algorithm is encoded in the computer program AMAS (Analysis of Multiply Aligned Sequences) which provides a textual summary of the analysis and an annotated (boxed, shaded and/or coloured) multiple sequence alignment. The algorithm is illustrated by application to an alignment of 67 SH2 domains where patterns of conserved hydrophobic residues that constitute the protein core are highlighted. The analysis of charge conservation across annexin domains identifies the locations at which conserved charges change sign. The algorithm simplifies the analysis of multiple sequence data by condensing the mass of information present, and thus allows the rapid identification of substitutions of structural and functional importance.

Algorithms

Multiple DNA and protein sequence alignment based on segment-to-segment comparison.

In this paper, a new way to think about, and to construct, pairwise as well as multiple alignments of DNA and protein sequences is proposed. Rather than forcing alignments to either align single residues or to introduce gaps by defining an alignment as a path running right from the source up to the sink in the associated dot-matrix diagram, we propose to consider alignments as consistent equivalence relations defined on the set of all positions occurring in all sequences under consideration. We also propose constructing alignments from whole segments exhibiting highly significant overall similarity rather than by aligning individual residues. Consequently, we present an alignment algorithm that (i) is based on segment-to-segment comparison instead of the commonly used residue-to-residue comparison and which (ii) avoids the well-known difficulties concerning the choice of appropriate gap penalties: gaps are not treated explicity, but remain as those parts of the sequences that do not belong to any of the aligned segments. Finally, we discuss the application of our algorithm to two test examples and compare it with commonly used alignment methods. As a first example, we aligned a set of 11 DNA sequences coding for functional helix-loop-helix proteins. Though the sequences show only low overall similarity, our program correctly aligned all of the 11 functional sites, which was a unique result among the methods tested. As a by-product, the reading frames of the sequences were identified. Next, we aligned a set of ribonuclease H proteins and compared our results with alignments produced by other programs as reported by McClure et al. [McClure, M. A., Vasi, T. K. & Fitch, W. M. (1994) Mol. Biol. Evol. 11, 571-592]. Our program was one of the best scoring programs. However, in contrast to other methods, our protein alignments are independent of user-defined parameters.

Algorithms

Analysis of the nucleotide and derived amino acid sequences of the SsoII restriction endonuclease and methyltransferase.

A 2648-bp fragment from the P4 plasmid of Shigella sonnei strain 47 coding for the SsoII restriction endonuclease (ENase) and methyltransferase (MTase) (recognition sequence 5'-CCNGG) was sequenced. Two divergently arranged open reading frames of 905 bp for the SsoII ENase (R.SsoII) and 1137 bp for the MTase (M.SsoII) were identified. The coding regions are separated by 110 bp. The calculated M(r) of R.SsoII (35937) and M.SsoII (42887) are in good agreement with values previously obtained by in vitro transcription-translation experiments, i.e., 35 and 43 kDa for the ENase and MTase, respectively. The M.SsoII amino acid (aa) sequence revealed a considerable similarity to m5C-MTases recognizing the related sequences--M.EcoRII, M.dcm, M.MspI, M.BsuFI, M.HpaII, and M.HhaI. Surprisingly, the greatest degree of homology has been observed between the aa sequences of M.SsoII and M.NlaX, with an unidentified recognition sequence. The multiple alignment of aa sequences helps to identify the blocks of conserved aa in variable regions of MTases. These conserved aa can play a key role in target recognition. Some aspects of evolution of m5C-MTases are discussed.

Amino Acid Sequence

Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.

A wealth of protein and DNA sequence data is being generated by genome projects and other sequencing efforts. A crucial barrier to deciphering these sequences and understanding the relations among them is the difficulty of detecting subtle local residue patterns common to multiple sequences. Such patterns frequently reflect similar molecular structures and biological properties. A mathematical definition of this "local multiple alignment" problem suitable for full computer automation has been used to develop a new and sensitive algorithm, based on the statistical method of iterative sampling. This algorithm finds an optimized local alignment model for N sequences in N-linear time, requiring only seconds on current workstations, and allows the simultaneous detection and optimization of multiple patterns and pattern repeats. The method is illustrated as applied to helix-turn-helix proteins, lipocalins, and prenyltransferases.

Algorithms

A method to recognize distant repeats in protein sequences.

An automated algorithm is presented that delineates protein sequence fragments which display similarity. The method incorporates a selection of a number of local nonoverlapping sequence alignments with the highest similarity scores and a graph-theoretical approach to elucidate the consistent start and end points of the fragments comprising one or more ensembles of related subsequences. The procedure allows the simultaneous identification of different types of repeats within one sequence. A multiple alignment of the resulting fragments is performed and a consensus sequence derived from the ensemble(s). Finally, a profile is constructed from the multiple alignment to detect possible and more distant members within the sequence. The method tolerates mutations in the repeats as well as insertions and deletions. The sequence spans between the various repeats or repeat clusters may be of different lengths. The technique has been applied to a number of proteins where the repeating fragments have been derived from information additional to the protein sequences.

Algorithms

Structure-based multiple alignment of extracellular pectate lyase sequences.

Pectate lyases are secreted virulence factors which degrade the pectate component of plant cell walls. The evolutionary-based multiple alignment of extracellular pectate lyases has been corrected using three-dimensional structural information derived from Erwinia chyrsanthemi pectate lyases C and E. The new multiple alignment reveals invariant amino acids likely to be involved in two different enzymatic functions.

Amino Acid Sequence

The orotidine-5'-monophosphate decarboxylase gene of Myxococcus xanthus. Comparison to the OMP decarboxylase gene family.

The nucleotide sequence of the Myxococcus xanthus orotidine-5'-monophosphate decarboxylase (OMP DCase) gene was determined. The derived protein sequence is not closely related to other prokaryotic OMP DCase sequences; nor is it closely related to any eukaryotic OMP DCase sequences. Progressive multiple alignment of the M. xanthus OMP DCase protein sequence with 19 other OMP DCase sequences revealed four conserved regions present in all 20 sequences. Ten entirely conserved residues were found in these four regions and one region contains a tight cluster of 5 conserved residues, certain of which may be catalytically active residues. A second open reading frame was found upstream of uraA and oriented in the same direction as uraA. A stretch of 21 consecutive pyrimidine (C or T) residues were found in the intercistronic region between the potential ribosome-binding site of uraA and the UGA stop codon of the upstream open reading frame. RNA directly upstream of the pyrimidine run, including the UGA stop codon of the upstream open reading frame, could be folded into a stable hairpin structure resembling Rho-independent terminators of Escherichia coli. Expression of the uraA gene may be regulated by an intercistronic transcription termination mechanism.

Amino Acid Sequence

Evolutionary divergence plots of homologous proteins.

A simple and efficient method is described for analyzing quantitatively multiple protein sequence alignments and finding the most conserved blocks as well as the maxima of divergence within the set of aligned sequences. It consists of calculating the mean distance and the root-mean-square distance in each column of the multiple alignment, averaging the values in a window of defined length and plotting the results as a function of the position of the window. Due attention is paid to the presence of gaps in the columns. Several examples are provided, using the sequences of several cytochromes c, serine proteases, lysozymes and globins. Two distance matrices are compared, namely the matrix derived by Gribskov and Burgess from the Dayhoff matrix, and the Risler Structural Superposition Matrix. In each case, the divergence plots effectively point to the specific residues which are known to be essential for the catalytic activity of the proteins. In addition, the regions of maximum divergence are clearly delineated. Interestingly, they are generally observed in positions immediately flanking the most conserved blocks. The method should therefore be useful for delineating the peptide segments which will be good candidates for site-directed mutagenesis and for visualizing the evolutionary constraints along homologous polypeptide chains.

Amino Acid Sequence

Phylogenetic relationships among megabats, microbats, and primates.

We present 744 nucleotide base positions from the mitochondrial 12S rRNA gene and 236 base positions from the mitochondrial cytochrome oxidase subunit I gene for a microbat, Brachyphylla cavernarum, and a megabat, Pteropus capestratus, in phylogenetic analyses with homologous DNA sequences from Homo sapiens, Mus musculus (house mouse), and Gallus gallus (chicken). We use information on evolutionary rate differences for different types of sequence change to establish phylogenetic character weights, and we consider alternative rRNA alignment strategies in finding that this mtDNA data set clearly supports bat monophyly. This result is found despite variations in outgroup used, gap coding scheme, and order of input for DNA sequences in multiple alignment bouts. These findings are congruent with morphological characters including details of wing structure as well as cladistic analyses of amino acid sequences for three globin genes and indicate that neurological similarities between megabats and primates are due to either retention of primitive characters or to convergent evolution rather than to inheritance from a common ancestor. This finding also indicates a single origin for flight among mammals.

Amino Acid Sequence

Flexible algorithm for direct multiple alignment of protein structures and sequences.

The recently described equivalence between the alignment of two proteins and a conformation of a lattice chain on a two-dimensional square lattice is extended to multiple alignments. The search for the optimal multiple alignment between several proteins, which is equivalent to finding the energy minimum in the conformational space of a multi-dimensional lattice chain, is studied by the Monte Carlo approach. This method, while not deterministic, and for two-dimensional problems slower than dynamic programming, can accept arbitrary scoring functions, including non-local ones, and its speed decreases slowly with increasing number of dimensions. For the local scoring functions, the MC algorithm can also reproduce known exact solutions for the direct multiple alignments. As illustrated by examples, both for structure- and sequence-based alignments, direct multi-dimensional alignments are able to capture weak similarities between divergent families much better than ones built from pairwise alignments by a hierarchical approach.

Algorithms

Calculating percent identity between protein or DNA sequences with a word processor.

Two macros, to calculate percentage identity between protein or DNA sequences using the Microsoft Word word processor, are described. The user prepares an alignment file of multiple sequences which is used by the macros to calculate number of matches, number of mismatches, total number of compared positions, and the percent identity. The macros are especially useful when alignment of multiple sequences is possible only by eye.

Algorithms

Evolutionary relationship between the TonB-dependent outer membrane transport proteins: nucleotide and amino acid sequences of the Escherichia coli colicin I receptor gene.

The nucleotide sequence of the Escherichia coli colicin I receptor gene (cir) has been determined. The predicted mature protein consists of 599 amino acids and has a molecular weight of 67,169. Several previously noted characteristics of other E. coli outer membrane protein sequences were also identified in the sequence of Cir. These include an overall acidic nature, the absence of long hydrophobic stretches of amino acids, and a lack of predicted alpha-helical secondary structure. Because two classes of outer membrane proteins (the TonB-dependent transport proteins and the porins) share some structural features, protein sequences from both of these groups were aligned pairwise and scored for sequence similarity. Statistical evidence suggested that the porins were not related to the proteins in the TonB-dependent group; however, there was a significant relationship between the proteins in the TonB-dependent group. On the basis of the multiple progressive sequence alignment and the similarity scores derived from it, a tree representing evolutionary distance between five TonB-dependent outer membrane transport proteins was generated.

Amino Acid Sequence