Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

GTPase domains of ras p21 oncogene protein and elongation factor Tu: analysis of three-dimensional structures, sequence families, and functional sites.

GTPase domains are functional and structural units employed as molecular switches in a variety of important cellular functions, such as growth control, protein biosynthesis, and membrane traffic. Amino acid sequences of more than 100 members of different subfamilies are known, but crystal structures of only mammalian ras p21 and bacterial elongation factor Tu have been determined. After optimal superposition of these remarkably similar structures, careful multiple sequence alignment, and calculation of residue-residue interactions, we analyzed the two subfamilies in terms of structural conservation, sequence conservation, and residue contact strength. There are three main results. (i) A structure-based alignment of p21 and elongation factor Tu. (ii) The definition of a common conserved structural core that may be useful as the basis of model building by homology of the three-dimensional structure of any GTPase domain. (iii) Identification of sequence regions, other than the effector loop and the nucleotide binding site, that may be involved in the functional cycle: they are loop L4, known to change conformation after GTP hydrolysis; helix alpha 2, especially Arg-73 and Met-67 in ras p21; loops L8 and L10, including ras p21 Arg-123, Lys-147, and Leu-120; and residues located spatially near the N and C termini. These regions are candidate sites for interaction either with the GTP/GDP exchange factor, with a GTPase-affected function, or with a molecule delivered to a destination site with the aid of the GTPase domain.

Amino Acid Sequence

Improved prediction of protein secondary structure by use of sequence profiles and neural networks.

The explosive accumulation of protein sequences in the wake of large-scale sequencing projects is in stark contrast to the much slower experimental determination of protein structures. Improved methods of structure prediction from the gene sequence alone are therefore needed. Here, we report a substantial increase in both the accuracy and quality of secondary-structure predictions, using a neural-network algorithm. The main improvements come from the use of multiple sequence alignments (better overall accuracy), from "balanced training" (better prediction of beta-strands), and from "structure context training" (better prediction of helix and strand lengths). This method, cross-validated on seven different test sets purged of sequence similarity to learning sets, achieves a three-state prediction accuracy of 69.7%, significantly better than previous methods. In addition, the predicted structures have a more realistic distribution of helix and strand segments. The predictions may be suitable for use in practice as a first estimate of the structural type of newly sequenced proteins.

Amino Acid Sequence

Identification of active-site residues of the adenovirus endopeptidase.

Multiple sequence alignment of the 12 adenovirus endopeptidases known to date identified a number of conserved residues which might be important for enzyme activity. Eleven mutants were created in the cloned gene by site-directed mutagenesis to identify the active site of this thiol endopeptidase. Analysis of the proteolytic activity in a crude system using viral precursor proteins, as well as in a purified system with activated proteinases using a new chromophoric octapeptide substrate, yielded results consistent with Cys-104 and His-54 being two members of the active site. This result was confirmed by the carboxymethylation of the reactive Cys-104 and its prevention by the active-thiol-specific agent E64. Although Cys-122 and Cys-126 were also reactive cysteines, mutation of these residues did not affect enzyme activity. Replacement of the active-site Cys-104 by serine converted the enzyme into a serine-like proteinase, sensitive to serine proteinase inhibitors. The absence of homology to other proteinases, particularly at the active-site cysteine, coupled with the requirement for activation by a substrate cleavage fragment, indicates that the adenovirus endoproteinase may represent a new subclass of cysteine proteinases.

Adenoviridae

Determinant for beta-subunit regulation in high-conductance voltage-activated and Ca(2+)-sensitive K+ channels: an additional transmembrane region at the N terminus.

The pore-forming alpha subunit of large conductance voltage- and Ca(2+)-sensitive K (MaxiK) channels is regulated by a beta subunit that has two membrane-spanning regions separated by an extracellular loop. To investigate the structural determinants in the pore-forming alpha subunit necessary for beta-subunit modulation, we made chimeric constructs between a human MaxiK channel and the Drosophila homologue, which we show is insensitive to beta-subunit modulation, and analyzed the topology of the alpha subunit. A comparison of multiple sequence alignments with hydrophobicity plots revealed that MaxiK channel alpha subunits have a unique hydrophobic segment (S0) at the N terminus. This segment is in addition to the six putative transmembrane segments (S1-S6) usually found in voltage-dependent ion channels. The transmembrane nature of this unique S0 region was demonstrated by in vitro translation experiments. Moreover, normal functional expression of signal sequence fusions and in vitro N-linked glycosylation experiments indicate that S0 leads to an exoplasmic N terminus. Therefore, we propose a new model where MaxiK channels have a seventh transmembrane segment at the N terminus (S0). Chimeric exchange of 41 N-terminal amino acids, including S0, from the human MaxiK channel to the Drosophila homologue transfers beta-subunit regulation to the otherwise unresponsive Drosophila channel. Both the unique S0 region and the exoplasmic N terminus are necessary for this gain of function.

Amino Acid Sequence

Molecular cloning of the isoquinoline 1-oxidoreductase genes from Pseudomonas diminuta 7, structural analysis of iorA and iorB, and sequence comparisons with other molybdenum-containing hydroxylases.

The iorA and iorB genes from the isoquinoline-degrading bacterium Pseudomonas diminuta 7, encoding the heterodimeric molybdo-iron-sulfur-protein isoquinoline 1-oxidoreductase, were cloned and sequenced. The deduced amino acid sequences IorA and IorB showed homologies (i) to the small (gamma) and large (alpha) subunits of complex molybdenum-containing hydroxylases (alpha beta gamma/alpha 2 beta 2 gamma 2) possessing a pterin molybdenum cofactor with a monooxo-monosulfido-type molybdenum center, (ii) to the N- and C-terminal regions of aldehyde oxidoreductase from Desulfovibrio gigas, and (iii) to the N- and C-terminal domains of eucaryotic xanthine dehydrogenases, respectively. The closest similarity to IorB was shown by aldehyde dehydrogenase (Adh) from the acetic acid bacterium Acetobacter polyoxogenes. Five conserved domains of IorB were identified by multiple sequence alignments. Whereas IorB and Adh showed an identical sequential arrangement of these conserved domains, in all other molybdenum-containing hydroxylases the relative position of "domain A" differed. IorA contained eight conserved cysteine residues. The amino acid pattern harboring the four cysteine residues proposed to ligate the Fe/S I cluster was homologous to the consensus binding site of bacterial and chloroplast-type [2Fe-2S] ferredoxins, whereas the pattern including the four cysteines assumed to ligate the Fe/S II center showed no similarities to any described [2Fe-2S] binding motif. The N-terminal region of IorB comprised a putative signal peptide similar to typical leader peptides, indicating that isoquinoline 1-oxidoreductase is associated with the cell membrane.

Amino Acid Sequence

Cloning of a novel family of mammalian GTP-binding proteins (RagA, RagBs, RagB1) with remote similarity to the Ras-related GTPases.

cDNA clones of two novel Ras-related GTP-binding proteins (RagA and RagB) were isolated from rat and human cDNA libraries. Their deduced amino acid sequences comprise four of the six known conserved GTP-binding motifs (PM1, -2, -3, G1), the remaining two (G2, G3) being strikingly different from those of the Ras family, and an unusually large C-terminal domain (100 amino acids) presumably unrelated to GTP binding. RagA and RagB differ by seven conservative amino acid substitutions (98% identity), and by 33 additional residues at the N terminus of RagB. In addition, two isoforms of RagB (RagBs and RagB1) were found that differed only by an insertion of 28 codons between the GTP-binding motifs PM2 and PM3, apparently generated by alternative mRNA splicing. Polymerase chain reaction amplification with specific primers indicated that both long and short form of RagB transcripts were present in adrenal gland, thymus, spleen, and kidney, whereas in brain, only the long form RagB1 was detected. A long splicing variant of RagA was not detected. Recombinant glutathione S-transferase (GST) fusion proteins of RagA and RagBs bound large amounts of radiolabeled GTP gamma S in a specific and saturable manner. In contrast, GTP gamma S binding of GST-RagB1 hardly exceeded that of recombinant GST. GTP gamma S bound to recombinant RagA, and RagBs was rapidly exchangeable for GTP, whereas no intrinsic GTPase activity was detected. A multiple sequence alignment indicated that RagA and RagB cannot be assigned to any of the known subfamilies of Ras-related GTPases but exhibit a 52% identity with a yeast protein (Gtr1) presumably involved in phosphate transport and/or cell growth. It is suggested that RagA and RagB are the mammalian homologues of Gtr1 and that they represent a novel subfamily of Ras-homologous GTP binding proteins.

Amino Acid Sequence

cDNA cloning, tissue distribution, and identification of the catalytic triad of monoglyceride lipase. Evolutionary relationship to esterases, lysophospholipases, and haloperoxidases.

Monoglyceride lipase catalyzes the last step in the hydrolysis of stored triglycerides in the adipocyte and presumably also complements the action of lipoprotein lipase in degrading triglycerides from chylomicrons and very low density lipoproteins. Monoglyceride lipase was cloned from a mouse adipocyte cDNA library. The predicted amino acid sequence consisted of 302 amino acids, corresponding to a molecular weight of 33,218. The sequence showed no extensive homology to other known mammalian proteins, but a number of microbial proteins, including two bacterial lysophospholipases and a family of haloperoxidases, were found to be distantly related to this enzyme. By means of multiple sequence alignment and secondary structure prediction, the structural elements in monoglyceride lipase, as well as the putative catalytic triad, were identified. The residues of the proposed triad, Ser-122, in a GXSXG motif, Asp-239, and His-269, were confirmed by site-directed mutagenesis experiments. Northern blot analysis revealed that monoglyceride lipase is ubiquitously expressed among tissues, with a transcript size of about 4 kilobases.

Adipocytes

Molecular modelling of mammalian CYP2B isoforms and their interaction with substrates, inhibitors and redox partners.

1. The construction of three-dimensional models of CYP2B isozymes from rat (CYP2B1), rabbit (CYP2B4) and man (CYP2B6), based on a multiple sequence alignment with CYP102, a unique eukaryotic-like bacterial P450 (in terms of possessing an NADPH-dependent FAD- and FMN-containing oxidoreductase redox partner) of known crystal structure, is reported. 2. The enzyme models described are shown to be consistent with experimental evidence from site-directed mutagenesis studies, antibody recognition sites and amino acid residues identified as being associated with redox partner interactions, together with the location of a key serine residue (Ser-128) likely to be involved in protein kinaseA-mediated phosphorylation. 3. A substantial number of known substrates and inhibitors of CYP2B isozymes are shown to fit the putative active sites of the enzyme models in agreement with their reported position of metabolism or mode of inhibition respectively. In particular, there is complementarity between the characteristic non-planar geometries of CYP2B substrates and key groups in the enzymes' active sites. 4. Molecular modelling of CYP2B isozymes appears to rationalize a number of the reported findings from quantitative structure-activity relationship investigations on series of CYP2B substrates and inhibitors.

Amino Acid Sequence

A new family of plasma membrane polypeptides differentially regulated during plant development.

Two cDNAs encoding polypeptides identified in a tobacco leaf plasma membrane fraction prepared by phase partitioning were cloned. The deduced polypeptides, P16 and P17, exhibit a striking primary structure, similar to that of P19, a previously cloned plasma membrane polypeptide. Antibodies raised to the recombinant proteins were used to probe the cellular location of P16, P17 and P19 by means of western blotting of sucrose density-gradient fractions; all three polypeptides were found to be located solely at the plasma membrane. Furthermore, P19 antigen accumulated transiently at the time of floral induction while P16 and P17 antigens accumulated towards the end of the life-cycle. These results together with sequence database searches and multiple-sequence alignments suggest that we have identified a new family of plasma membrane polypeptides that (i) are putatively plant specific and (ii) are differentially regulated during plant development. These polypeptides are termed DREPPs for developmentally regulated plasma membrane polypeptides.

Amino Acid Sequence

Statistical modeling, phylogenetic analysis and structure prediction of a protein splicing domain common to inteins and hedgehog proteins.

Inteins, introns spliced at the protein level, and the hedgehog family of proteins involved in eucaryotic development both undergo autocatalytic proteolysis. Here, a specific and sensitive hidden Markov model (HMM) of protein splicing domain shared by inteins and the hedgehog proteins has been trained and employed for further analysis. The HMM characterizes the common features of this domain including the position where a site-specific DNA endonuclease domain is inserted in the majority of the inteins. The HMM was used to identify several new putative inteins, such as that in the Methanococcus jannaschii klbA protein, and to generate a multiple sequence alignment of sequences possessing this domain. Phylogenetic analysis suggests that hedgehog proteins evolved from inteins. Secondary and tertiary structure predictions suggest that the domain has a structure similar to a beta-sandwich. Similarities between the serine protease cleavage mechanism and the protein splicing reaction mechanism are discussed. Examination of the locations of inteins indicates that they are not inserted randomly in an extein, but are often inserted at functionally important positions in the host proteins. A specific and sensitive HMM for a domain present in klbA proteins identified several additional bacterial and archaeal family members, and analysis of the site of insertion of the intein suggests residues that may be functionally important. This domain may play a role in formation of surface-associated protein complexes.

Algorithms

PHD--an automatic mail server for protein secondary structure prediction.

By the middle of 1993, > 30,000 protein sequences has been listed. For 1000 of these, the three-dimensional (tertiary) structure has been experimentally solved. Another 7000 can be modelled by homology. For the remaining 21,000 sequences, secondary structure prediction provides a rough estimate of structural features. Predictions in three states range between 35% (random) and 88% (homology modelling) overall accuracy. Using information about evolutionary conservation as contained in multiple sequence alignments, the secondary structure of 4700 protein sequences was predicted by the automatic e-mail server PHD. For proteins with at least one known homologue, the method has an expected overall three-state accuracy of 71.4% for proteins with at least one known homologue (evaluated on 126 unique protein chains).

Algorithms

SEQSEE: a comprehensive program suite for protein sequence analysis.

SEQSEE (SEQuence SEEker) is a multi-purpose, menu-driven suite of programs designed to provide a fully integrated, state-of-the-art package for the analysis and display of protein sequences and protein databases. It is currently configured to run on most UNIX-based machines including Sun, SGI and NeXT workstations with conversion to other architectures (e.g. Vax or Cray) being a relatively simple task. SEQSEE is capable of performing nearly all of the analytical and comparative tasks found in most comprehensive commercially available software packages. These include sequence/database searching, sequence retrieval, sequence entry and editing, statistical sequence analysis, multiple sequence alignment, flexible pattern matching, and secondary structure prediction. SEQSEE also integrates a number of unique databases which allow it to perform many additional functions such as structure-based sequence alignments and homology-based secondary structure prediction. Additional enhancements to many previously published algorithms have substantially improved the performance of SEQSEE over that found for most other commercial products. The source code, the documentation and all of the required databases for SEQSEE are freely available and may be obtained by anonymous ftp.

Algorithms

Comparison of side chain interactions performed by structurally equivalent residues in homologous protein structures.

The present work describes the computer program Hom-Bond, which allows to identify and compare intra-molecular interactions performed by side chain polar atoms as observed in a family of homologous protein structures with known and conserved 3-D conformation. For this purpose, the side chain to side chain and the side chain to main chain hydrogen bonds, the disulfide and the salt bridges are identified in each considered protein structure. Subsequently, the side chain interactions are displayed according to the multiple sequence alignment. The presented approach allows to easily identify bonds which are conserved in homologous proteins and to analyse rearrangements of the network of side chain interactions that characterize each protein structure.

Amino Acid Sequence

Sisyphus and prediction of protein structure.

The problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and 1D prediction have reopened the field. Mean-force-potentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue contacts (2D) can be detected by analysis of correlated mutations, albeit with low accuracy. Secondary structure, solvent accessibility and transmembrane helices (1D) can be predicted with significantly improved accuracy using multiple sequence alignments. Some of these new prediction methods have proven accurate and reliable enough to be useful in genome analysis, and in experimental structure determination. Moreover, the new generation of theoretical methods is increasingly influencing experiments in molecular biology.

Computers

Efficient discovery of conserved patterns using a pattern graph.

MOTIVATION: We have previously reported an algorithm for discovering patterns conserved in sets of related unaligned protein sequences. The algorithm was implemented in a program called Pratt. Pratt allows the user to define a class of patterns (e.g. the degree of ambiguity allowed and the length and number of gaps), and is then guaranteed to find the conserved patterns in this class scoring highest according to a defined fitness measure. In many cases, this version of Pratt was very efficient, but in other cases it was too time consuming to be applied. Hence, a more efficient algorithm was needed. RESULTS: In this paper, we describe a new and improved searching strategy that has two main advantages over the old strategy. First, it allows for easier integration with programs for multiple sequence alignment and data base search. Secondly, it makes it possible to use branch-and-bound search, and heuristics, to speed up the search. The new search strategy has been implemented in a new version of the Pratt program.

Algorithms

VHMPT: a graphical viewer and editor for helical membrane protein topologies.

MOTIVATION: Lacking structures resolved at atomic resolution, the great majority of membrane proteins have typically been depicted in a schematic two-dimensional (2D) topology consisting of putative transmembrane domains predicted from hydropathy plots. As more and more sequences of membrane proteins become available from genome projects, there is a need to automate the process of generating the schematic topology while allowing important information, such as the individual amino acid and the extent to which it is conserved in evolution, to be conveniently inspected. We addressed this need by developing a program called VHMPT. RESULTS: VHMPT (a graphical V iewer and editor for H elical line M embrane P rotein T opologies) can automatically generate a schematic 2D topology for a protein with transmembrane helices. Through an interactive graphical interface, VHMPT allows users to modify the layout of the generated topology, label specific amino acid or amino acid groups, and annotate with arrows and texts. Given a multiple sequence alignment file, VHMPT can also color code a normalized conservation score for each amino acid on the generated topology, allowing ready visual recognition of highly conserved (or variable) topological regions. VHMPT is written in Tcl/Tk and can run on platforms that have installed the Tcl/Tk interpreter. AVAILABILITY: The source code and a user manual for VHMPT are available for download at http://www. ibms.sinica.edu.tw/mjhwang/vhmpt. CONTACT: mjhwang@mail.ibms.sinica.edu.tw

Computational Biology

TOPAL: recombination detection in DNA and protein sequences.

UNLABELLED: TOPAL scans a multiple sequence alignment for evidence of recombinant sequences, prior to phylogenetic analysis. AVAILABILITY: The TOPAL package may be accessed at http://www.bioss.sari.ac.uk/grainne, and by anonymous ftp at ftp.bioss. sari.ac.uk in the directory pub/phylogeny/topal. CONTACT: grainne@bioss.sari.ac.uk

Computational Biology

Profile hidden Markov models.

The recent literature on profile hidden Markov model (profile HMM) methods and software is reviewed. Profile HMMs turn a multiple sequence alignment into a position-specific scoring system suitable for searching databases for remotely homologous sequences. Profile HMM analyses complement standard pairwise comparison methods for large-scale sequence analysis. Several software implementations and two large libraries of profile HMMs of common protein domains are available. HMM methods performed comparably to threading methods in the CASP2 structure prediction exercise.

Humans