Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

New method for accurate prediction of solvent accessibility from protein sequence.

A novel method was developed for predicting the solvent accessibility. Based on single sequence data, this method achieved 71.5% accuracy with a correlation coefficient of 0.42 in a database of 704 proteins with threshold of 20% for a two-state-defining solvent accessibility. Prediction in a data subset of 341 monomeric proteins achieved 72.7% accuracy with a correlation coefficient of 0. 43. On the average, prediction over short chains gives better results than that over long chains. With a solvent accessibility threshold of 20%, prediction over 236 monomeric proteins with chain length < 300 amino acid residues achieved 75.3% accuracy with a correlation coefficient of 0.44 by jackknife analysis, which is higher than that obtained by previous methods using multiple sequence alignments.

Algorithms↗

Amino acid recognition by Venus flytrap domains is encoded in an 8-residue motif.

A motif foramino acid recognition by proteins or domains of the periplasmic binding protein-like I superfamily has been identified. An initial pattern of 5 residues was based on a multiple sequence alignment of selected proteins of that fold family and on common structural features observed in the crystal structure of some members of the family [leucine isoleucine valine binding protein (LIVBP), leucine binding protein (LBP), and metabotropic glutamate receptor type 1 (mGlu1R) amino terminal domain)]. This pattern was used against the PIR-NREF sequence database and further refined to retrieve all sequences of proteins that belong to the family and eliminate those that do not belong to it. A motif of 8 residues was finally selected to build up the general signature. A total of 232 sequences were retrieved. They were found to belong to only three families of proteins: bacterial periplasmic binding proteins (PBP, 71 sequences), family 3 (or C) of G-protein coupled receptor (GPCR) (146 sequences), and plant putative ionotropic glutamate receptors (iGluR, 15 sequences). PBPs are known to adopt a bilobate structure also named Venus flytrap domain, or LIVBP domain in the present case. Family 3/C GPCRs are also known to hold such a domain. However, for plant iGluRs, it was previously detected by classical similarity searches but not specifically described. Thus plant iGluRs carry two Venus flytrap domains, one that binds glutamate and an additional one that would be a modulatory LIVBP domain. In some cases, the modulator binding to that domain would be an amino acid.

Amino Acid Motifs↗

Investigation of mechanism of desmopressin binding in vasopressin V2 receptor versus vasopressin V1a and oxytocin receptors: molecular dynamics simulation of the agonist-bound state in the membrane-aqueous system.

The vasopressin V2 receptor (V2R) belongs to the Class A G protein-coupled receptors (GPCRs). V2R is expressed in the renal collecting duct (CD), where it mediates the antidiuretic action of the neurohypophyseal hormone arginine vasopressin (CYFQNCPRG-NH2, AVP). Desmopressin ([1-deamino, 8-D]AVP, dDAVP) is strong selective V2R agonist with negligible pressor and uterotonic activity. In this paper, the interactions responsible for binding of dDAVP to vasopressin V2 receptor versus vasopressin V1a and oxytocin receptors has been examined. Three-dimensional activated models of the receptors were constructed using the multiple sequence alignment and the complex of activated rhodopsin with Gt(alpha) C-terminal peptide of transducin MII-Gt(alpha) (338-350) prototype (Slusarz, R.; Ciarkowski, J. Acta Biochim Pol 2004 51, 129-136) as a template. The 1-ns unconstrained molecular dynamics (MD) of receptor-dDAVP complexes immersed in the fully hydrated 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphatidylcholine (POPC) membrane model was conducted in an Amber 7.0 force field. Highly conserved transmembrane residues have been proposed as being responsible for V2R activation and G protein coupling. Molecular mechanism of the dDAVP binding has been suggested. The internal water molecules involved in an intricate network of the hydrogen bonds inside the receptor cavity have been identified and their role in the stabilization of the agonist-bound state proposed.

Amino Acid Motifs↗

Combined sequence and structure analysis of the fungal laccase family.

Plant and fungal laccases belong to the family of multi-copper oxidases and show much broader substrate specificity than other members of the family. Laccases have consequently been of interest for potential industrial applications. We have analyzed the essential sequence features of fungal laccases based on multiple sequence alignments of more than 100 laccases. This has resulted in identification of a set of four ungapped sequence regions, L1-L4, as the overall signature sequences that can be used to identify the laccases, distinguishing them within the broader class of multi-copper oxidases. The 12 amino acid residues in the enzymes serving as the copper ligands are housed within these four identified conserved regions, of which L2 and L4 conform to the earlier reported copper signature sequences of multi-copper oxidases while L1 and L3 are distinctive to the laccases. The mapping of regions L1-L4 on to the three-dimensional structure of the Coprinus cinerius laccase indicates that many of the non-copper-ligating residues of the conserved regions could be critical in maintaining a specific, more or less C-2 symmetric, protein conformational motif characterizing the active site apparatus of the enzymes. The observed intraprotein homologies between L1 and L3 and between L2 and L4 at both the structure and the sequence levels suggest that the quasi C-2 symmetric active site conformational motif may have arisen from a structural duplication event that neither the sequence homology analysis nor the structure homology analysis alone would have unraveled. Although the sequence and structure homology is not detectable in the rest of the protein, the relative orientation of region L1 with L2 is similar to that of L3 with L4. The structure duplication of first-shell and second-shell residues has become cryptic because the intraprotein sequence homology noticeable for a given laccase becomes significant only after comparing the conservation pattern in several fungal laccases. The identified motifs, L1-L4, can be useful in searching the newly sequenced genomes for putative laccase enzymes.

Amino Acid Sequence↗

Identification and characterization of Uss1p (Sdb23p): a novel U6 snRNA-associated protein with significant similarity to core proteins of small nuclear ribonucleoproteins.

The SDB23 gene of Saccharomyces cerevisiae was isolated in a search for high copy-number suppressors of mutations in a cell cycle gene, DBF2, SDB23 encodes a 21,276 Da protein with significant sequence similarity to characterized mammalian snRNP core proteins. Examination of multiple sequence alignments of snRNP core proteins with Sdb23p indicates that all of these proteins share a number of highly conserved residues, and identifies a novel motif for snRNP core proteins. Sdb23p is essential for cell viability and is required for nuclear pre-mRNA splicing both in vivo and in vitro. Extracts prepared from Sdb23p-depleted cells are unable to support splicing and have vastly reduced levels of U6 snRNA. The stability of U1, U2, U4 and U5 spliceosomal snRNAs is not affected by the loss of Sdb23p. Antibodies raised against Sdb23p strongly coimmunoprecipitate free U6 snRNA and U4/U6 base-paired snRNAs. These results establish that SDB23 encodes a novel U6 snRNA-associated protein that is essential for the stability of U6 snRNA. We therefore propose the more logical name USS1 (U-Six SnRNP) for this gene.

Amino Acid Sequence↗

Molecular analysis of pleckstrin: the major protein kinase C substrate of platelets.

Activation of protein kinase C (PKC) in platelets causes the immediate phosphorylation of pleckstrin, an apparent Mr 40-47,000 protein previously called 40K or P47. Pleckstrin presumably plays an important but as yet unknown role in mediating cellular responses evoked by agonist-induced phosphoinositide turnover. We have cloned the cDNA for pleckstrin from the HL-60 human promyelocytic leukemia cell line by immunological screening of a lambda gt11 expression library (Tyers et al.: Nature 333:470-473, 1988) and now report further analysis of the pleckstrin sequence. Pleckstrin has a deduced Mr of 40,087 and is encoded by a 1,050-bp open reading frame which is preceded by a short open reading frame that terminates before the correct initiator methionine. A single polymorphic site was found in the coding region. An unusual pattern of sequence heterogeneity occurred about a poly(A) tract in the 3' untranslated region. The 3.0-kb pleckstrin mRNA induced upon differentiation of HL-60 cells apparently has heterogeneous 5' ends which undergo differential regulation during HL-60 cell maturation. Analysis by multiple sequence alignment with known PKC substrates identified a strong candidate site for phosphorylation by PKC and a potential Ca2+-binding EF-hand motif. No other similarities to proteins in current databases were found.

Amino Acid Sequence↗

UCSF Chimera--a visualization system for exploratory research and analysis.

The design, implementation, and capabilities of an extensible visualization system, UCSF Chimera, are discussed. Chimera is segmented into a core that provides basic services and visualization, and extensions that provide most higher level functionality. This architecture ensures that the extension mechanism satisfies the demands of outside developers who wish to incorporate new features. Two unusual extensions are presented: Multiscale, which adds the ability to visualize large-scale molecular assemblies such as viral coats, and Collaboratory, which allows researchers to share a Chimera session interactively despite being at separate locales. Other extensions include Multalign Viewer, for showing multiple sequence alignments and associated structures; ViewDock, for screening docked ligand orientations; Movie, for replaying molecular dynamics trajectories; and Volume Viewer, for display and analysis of volumetric data. A discussion of the usage of Chimera in real-world situations is given, along with anticipated future directions. Chimera includes full user documentation, is free to academic and nonprofit users, and is available for Microsoft Windows, Linux, Apple Mac OS X, SGI IRIX, and HP Tru64 Unix from http://www.cgl.ucsf.edu/chimera/.

Amino Acid Sequence↗

Template-based recognition of protein fold within the midnight and twilight zones of protein sequence similarity.

Most homologous pairs of proteins have no significant sequence similarity to each other and are not identified by direct sequence comparison or profile-based strategies. However, multiple sequence alignments of low similarity homologues typically reveal a limited number of positions that are well conserved despite diversity of function. It may be inferred that conservation at most of these positions is the result of the importance of the contribution of these amino acids to the folding and stability of the protein. As such, these amino acids and their relative positions may define a structural signature. We demonstrate that extraction of this fold template provides the basis for the sequence database to be searched for patterns consistent with the fold, enabling identification of homologs that are not recognized by global sequence analysis. The fold template method was developed to address the need for a tool that could comprehensively search the midnight and twilight zones of protein sequence similarity without reliance on global statistical significance. Manual implementations of the fold template method were performed on three folds--immunoglobulin, c-lectin and TIM barrel. Following proof of concept of the template method, an automated version of the approach was developed. This automated fold template method was used to develop fold templates for 10 of the more populated folds in the SCOP database. The fold template method developed three-dimensional structural motifs or signatures that were able to return a diverse collection of proteins, while maintaining a low false positive rate. Although the results of the manual fold template method were more comprehensive than the automated fold template method, the diversity of the results from the automated fold template method surpassed those of current methods that rely on statistical significance to infer evolutionary relationships among divergent proteins.

Amino Acid Sequence↗

Characterization of covalently inhibited extracellular lipase from Streptomyces rimosus by matrix-assisted laser desorption/ionization time-of-flight and matrix-assisted laser desorption/ionization quadrupole ion trap reflectron time-of-flight mass spectrometry: localization of the active site serine.

A chemical modification approach combined with matrix-assisted laser desorption/ionization (MALDI) mass spectrometry was used to identify the active site serine residue of an extracellular lipase from Streptomyces rimosus R6-554W. The lipase, purified from a high-level overexpressing strain, was covalently modified by incubation with 3,4-dichloroisocoumarin, a general mechanism-based serine protease inhibitor. MALDI time-of-flight (TOF) mass spectrometry was used to probe the nature of the intact inhibitor-modified lipase and to clarify the mechanism of lipase inhibition by 3,4-dichloroisocoumarin. The stoichiometry of the inhibition reaction revealed that specifically one molecule of inhibitor was bound to the lipase. The MALDI matrix 2,6-dihydroxyacetophenone facilitated the formation of highly abundant [M + 2H](2+) ions with good resolution compared to other matrices in a linear TOF instrument. This allowed the detection of two different inhibitor-modified lipase species. Exact localization of the modified amino acid residue was accomplished by tryptic digestion followed by low-energy collision-induced dissociation peptide sequencing of the detected 2-(carboxychloromethyl)benzoylated peptide by means of a MALDI quadrupole ion trap reflectron TOF instrument. The high sequence coverage obtained by this approach allowed the confirmation of the site specificity of the inhibition reaction and the unambiguous identification of the serine at position 10 as the nucleophilic amino acid residue in the active site of the enzyme. This result is in agreement with the previously obtained data from multiple sequence alignment of S. rimosus lipase with different esterases, which indicated that this enzyme exhibits a characteristic Gly-Asp-Ser-(Leu) motif located close to the N-terminus and is harboring the catalytically active serine residue. Therefore, this study experimentally proves the classification of the S. rimosus lipase as GDS(L) lipolytic enzyme.

Amino Acid Sequence↗

Modular arrangement of proteins as inferred from analysis of homology.

The structure of many proteins consists of a combination of discrete modules that have been shuffled during evolution. Such modules can frequently be recognized from the analysis of homology. Here we present a systematic analysis of the modular organization of all sequenced proteins. To achieve this we have developed an automatic method to identify protein domains from sequence comparisons. Homologous domains can then be clustered into consistent families. The method was applied to all 21,098 nonfragment protein sequences in SWISS-PROT 21.0, which was automatically reorganized into a comprehensive protein domain database, ProDom. We have constructed multiple sequence alignments for each domain family in ProDom, from which consensus sequences were generated. These nonreduntant domain consensuses are useful for fast homology searches. Domain organization in ProDom is exemplified for proteins of the phosphoenolpyruvate:sugar phosphotransferase system (PEP:PTS) and for bacterial 2-component regulators. We provide 2 examples of previously unrecognized domain arrangements discovered with the help of ProDom.

Amino Acid Sequence↗

MIF proteins are not glutathione transferase homologs.

Although macrophage migration inhibitory factor (MIF) proteins conjugate glutathione, sequence analysis does not support their homology to other glutathione transferases. Glutathione transferases are not detected with MIF proteins in searches of protein sequence databases, and MIF proteins do not share significant sequence similarity with glutathione transferases. Homology cannot be demonstrated by multiple sequence alignment or evolutionary tree construction; such methods assume that the proteins being analyzed are homologous.

Animals↗

Similarity between pyridoxal/pyridoxamine phosphate-dependent enzymes involved in dideoxy and deoxyaminosugar biosynthesis and other pyridoxal phosphate enzymes.

A multiple sequence alignment among aspartate aminotransferase, dialkylglycine decarboxylase, and serine hydroxymethyltransferase (DAS) was used for profile databank search. The DAS profile could detect similarities to other pyridoxal or pyridoxamine phosphate-dependent enzymes, like several gene products involved in dideoxysugar and deoxyaminosugar synthesis. The alignment among DAS and such gene products shows the conservation of aspartate 222 and lysine 258, which, in aspartate aminotransferase, interacts with the N1 of the coenzyme pyridine ring and forms the internal Schiff base, respectively. The lysine is replaced by histidine in the pyridoxamine phosphate-dependent gene products. The alignment indicates also that the region encompassing the coenzyme binding site is the most conserved.

Amino Acid Sequence↗

The sequence of a subtilisin-type protease (aerolysin) from the hyperthermophilic archaeum Pyrobaculum aerophilum reveals sites important to thermostability.

The hyperthermophilic archaeum Pyrobaculum aerophilum grows optimally at 100 degrees C and pH 7.0. Cell homogenates exhibit strong proteolytic activity within a temperature range of 80-130 degrees C. During an analysis of cDNA and genomic sequence tags, a genomic clone was recovered showing strong sequence homology to alkaline subtilisins of Bacillus sp. The total DNA sequence of the gene encoding the protease (named "aerolysin") was determined. Multiple sequence alignment with 15 different serine-type proteases showed greatest homology with subtilisins from gram-positive bacteria rather than archaeal or eukaryal serine proteases. Models of secondary and tertiary structure based on sequence alignments and the tertiary structures of subtilisin Carlsberg, BPN', thermitase, and protease K were generated for P. aerophilum subtilisin. This allowed identification of sites potentially contributing to the thermostability of the protein. One common transition put alanines at the beginning and end of surface alpha-helices. Aspartic acids were found at the N-terminus of several surface helices, possibly increasing stability by interacting with the helix dipole. Several of the substitutions in regions expected to form surface loops were adjacent to each other in the tertiary structure model.

Amino Acid Sequence↗

Eukaryotic translation elongation factor 1 gamma contains a glutathione transferase domain--study of a diverse, ancient protein superfamily using motif search and structural modeling.

Using computer methods for multiple alignment, sequence motif search, and tertiary structure modeling, we show that eukaryotic translation elongation factor 1 gamma (EF1 gamma) contains an N-terminal domain related to class theta glutathione S-transferases (GST). GST-like proteins related to class theta comprise a large group including, in addition to typical GSTs and EF1 gamma, stress-induced proteins from bacteria and plants, bacterial reductive dehalogenases and beta-etherases, and several uncharacterized proteins. These proteins share 2 conserved sequence motifs with GSTs of other classes (alpha, mu, and pi). Tertiary structure modeling showed that in spite of the relatively low sequence similarity, the GST-related domain of EF1 gamma is likely to form a fold very similar to that in the known structures of class alpha, mu, and pi GSTs. One of the conserved motifs is implicated in glutathione binding, whereas the other motif probably is involved in maintaining the proper conformation of the GST domain. We predict that the GST-like domain in EF1 gamma is enzymatically active and that to exhibit GST activity, EF1 gamma has to form homodimers. The GST activity may be involved in the regulation of the assembly of multisubunit complexes containing EF1 and aminoacyl-tRNA synthetases by shifting the balance between glutathione, disulfide glutathione, thiol groups of cysteines, and protein disulfide bonds. The GST domain is a widespread, conserved enzymatic module that may be covalently or noncovalently complexed with other proteins. Regulation of protein assembly and folding may be 1 of the functions of GST.

Amino Acid Sequence↗

Comparative modeling of the three-dimensional structure of type II antifreeze protein.

Type II antifreeze proteins (AFP), which inhibit the growth of seed ice crystals in the blood of certain fishes (sea raven, herring, and smelt), are the largest known fish AFPs and the only class for which detailed structural information is not yet available. However, a sequence homology has been recognized between these proteins and the carbohydrate recognition domain of C-type lectins. The structure of this domain from rat mannose-binding protein (MBP-A) has been solved by X-ray crystallography (Weis WI, Drickamer K, Hendrickson WA, 1992, Nature 360:127-134) and provided the coordinates for constructing the three-dimensional model of the 129-amino acid Type II AFP from sea raven, to which it shows 19% sequence identity. Multiple sequence alignments between Type II AFPs, pancreatic stone protein, MBP-A, and as many as 50 carbohydrate-recognition domain sequences from various lectins were performed to determine reliably aligned sequence regions. Successive molecular dynamics and energy minimization calculations were used to relax bond lengths and angles and to identify flexible regions. The derived structure contains two alpha-helices, two beta-sheets, and a high proportion of amino acids in loops and turns. The model is in good agreement with preliminary NMR spectroscopic analyses. It explains the observed differences in calcium binding between sea raven Type II AFP and MBP-A. Furthermore, the model proposes the formation of five disulfide bridges between Cys 7 and Cys 18, Cys 35 and Cys 125, Cys 69 and Cys 100, Cys 89 and Cys 111, and Cys 101 and Cys 117.(ABSTRACT TRUNCATED AT 250 WORDS)

Adaptation, Physiological↗

Predicting the structure of the light-harvesting complex II of Rhodospirillum molischianum.

We attempted to predict through computer modeling the structure of the light-harvesting complex II (LH-II) of Rhodospirillum molischianum, before the impending publication of the structure of a homologous protein solved by means of X-ray diffraction. The protein studied is an integral membrane protein of 16 independent polypeptides, 8 alpha-apoproteins and 8 beta-apoproteins, which aggregate and bind to 24 bacteriochlorophyll-a's and 12 lycopenes. Available diffraction data of a crystal of the protein, which could not be phased due to a lack of heavy metal derivatives, served to test the predicted structure, guiding the search. In order to determine the secondary structure, hydropathy analysis was performed to identify the putative transmembrane segments and multiple sequence alignment propensity analyses were used to pinpoint the exact sites of the 20-residue-long transmembrane segment and the 4-residue-long terminal sequence at both ends, which were independently verified and improved by homology modeling. A consensus assignment for the secondary structure was derived from a combination of all the prediction methods used. Three-dimensional structures for the alpha- and the beta-apoprotein were built by comparative modeling. The resulting tertiary structures are combined, using X-PLOR, into an alpha beta dimer pair with bacteriochlorophyll-a's attached under constraints provided by site-directed mutagenesis and spectral data. The alpha beta dimer pairs were then aggregated into a quaternary structure through further molecular dynamics simulations and energy minimization. The structure of LH-II so determined is an octamer of alpha beta heterodimers forming a ring with a diameter of 70 A.

Amino Acid Sequence↗

Active site model for gamma-aminobutyrate aminotransferase explains substrate specificity and inhibitor reactivities.

A homology model for the pig isozyme of the pyridoxal phosphate-dependent enzyme gamma-aminobutyrate (GABA) aminotransferase has been built based mainly on the structure of dialkylglycine decarboxylase and on a multiple sequence alignment of 28 evolutionarily related enzymes. The proposed active site structure is presented and analyzed. Hypothetical structures for external aldimine intermediates explain several characteristics of the enzyme. In the GABA external aldimine model, the pro-S proton at C4 of GABA, which abstracted in the 1,3-azaallylic rearrangement interconverting the aldimine and ketimine intermediates, is oriented perpendicular to the plane of the pyridoxal phosphate ring. Lys 329 is in close proximity and is probably the general base catalyst for the proton transfer reaction. The carboxylate group of GABA interacts with Arg 192 and Lys 203, which determine the specificity of the enzyme for monocarboxylic omega-amino acids such as GABA. In the proposed structure for the L-glutamate external aldimine, the alpha-carboxylate interacts with Arg 445. Glu 265 is proposed to interact with this same arginine in the GABA external aldimine, enabling the enzyme to act on omega-amino acids in one half-reaction and on alpha-amino acids in the other. The reactivities of inhibitors are well explained by the proposed active site structure. The R and S isomers of beta-substituted phenyl and p-chlorophenyl GABA would bind in very different modes due to differential steric interactions, with the reactive S isomer leaving the orientation of the GABA moiety relatively unperturbed compared to that of the natural substrate. In our model, only the reactive S isomer of the mechanism-based inhibitor vinyl-GABA, an effective anti-epileptic drug known clinically as Vigabatrin, would orient the scissile C4-H bond perpendicular to the coenzyme ring plane and present the proton to Lys 329, the proposed general base catalyst of the reaction. The R isomer would direct the vinyl group toward Lys 329 and the C4-H bond toward Arg 445. The active site model presented provides a basis for site-directed mutagenesis and drug design experiments.

4-Aminobutyrate Transaminase↗

Subtilases: the superfamily of subtilisin-like serine proteases.

Subtilases are members of the clan (or superfamily) of subtilisin-like serine proteases. Over 200 subtilases are presently known, more than 170 of which with their complete amino acid sequence. In this update of our previous overview (Siezen RJ, de Vos WM, Leunissen JAM, Dijkstra BW, 1991, Protein Eng 4:719-731), details of more than 100 new subtilases discovered in the past five years are summarized, and amino acid sequences of their catalytic domains are compared in a multiple sequence alignment. Based on sequence homology, a subdivision into six families is proposed. Highly conserved residues of the catalytic domain are identified, as are large or unusual deletions and insertions. Predictions have been updated for Ca(2+)-binding sites, disulfide bonds, and substrate specificity, based on both sequence alignment and three-dimensional homology modeling.

Amino Acid Sequence↗