Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Predicting protein secondary structure with probabilistic schemata of evolutionarily derived information.

We demonstrate the applicability of our previously developed Bayesian probabilistic approach for predicting residue solvent accessibility to the problem of predicting secondary structure. Using only single-sequence data, this method achieves a three-state accuracy of 67% over a database of 473 non-homologous proteins. This approach is more amenable to inspection and less likely to overlearn specifics of a dataset than "black box" methods such as neural networks. It is also conceptually simpler and less computationally costly. We also introduce a novel method for representing and incorporating multiple-sequence alignment information within the prediction algorithm, achieving 72% accuracy over a dataset of 304 non-homologous proteins. This is accomplished by creating a statistical model of the evolutionarily derived correlations between patterns of amino acid substitution and local protein structure. This model consists of parameter vectors, termed "substitution schemata," which probabilistically encode the structure-based heterogeneity in the distributions of amino acid substitutions found in alignments of homologous proteins. The model is optimized for structure prediction by maximizing the mutual information between the set of schemata and the database of secondary structures. Unlike "expert heuristic" methods, this approach has been demonstrated to work well over large datasets. Unlike the opaque neural network algorithms, this approach is physicochemically intelligible. Moreover, the model optimization procedure, the formalism for predicting one-dimensional structural features and our previously developed method for tertiary structure recognition all share a common Bayesian probabilistic basis. This consistency starkly contrasts with the hybrid and ad hoc nature of methods that have dominated this field in recent years.

Algorithms↗

Sequence analysis of the AAA protein family.

The AAA protein family, a recently recognized group of Walker-type ATPases, has been subjected to an extensive sequence analysis. Multiple sequence alignments revealed the existence of a region of sequence similarity, the so-called AAA cassette. The borders of this cassette were localized and within it, three boxes of a high degree of conservation were identified. Two of these boxes could be assigned to substantial parts of the ATP binding site (namely, to Walker motifs A and B); the third may be a portion of the catalytic center. Phylogenetic trees were calculated to obtain insights into the evolutionary history of the family. Subfamilies with varying degrees of intra-relatedness could be discriminated; these relationships are also supported by analysis of sequences outside the canonical AAA boxes: within the cassette are regions that are strongly conserved within each subfamily, whereas little or even no similarity between different subfamilies can be observed. These regions are well suited to define fingerprints for subfamilies. A secondary structure prediction utilizing all available sequence information was performed and the result was fitted to the general 3D structure of a Walker A/GTPase. The agreement was unexpectedly high and strongly supports the conclusion that the AAA family belongs to the Walker superfamily of A/GTPases.

Adenosine Triphosphatases↗

Motifs and structural fold of the cofactor binding site of human glutamate decarboxylase.

The pyridoxal-P binding sites of the two isoforms of human glutamate decarboxylase (GAD65 and GAD67) were modeled by using PROBE (a recently developed algorithm for multiple sequence alignment and database searching) to align the primary sequence of GAD with pyridoxal-P binding proteins of known structure. GAD's cofactor binding site is particularly interesting because GAD activity in the brain is controlled in part by a regulated interconversion of the apo- and holoenzymes. PROBE identified six motifs shared by the two GADs and four proteins of known structure: bacterial ornithine decarboxylase, dialkylglycine decarboxylase, aspartate aminotransferase, and tyrosine phenol-lyase. Five of the motifs corresponded to the alpha/beta elements and loops that form most of the conserved fold of the pyridoxal-P binding cleft of the four enzymes of known structure; the sixth motif corresponded to a helical element of the small domain that closes when the substrate binds. Eight residues that interact with pyridoxal-P and a ninth residue that lies at the interface of the large and small domains were also identified. Eleven additional conserved residues were identified and their functions were evaluated by examining the proteins of known structure. The key residues that interact directly with pyridoxal-P were identical in ornithine decarboxylase and the two GADs, thus allowing us to make a specific structural prediction of the cofactor binding site of GAD. The strong conservation of the cofactor binding site in GAD indicates that the highly regulated transition between apo- and holoGAD is accomplished by modifications in this basic fold rather than through a novel folding pattern.

Amino Acid Sequence↗

Evidence that peroxiredoxins are novel members of the thioredoxin fold superfamily.

Peroxiredoxins catalyze reduction of hydrogen peroxide or alkyl peroxide, to water or the corresponding alcohol. Detailed analysis of their sequences indicates that these enzymes possess a thioredoxin (Trx)-like fold and consequently are homologues of both thioredoxin and glutathione peroxidase (GPx). Sequence- and structure-based multiple sequence alignments indicate that the peroxiredoxin active site cysteine and GPx active site selenocysteine are structurally equivalent. Homologous peroxiredoxin and GPx enzymes are predicted to catalyze equivalent reactions via similar reaction intermediates.

Amino Acid Sequence↗

Probing the ligand binding domain of the GluR2 receptor by proteolysis and deletion mutagenesis defines domain boundaries and yields a crystallizable construct.

Ionotropic glutamate receptors constitute an important family of ligand-gated ion channels for which there is little biochemical or structural data. Here we probe the domain structure and boundaries of the ligand binding domain of the AMPA-sensitive GluR2 receptor by limited proteolysis and deletion mutagenesis. To identify the proteolytic fragments, Maldi mass spectrometry and N-terminal amino acid sequencing were employed. Trypsin digestion of HS1S2 (Chen GQ, Gouaux E. 1997. Proc Natl Acad Sci USA 94:13431-13436) in the presence and absence of glutamate showed that the ligand stabilized the S1 and S2 fragments against complete digestion. Using limited proteolysis and multiple sequence alignments of glutamate receptors as guides, nine constructs were made, folded, and screened for ligand binding activity. From this screen, the S1S21 construct proved to be trypsin- and chymotrypsin-resistant, stable to storage at 4 degrees C, and amenable to three-dimensional crystal formation. The HS1S21 variant was readily prepared on a large scale, the His tag was easily removed by trypsin, and crystals were produced that diffracted to beyond 1.5 A resolution. These experiments, for the first time, pave the way to economical overproduction of the ligand binding domains of glutamate receptors and more accurately map the boundaries of the ligand binding domain.

Amino Acid Sequence↗

In silico two-hybrid system for the selection of physically interacting protein pairs.

Deciphering the interaction links between proteins has become one of the main tasks of experimental and bioinformatic methodologies. Reconstruction of complex networks of interactions in simple cellular systems by integrating predicted interaction networks with available experimental data is becoming one of the most demanding needs in the postgenomic era. On the basis of the study of correlated mutations in multiple sequence alignments, we propose a new method (in silico two-hybrid, i2h) that directly addresses the detection of physically interacting protein pairs and identifies the most likely sequence regions involved in the interactions. We have applied the system to several test sets, showing that it can discriminate between true and false interactions in a significant number of cases. We have also analyzed a large collection of E. coli protein pairs as a first step toward the virtual reconstruction of its complete interaction network.

Animals↗

Scoring residue conservation.

The importance of a residue for maintaining the structure and function of a protein can usually be inferred from how conserved it appears in a multiple sequence alignment of that protein and its homologues. A reliable metric for quantifying residue conservation is desirable. Over the last two decades many such scores have been proposed, but none has emerged as a generally accepted standard. This work surveys the range of scores that biologists, biochemists, and, more recently, bioinformatics workers have developed, and reviews the intrinsic problems associated with developing and evaluating such a score. A general formula is proposed that may be used to compare the properties of different particular conservation scores or as a measure of conservation in its own right.

Amino Acids↗

Dynamic fluorescence studies of beta-glycosidase mutants from Sulfolobus solfataricus: effects of single mutations on protein thermostability.

Multiple sequence alignment on 73 proteins belonging to glycosyl hydrolase family 1 reveals the occurrence of a segment (83-124) in the enzyme sequences from hyperthermophilic archaea bacteria, which is absent in all the mesophilic members of the family. The alignment of the known three-dimensional structures of hyperthermophilic glycosidases with the known ones from mesophilic organisms shows a similar spatial organizations of beta-glycosidases except for this sequence segment whose structure is located on the external surface of each of four identical subunits, where it overlaps two alpha-helices. Site-directed mutagenesis substituting N97 or S101 with a cysteine residue in the sequence of beta-glycosidase from hyperthermophilic archaeon Sulfolobus solfataricus caused some changes in the structural and dynamic properties as observed by circular dichroism in far- and near-UV light, as well as by frequency domain fluorometry, with a simultaneous loss of thermostability. The results led us to hypothesize an important role of the sequence segment present only in hyperthermophilic beta-glycosidases, in the thermal adaptation of archaea beta-glycosidases. The thermostabilization mechanism could occur as a consequence of numerous favorable ionic interactions of the 83-124 sequence with the other part of protein matrix that becomes more rigid and less accessible to the insult of thermal-activated solvent molecules.

Adaptation, Physiological↗

Multiple linear regression for protein secondary structure prediction.

In the present work, a novel method was proposed for prediction of secondary structure. Over a database of 396 proteins (CB396) with a three-state-defining secondary structure, this method with jackknife procedure achieved an accuracy of 68.8% and SOV score of 71.4% using single sequence and an accuracy of 73.7% and SOV score of 77.3% using multiple sequence alignments. Combination of this method with DSC, PHD, PREDATOR, and NNSSP gives Q3 = 76.2% and SOV = 79.8%.

Databases, Factual↗

Conserved residue clustering and protein structure prediction.

Protein residues that are critical for structure and function are expected to be conserved throughout evolution. Here, we investigate the extent to which these conserved residues are clustered in three-dimensional protein structures. In 92% of the proteins in a data set of 79 proteins, the most conserved positions in multiple sequence alignments are significantly more clustered than randomly selected sets of positions. The comparison to random subsets is not necessarily appropriate, however, because the signal could be the result of differences in the amino acid composition of sets of conserved residues compared to random subsets (hydrophobic residues tend to be close together in the protein core), or differences in sequence separation of the residues in the different sets. In order to overcome these limits, we compare the degree of clustering of the conserved positions on the native structure and on alternative conformations generated by the de novo structure prediction method Rosetta. For 65% of the 79 proteins, the conserved residues are significantly more clustered in the native structure than in the alternative conformations, indicating that the clustering of conserved residues in protein structures goes beyond that expected purely from sequence locality and composition effects. The differences in the spatial distribution of conserved residues can be utilized in de novo protein structure prediction: We find that for 79% of the proteins, selection of the Rosetta generated conformations with the greatest clustering of the conserved residues significantly enriches the fraction of close-to-native structures.

Acyl Coenzyme A↗

Weighted geometric docking: incorporating external information in the rotation-translation scan.

Weighted geometric docking is a prediction algorithm that matches weighted molecular surfaces. Each molecule is represented by a grid of complex numbers, storing information about the shape of the molecule in the real part and weight information in the imaginary part. The weights are based on experimental biochemical and biophysical data or on theoretical analyses of amino acid conservation or correlation patterns in multiple-sequence alignments of homologous proteins. Only a few surface residues on either one or both molecules are weighted. In contrast to methods that use postscan filtering based on biochemical information, our method incorporates the external data in the rotation-translation search, producing a different set of docking solutions biased toward solutions in which the up-weighted residues are at the interface. Similarly, interactions involving specified residues can be impeded. The weighted geometric algorithm was applied to five systems for which regular geometric docking of the unbound molecules gave poor results. We obtained much better ranking of the nearly correct prediction and higher statistical significance when weighted geometric docking was used. The method was successful even when the weighted portion of the surface corresponded only partially and approximately to the binding site.

Algorithms↗

Deciphering a novel thioredoxin-like fold family.

Sequence--and structure-based searching strategies have proven useful in the identification of remote homologs and have facilitated both structural and functional predictions of many uncharacterized protein families. We implement these strategies to predict the structure of and to classify a previously uncharacterized cluster of orthologs (COG3019) in the thioredoxin-like fold superfamily. The results of each searching method indicate that thioltransferases are the closest structural family to COG3019. We substantiate this conclusion using the ab initio structure prediction method rosetta, which generates a thioredoxin-like fold similar to that of the glutaredoxin-like thioltransferase (NrdH) for a COG3019 target sequence. This structural model contains the thiol-redox functional motif CYS-X-X-CYS in close proximity to other absolutely conserved COG3019 residues, defining a novel thioredoxin-like active site that potentially binds metal ions. Finally, the rosetta-derived model structure assists us in assembling a global multiple-sequence alignment of COG3019 with two other thioredoxin-like fold families, the thioltransferases and the bacterial arsenate reductases (ArsC).

Amino Acid Sequence↗

Evolutionary analysis reveals collective properties and specificity in the C-type lectin and lectin-like domain superfamily.

Members of the C-type lectin/C-type lectin-like domain (CTL/CTLD) superfamily share a common fold and are involved in a variety of functions, such as generalized defense mechanisms against foreign agents, discrimination between healthy and pathogen-infected cells, and endocytosis and blood coagulation. In this work we used ConSurf, a computer program recently developed in our lab, to perform an evolutionary analysis of this superfamily in order to further identify characteristics of all or part of its members. Given a set of homologous proteins in the form of multiple sequence alignment (MSA) and an inferred phylogenetic tree, ConSurf calculates the conservation score in every alignment position, taking into account the relationships between the sequences and the physicochemical similarity between the amino acids. The scores are then color-coded onto the three-dimensional structure of one of the homologous proteins. We provide here and at http://ashtoret.tau.ac.il/ approximately sharon a detailed analysis of the conservation pattern obtained for the entire superfamily and for two subgroups of proteins: (a) 21 CTLs and (b) 11 heterodimeric CTLD toxins. We show that, in general, proteins of the superfamily have one face that is constructed mostly of conserved residues and another that is not, and we suggest that the former face is involved in binding to other proteins or domains. In the CTLs examined we detected a region of highly conserved residues, corresponding to the known calcium- and carbohydrate-binding site of the family, which is not conserved throughout the entire superfamily, and in the CTLD toxins we found a patch of highly conserved residues, corresponding to the known dimerization region of these proteins. Our analysis also detected patches of conserved residues with yet unknown function(s).

Amino Acid Sequence↗

Conserved structural elements in glutathione transferase homologues encoded in the genome of Escherichia coli.

Multiple sequence alignments of the eight glutathione (GSH) transferase homologues encoded in the genome of Escherichia coli were used to define a consensus sequence for the proteins. The consensus sequence was analyzed in the context of the three-dimensional structure of the gst gene product (EGST) obtained from two different crystal forms of the enzyme. The enzyme consists of two domains. The N-terminal region (domain I) has a thioredoxin-like alpha/beta-fold, while the C-terminal domain (domain II) is all alpha-helical. The majority of the consensus residues (12/17) reside in the N-terminal domain. Fifteen of the 17 residues are involved in hydrophobic core interactions, turns, or electrostatic interactions between the two domains. The results suggest that all of the homologues retain a well-defined group of structural elements both in and between the N-terminal alpha/beta domain and the C-terminal domain. The conservation of two key residues for the recognition motif for the gamma-glutamyl-portion of GSH indicates that the homologues may interact with GSH or GSH analogues such as glutathionylspermidine or alpha-amino acids. The genome context of two of the homologues forms the basis for a hypothesis that the b2989 and yibF gene products are involved in glutathionylspermidine and selenium biochemistry, respectively.

Amino Acid Sequence↗

3D structural model of the G-protein-coupled cannabinoid CB2 receptor.

The potential for therapeutic specificity in regulating diseases and for reduced side effects has made cannabinoid (CB) receptors one of the most important G-protein-coupled receptor (GPCR) targets for drug discovery. The cannabinoid (CB) receptor subtype CB2 is of particular interest due to its involvement in signal transduction in the immune system and its increased characterization by mutational and other studies. However, our understanding of their mode of action has been limited by the absence of an experimental receptor structure. In this study, we have developed a 3D model of the CB2 receptor based on the recent crystal structure of a related GPCR, bovine rhodopsin. The model was developed using multiple sequence alignment of homologous receptor sub-types in humans and mammals, and compared with other GPCRs. Alignments were analyzed with mutation scores, pairwise hydrophobicity profiles and Kyte-Doolittle plots. The 3D model of the transmembrane segment was generated by mapping the CB2 sequence onto the homologous residues of the rhodopsin structure. The extra- and intracellular loop regions of the CB2 were generated by searching for homologous C(alpha) backbone sequences in published structures in the Brookhaven Protein Databank (PDB). Residue side chains were positioned through a combination of rotamer library searches, simulated annealing and minimization. Intermediate models of the 7TM helix bundles were analyzed in terms of helix tilt angles, hydrogen-bond networks, conserved residues and motifs, possible disulfide bonds. The amphipathic cytoplasmic helix domain was also correlated with biological and site-directed mutagenesis data. Finally, the model receptor-binding cavity was characterized using solvent-accessible surface approach.

Amino Acid Sequence↗

Three acidic residues are at the active site of a beta-propeller architecture in glycoside hydrolase families 32, 43, 62, and 68.

Multiple-sequence alignment of glycoside hydrolase (GH) families 32, 43, 62, and 68 revealed three conserved blocks, each containing an acidic residue at an equivalent position in all the enzymes. A detailed analysis of the site-directed mutations so far performed on invertases (GH32), arabinanases (GH43), and bacterial fructosyltransferases (GH68) indicated a direct implication of the conserved residues Asp/Glu (block I), Asp (block II), and Glu (block III) in substrate binding and hydrolysis. These residues are close in space in the 5-bladed beta-propeller fold determined for Cellvibrio japonicus alpha-L-arabinanase Arb43A [Nurizzo et al., Nat Struct Biol 2002;9:665-668] and Bacillus subtilis endo-1,5-alpha-L-arabinanase. A sequence-structure compatibility search using 3D-PSSM, mGenTHREADER, INBGU, and SAM-T02 programs predicted indistinctly the 5-bladed beta-propeller fold of Arb43A and the 6-bladed beta-propeller fold of sialidase/neuraminidase (GH33, GH34, and GH83) as the most reliable topologies for GH families 32, 62, and 68. We conclude that the identified acidic residues are located at the active site of a beta-propeller architecture in GH32, GH43, GH62, and GH68, operating with a canonical reaction mechanism of either inversion (GH43 and likely GH62) or retention (GH32 and GH68) of the anomeric configuration. Also, we propose that the beta-propeller architecture accommodates distinct binding sites for the acceptor saccharide in glycosyl transfer reaction.

Amino Acid Sequence↗

Prediction of transmembrane regions of beta-barrel proteins using ANN- and SVM-based methods.

This article describes a method developed for predicting transmembrane beta-barrel regions in membrane proteins using machine learning techniques: artificial neural network (ANN) and support vector machine (SVM). The ANN used in this study is a feed-forward neural network with a standard back-propagation training algorithm. The accuracy of the ANN-based method improved significantly, from 70.4% to 80.5%, when evolutionary information was added to a single sequence as a multiple sequence alignment obtained from PSI-BLAST. We have also developed an SVM-based method using a primary sequence as input and achieved an accuracy of 77.4%. The SVM model was modified by adding 36 physicochemical parameters to the amino acid sequence information. Finally, ANN- and SVM-based methods were combined to utilize the full potential of both techniques. The accuracy and Matthews correlation coefficient (MCC) value of SVM, ANN, and combined method are 78.5%, 80.5%, and 81.8%, and 0.55, 0.63, and 0.64, respectively. These methods were trained and tested on a nonredundant data set of 16 proteins, and performance was evaluated using "leave one out cross-validation" (LOOCV). Based on this study, we have developed a Web server, TBBPred, for predicting transmembrane beta-barrel regions in proteins (available at http://www.imtech.res.in/raghava/tbbpred).

Algorithms↗

Protein contact prediction using patterns of correlation.

We describe a new method for using neural networks to predict residue contact pairs in a protein. The main inputs to the neural network are a set of 25 measures of correlated mutation between all pairs of residues in two "windows" of size 5 centered on the residues of interest. While the individual pair-wise correlations are a relatively weak predictor of contact, by training the network on windows of correlation the accuracy of prediction is significantly improved. The neural network is trained on a set of 100 proteins and then tested on a disjoint set of 1033 proteins of known structure. An average predictive accuracy of 21.7% is obtained taking the best L/2 predictions for each protein, where L is the sequence length. Taking the best L/10 predictions gives an average accuracy of 30.7%. The predictor is also tested on a set of 59 proteins from the CASP5 experiment. The accuracy is found to be relatively consistent across different sequence lengths, but to vary widely according to the secondary structure. Predictive accuracy is also found to improve by using multiple sequence alignments containing many sequences to calculate the correlations.

Amino Acids↗