Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Predicting absolute contact numbers of native protein structure from amino acid sequence.

The contact number of an amino acid residue in a protein structure is defined by the number of C(beta) atoms around the C(beta) atom of the given residue, a quantity similar to, but different from, solvent accessible surface area. We present a method to predict the contact numbers of a protein from its amino acid sequence. The method is based on a simple linear regression scheme and predicts the absolute values of contact numbers. When single sequences are used for both parameter estimation and cross-validation, the present method predicts the contact numbers with a correlation coefficient of 0.555 on average. When multiple sequence alignments are used, the correlation increases to 0.627, which is a significant improvement over previous methods. In terms of discrete states prediction, the accuracies for 2-, 3-, and 10-state predictions are, respectively, 71.4%, 54.1%, and 18.9% with residue type-dependent unbiased thresholds, and 76.3%, 59.2%, and 21.8% with residue type-independent unbiased thresholds. The difference between accessible surface area and contact number from a prediction viewpoint and the application of contact number prediction to three-dimensional structure prediction are discussed.

Amino Acid Sequence↗

The ConSurf-HSSP database: the mapping of evolutionary conservation among homologs onto PDB structures.

The HSSP (Homology-Derived Secondary Structure of Proteins) database provides multiple sequence alignments (MSAs) for proteins of known three-dimensional (3D) structure in the Protein Data Bank (PDB). The database also contains an estimate of the degree of evolutionary conservation at each amino acid position. This estimate, which is based on the relative entropy, correlates with the functional importance of the position; evolutionarily conserved positions (i.e., positions with limited variability and low entropy) are occasionally important to maintain the 3D structure and biological function(s) of the protein. We recently developed the Rate4Site algorithm for scoring amino acid conservation based on their calculated evolutionary rate. This algorithm takes into account the phylogenetic relationships between the homologs and the stochastic nature of the evolutionary process. Here we present the ConSurf-HSSP database of Rate4Site estimates of the evolutionary rates of the amino acid positions, calculated using HSSP's MSAs. The database provides precalculated evolutionary rates for nearly all of the PDB. These rates are projected, using a color code, onto the protein structure, and can be viewed online using the ConSurf server interface. To exemplify the database, we analyzed in detail the conservation pattern obtained for pyruvate kinase and compared the results with those observed using the relative entropy scores of the HSSP database. It is reassuring to know that the main functional region of the enzyme is detectable using both conservation scores. Interestingly, the ConSurf-HSSP calculations mapped additional functionally important regions, which are moderately conserved and were overlooked by the original HSSP estimate. The ConSurf-HSSP database is available online (http://consurf-hssp.tau.ac.il).

Algorithms↗

Structure and dynamics of the human pleckstrin DEP domain: distinct molecular features of a novel DEP domain subfamily.

Pleckstrin1 is a major substrate for protein kinase C in platelets and leukocytes, and comprises a central DEP (disheveled, Egl-10, pleckstrin) domain, which is flanked by two PH (pleckstrin homology) domains. DEP domains display a unique alpha/beta fold and have been implicated in membrane binding utilizing different mechanisms. Using multiple sequence alignments and phylogenetic tree reconstructions, we find that 6 subfamilies of the DEP domain exist, of which pleckstrin represents a novel and distinct subfamily. To clarify structural determinants of the DEP fold and to gain further insight into the role of the DEP domain, we determined the three-dimensional structure of the pleckstrin DEP domain using heteronuclear NMR spectroscopy. Pleckstrin DEP shares main structural features with the DEP domains of disheveled and Epac, which belong to different DEP subfamilies. However, the pleckstrin DEP fold is distinct from these structures and contains an additional, short helix alpha4 inserted in the beta4-beta5 loop that exhibits increased backbone mobility as judged by NMR relaxation measurements. Based on sequence conservation, the helix alpha4 may also be present in the DEP domains of regulator of G-protein signaling (RGS) proteins, which are members of the same DEP subfamily. In pleckstrin, the DEP domain is surrounded by two PH domains. Structural analysis and charge complementarity suggest that the DEP domain may interact with the N-terminal PH domain in pleckstrin. Phosphorylation of the PH-DEP linker, which is required for pleckstrin function, could regulate such an intramolecular interaction. This suggests a role of the pleckstrin DEP domain in intramolecular domain interactions, which is distinct from the functions of other DEP domain subfamilies found so far.

Amino Acid Sequence↗

Prediction of distant residue contacts with the use of evolutionary information.

In this work we present a novel correlated mutations analysis (CMA) method that is significantly more accurate than previously reported CMA methods. Calculation of correlation coefficients is based on physicochemical properties of residues (predictors) and not on substitution matrices. This results in reliable prediction of pairs of residues that are distant in protein sequence but proximal in its three dimensional tertiary structure. Multiple sequence alignments (MSA) containing a sequence of known structure for 127 families from PFAM database have been selected so that all major protein architectures described in CATH classification database are represented. Protein sequences in the selected families were filtered so that only those evolutionarily close to the target protein remain in the MSA. The average accuracy obtained for the alpha beta class of proteins was 26.8% of predicted proximal pairs with average improvement over random accuracy (IOR) of 6.41. Average accuracy is 20.6% for the mainly beta class and 14.4% for the mainly alpha class. The optimum correlation coefficient cutoff (cc cutoff) was found to be around 0.65. The first predictor, which correlates to hydrophobicity, provides the most reliable results. The other two predictors give good predictions which can be used in conjunction to those of the first one. When stricter cc cutoff is chosen, the average accuracy increases significantly (38.76% for alpha beta class), but the trade off is a smaller number of predictions. The use of solvent accessible area estimations for filtering false positives out of the predictions is promising.

Algorithms↗

S-adenosylhomocysteine hydrolase from the archaeon Pyrococcus furiosus: biochemical characterization and analysis of protein structure by comparative molecular modeling.

S-adenosylhomocysteine hydrolase (AdoHcyHD) is an ubiquitous enzyme that catalyzes the breakdown of S-adenosylhomocysteine, a powerful inhibitor of most transmethylation reactions, to adenosine and L-homocysteine. AdoHcyHD from the hyperthermophilic archaeon Pyrococcus furiosus (PfAdoHcyHD) was cloned, expressed in Escherichia coli, and purified. The enzyme is thermoactive with an optimum temperature of 95 degrees C, and thermostable retaining 100% residual activity after 1 h at 90 degrees C and showing an apparent melting temperature of 98 degrees C. The enzyme is a homotetramer of 190 kDa and contains four cysteine residues per subunit. Thiol groups are not involved in the catalytic process whereas disulfide bond(s) could be present since incubation with 0.8 M dithiothreitol reduces enzyme activity. Multiple sequence alignment of hyperthermophilic AdoHcyHD reveals the presence of two cysteine residues in the N-terminus of the enzyme conserved only in members of Pyrococcus species, and shows that hyperthermophilic AdoHcyHD lack eight C-terminal residues, thought to be important for structural and functional properties of the eukaryotic enzyme. The homology-modeled structure of PfAdoHcyHD shows that Trp220, Tyr181, Tyr184, and Leu185 of each subunit and Ile244 from a different subunit form a network of hydrophobic and aromatic interactions in the central channel formed at the subunits interface. These contacts partially replace the interactions of the C-terminal tail of the eukaryotic enzyme required for tetramer stability. Moreover, Cys221 and Lys245 substitute for Thr430 and Lys426, respectively, of the human enzyme in NAD-binding. Interestingly, all these residues are fairly well conserved in hyperthermophilic AdoHcyHDs but not in mesophilic ones, thus suggesting a common adaptation mechanism at high temperatures.

Adenosine↗

Large-scale prediction of function shift in protein families with a focus on enzymatic function.

Protein function shift can be predicted from sequence comparisons, either using positive selection signals or evolutionary rate estimation. None of the methods have been validated on large datasets, however. Here we investigate existing and novel methods for protein function shift prediction, and benchmark the accuracy against a large dataset of proteins with known enzymatic functions. Function change was predicted between subfamilies by identifying two kinds of sites in a multiple sequence alignment: Conservation-Shifting Sites (CSS), which are conserved in two subfamilies using two different amino acid types, and Rate-Shifting Sites (RSS), which have different evolutionary rates in two subfamilies. CSS were predicted by a new entropy-based method, and RSS using the Rate-Shift program. In principle, the more CSS and RSS between two subfamilies, the more likely a function shift between them. A test dataset was built by extracting subfamilies from Pfam with different EC numbers that belong to the same domain family. Subfamilies were generated automatically using a phylogenetic tree-based program, BETE. The dataset comprised 997 subfamily pairs with four or more members per subfamily. We observed a significant increase in CSS and RSS for subfamily comparisons with different EC numbers compared to cases with same EC numbers. The discrimination was better using RSS than CSS, and was more pronounced for larger families. Combining RSS and CSS by discriminant analysis improved classification accuracy to 71%. The method was applied to the Pfam database and the results are available at http://FunShift.cgb.ki.se. A closer examination of some superfamily comparisons showed that single EC numbers sometimes embody distinct functional classes. Hence, the measured accuracy of function shift is underestimated.

Amino Acid Sequence↗

Predicting protein secondary structure and solvent accessibility with an improved multiple linear regression method.

We have improved the multiple linear regression (MLR) algorithm for protein secondary structure prediction by combining it with the evolutionary information provided by multiple sequence alignment of PSI-BLAST. On the CB513 dataset, the three states average overall per-residue accuracy, Q(3), reached 76.4%, while segment overlap accuracy, SOV99, reached 73.2%, using a rigorous jackknife procedure and the strictest reduction of eight states DSSP definition to three states. This represents an improvement of approximately 5% on overall per-residue accuracy compared with previous work. The relative solvent accessibility prediction also benefited from this combination of methods. The system achieved 77.7% average jackknifed accuracy for two states prediction based on a 25% relative solvent accessibility mode, with a Mathews' correlation coefficient of 0.548. The improved MLR secondary structure and relative solvent accessibility prediction server is available at http://spg.biosci.tsinghua.edu.cn/.

Algorithms↗

The Histone Database: a comprehensive resource for histones and histone fold-containing proteins.

The Histone Database is a curated and searchable collection of full-length sequences and structures of histones and nonhistone proteins containing histone-like folds, compiled from major public databases. Several new histone fold-containing proteins have been identified, including the huntingtin-interacting protein HYPM. Additionally, based on the recent crystal structure of the Son of Sevenless protein, an interpretation of the sequence analysis of the histone fold domain is presented. The database contains an updated collection of multiple sequence alignments for the four core histones (H2A, H2B, H3, and H4) and the linker histones (H1/H5) from a total of 975 organisms. The database also contains information on the human histone gene complement and provides links to three-dimensional structures of histone and histone fold-containing proteins. The Histone Database is a comprehensive bioinformatics resource for the study of structure and function of histones and histone fold-containing proteins. The database is available at http://research.nhgri.nih.gov/histones/.

Amino Acid Sequence↗

Evolutionary coupling of structural and functional sequence information in the intracellular lipid-binding protein family.

We have mined the evolutionary record for the large family of intracellular lipid-binding proteins (iLBPs) by calculating the statistical coupling of residue variations in a multiple sequence alignment using methods developed by Ranganathan and coworkers (Lockless and Ranganathan, Science 1999:286;295-299). The 213 sequences analyzed have a wide range of ligand-binding functions as well as highly divergent phylogenetic origins, assuring broad sampling of sequence space. Emerging from this analysis were two major clusters of coupled residues, which when mapped onto the structure of a representative iLBP under study in our laboratory, cellular retinoic-acid binding protein I, are largely contiguous and provide useful points of comparison to available data for the folding of this protein. One cluster comprises a predominantly hydrophobic core away from the ligand-binding site and likely represents key structural information for the iLBP fold. The other cluster includes the portal region where ligand enters its binding site, regions of the ligand-binding cavity, and the region where the 10-stranded beta-barrel characteristic of this family closes (between strands 1' and 10). Linkages between these two clusters suggest that evolutionary pressures on this family constrain structural and functional sequence information in an interdependent fashion. The necessity of the structure to wrap around a hydrophobic ligand confounds the typical sequestration of hydrophobic side chains. Additionally, ligand entry and exit require these structures to have a capacity for specific conformational change during binding and release. We conclude that an essential and structurally apparent separation of local and global sequence information is conserved throughout the iLBP family.

Amino Acid Sequence↗

Evolutionary and structural feedback on selection of sequences for comparative analysis of proteins.

It has been noted that slowly evolving protein residues have two properties: (a) they tend to cluster in the native fold, and (b) they delineate functional surfaces-parts of the surface through which the protein interacts with other proteins or small ligands. Herein, we demonstrate that the two are coupled sufficiently strongly that one effect, when observed, statistically implies the other. Detection of both can be accomplished in multiple sequence alignment related methods by the careful selection of relevant sequences. For the demonstration, we use two sets of protein families: a small set of diverse proteins with diverse functional surfaces, and a large set of homodimerizing enzymes. A practical outcome of our considerations is a simple prescriptive rule for the selection of homologous sequences for the comparative analysis of proteins: in order to optimize the detection of (potentially unknown) functional surfaces, it is sufficient to select sequences in such a way that the residues observed at any level of evolutionary divergence, as implied by the alignment, cluster on the folded protein.

Animals↗

Codep: maximizing co-evolutionary interdependencies to discover interacting proteins.

Approaches for the determination of interacting partners from different protein families (such as ligands and their receptors) have made use of the property that interacting proteins follow similar patterns and relative rates of evolution. Interacting protein partners can then be predicted from the similarity of their phylogenetic trees or evolutionary distances matrices. We present a novel method called Codep, for the determination of interacting protein partners by maximizing co-evolutionary signals. The order of sequences in the multiple sequence alignments from two protein families is determined in such a manner as to maximize the similarity of substitution patterns at amino acid sites in the two alignments and, thus, phylogenetic congruency. This is achieved by maximizing the total number of interdependencies of amino acids sites between the alignments. Once ordered, the corresponding sequences in the two alignments indicate the predicted interacting partners. We demonstrate the efficacy of this approach with computer simulations and in analyses of several protein families. A program implementing our method, Codep, is freely available to academic users from our website: http://www.uhnresearch.ca/labs/tillier/.

Computer Simulation↗

Comparative model of EutB from coenzyme B12-dependent ethanolamine ammonia-lyase reveals a beta8alpha8, TIM-barrel fold and radical catalytic site structural features.

The structure of the EutB protein from Salmonella typhimurium, which contains the active site of the coenzyme B12 (adenosylcobalamin)-dependent enzyme, ethanolamine ammonia-lyase, has been predicted by using structural proteomics techniques of comparative modelling. The 453-residue EutB protein displays no significant sequence identity with proteins of known structure. Therefore, secondary structure prediction and fold recognition algorithms were used to identify templates. Multiple three-dimensional template matching (threading) servers identified predominantly beta8alpha8, TIM-barrel proteins, and in particular, the large subunits of diol dehydratase (PDB: 1eex:A, 1dio:A) and glycerol dehydratase (PDB: 1mmf:A), as templates. Consistent with this identification, the dehydratases are, like ethanolamine ammonia-lyase, Class II coenzyme B12-dependent enzymes. Model building was performed by using MODELLER. Models were evaluated by using different programs, including PROCHECK and VERIFY3D. The results identify a beta8alpha8, TIM-barrel fold for EutB. The beta8alpha8, TIM-barrel fold is consistent with a central role of the alpha/beta-barrel structures in radical catalysis conducted by the coenzyme B12- and S-adenosylmethionine-dependent (radical SAM) enzyme superfamilies. The EutB model and multiple sequence alignment among ethanolamine ammonia-lyase, diol dehydratase, and glycerol dehydratase from different species reveal the following protein structural features: (1) a "cap" loop segment that closes the N-terminal region of the barrel, (2) a common cobalamin cofactor binding topography at the C-terminal region of the barrel, and (3) a beta-barrel-internal guanidinium group from EutB R160 that overlaps the position of the active-site potassium ion found in the dehydratases. R160 is proposed to have a role in substrate binding and radical catalysis.

Amino Acid Sequence↗

Better prediction of the location of alpha-turns in proteins with support vector machine.

We have developed a novel method named AlphaTurn to predict alpha-turns in proteins based on the support vector machine (SVM). The prediction was done on a data set of 469 nonhomologous proteins containing 967 alpha-turns. A great improvement in prediction performance was achieved by using multiple sequence alignment generated by PSI-BLAST as input instead of the single amino acid sequence. The introduction of secondary structure information predicted by PSIPRED also improved the prediction performance. Moreover, we handled the very uneven data set by combining the cost factor j with the "state-shifting" rule. This further promoted the prediction quality of our method. The final SVM model yielded a Matthews correlation coefficient (MCC) of 0.25 by a 10-fold cross-validation. To our knowledge, this MCC value is the highest obtained so far for predicting alpha-turns. An online Web server based on this method has been developed and can be freely accessed at http://bmc.hust.edu.cn/bioinformatics/ or http://210.42.106.80/.

Amino Acid Sequence↗

Achieving 80% ten-fold cross-validated accuracy for secondary structure prediction by large-scale training.

An integrated system of neural networks, called SPINE, is established and optimized for predicting structural properties of proteins. SPINE is applied to three-state secondary-structure and residue-solvent-accessibility (RSA) prediction in this paper. The integrated neural networks are carefully trained with a large dataset of 2640 chains, sequence profiles generated from multiple sequence alignment, representative amino acid properties, a slow learning rate, overfitting protection, and an optimized sliding-widow size. More than 200,000 weights in SPINE are optimized by maximizing the accuracy measured by Q(3) (the percentage of correctly classified residues). SPINE yields a 10-fold cross-validated accuracy of 79.5% (80.0% for chains of length between 50 and 300) in secondary-structure prediction after one-month (CPU time) training on 22 processors. An accuracy of 87.5% is achieved for exposed residues (RSA >95%). The latter approaches the theoretical upper limit of 88-90% accuracy in assigning secondary structures. An accuracy of 73% for three-state solvent-accessibility prediction (25%/75% cutoff) and 79.3% for two-state prediction (25% cutoff) is also obtained.

Algorithms↗

Correlated mutations and residue contacts in proteins.

The maintenance of protein function and structure constrains the evolution of amino acid sequences. This fact can be exploited to interpret correlated mutations observed in a sequence family as an indication of probable physical contact in three dimensions. Here we present a simple and general method to analyze correlations in mutational behavior between different positions in a multiple sequence alignment. We then use these correlations to predict contact maps for each of 11 protein families and compare the result with the contacts determined by crystallography. For the most strongly correlated residue pairs predicted to be in contact, the prediction accuracy ranges from 37 to 68% and the improvement ratio relative to a random prediction from 1.4 to 5.1. Predicted contact maps can be used as input for the calculation of protein tertiary structure, either from sequence information alone or in combination with experimental information.

Amino Acid Sequence↗

Progress of 1D protein structure prediction at last.

Accuracy of predicting protein secondary structure and solvent accessibility from sequence information has been improved significantly by using information contained in multiple sequence alignments as input to a neural network system. For the Asilomar meeting, predictions for 13 proteins were generated automatically using the publicly available prediction method PHD. The results confirm the estimate of 72% three-state prediction accuracy. The fairly accurate predictions of secondary structure segments made the tool useful as a starting point for modeling of higher dimensional aspects of protein structure.

Amino Acid Sequence↗

Molecular dynamics simulation of human neurohypophyseal hormone receptors complexed with oxytocin-modeling of an activated state.

The neurohypophyseal hormone oxytocin (CYIQNCPLG-NH(2), OT) is involved in the control of labor, secretion of milk and many social and behavioral functions via interaction with its receptors (OTR) located in the uterus, mammary glands and peripheral tissues, respectively. In this paper we propose the interactions responsible for OT binding and selectivity to OTR versus vasopressin ([F3,R8]OT, AVP) receptors: V1aR and V2R, all three belonging to the Class A G protein-coupled receptors (GPCRs). Three-dimensional models of the activated receptors were constructed using a multiple sequence alignment and the activated rhodopsin-transducin (MII-Gt) prototype [Slusarz and Ciarkowski, 2004] as a template. The 1 ns unconstrained molecular dynamics (MD) of three pairs of receptor-OT complexes (two complexes per each receptor) immersed in the fully hydrated 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphatidylcholine (POPC) lipid bilayer was conducted in the AMBER 7.0 force field. The relaxed models of ligand-receptor complexes were used to identify the putative binding sites of OT. The stabilizing interactions with conserved Gln residues in all complexes were identified. The nonconserved hydrophobic residues were proposed as responsible for OTR-OT selectivity and ligand recognition. These results provide guidelines for experimental site-directed mutagenesis and if confirmed, they may be helpful in designing new selective OT analogs with both agonistic or antagonistic properties.

Amino Acid Sequence↗

Analysis of interactions responsible for vasopressin binding to human neurohypophyseal hormone receptors-molecular dynamics study of the activated receptor-vasopressin-G(alpha) systems.

Vasopressin (CYFQNCPRG-NH(2), AVP) is a semicyclic endogenous peptide, which exerts a variety of biological effects in mammals. The main physiological roles of AVP are the regulation of water balance and the control of blood pressure and adrenocorticotropin hormone (ACTH) secretion, mediated via three different subtypes of vasopressin receptors: V1a, V1b and V2 receptors (V1aR, V1bR and V2R, respectively). They are the members of the class A, G-protein-coupled receptors (GPCRs). AVP also modulates several behavioral and social functions. In this study, the interactions responsible for AVP binding to vasopressin V1a and V2 receptors versus the closely related oxytocin ([I3,L8]AVP, OT) receptor (OTR) have been investigated. Three-dimensional models of the activated receptors were constructed using multiple sequence alignment, followed by homology modeling using the complex of activated rhodopsin with Gt(alpha) C-terminal peptide of transducin MII-Gt(338-350) prototype as a template. AVP was docked into the receptor-G(alpha) systems. The three lowest-energy pairs of receptor-AVP-G(alpha) (two complexes per each receptor) were selected. The 1-ns unconstrained molecular dynamics (MD) of complexes embedded into the fully hydrated 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphatidylcholine (POPC) lipid bilayer was conducted in the AMBER 7.0 force field. Six relaxed receptor-AVP-G(alpha) models were obtained. The residues responsible for AVP binding to vasopressin receptors have been identified and a different mechanism of AVP binding to V2R than to V1aR has been proposed.

Amino Acid Sequence↗