Search PubMed⌕ Search

Biomedical subjects

J Skolnick

Publications and source records attributed to J Skolnick.

At least 55 records · Page 3Linked to original sources

What should the Z-score of native protein structures be?

The Z-score of a protein is defined as the energy separation between the native fold and the average of an ensemble of misfolds in the units of the standard deviation of the ensemble. The Z-score is often used as a way of testing the knowledge-based potentials for their ability to recognize the native fold from other alternatives. However, it is not known what range of values the Z-scores should have if one had a correct potential. Here, we offer an estimate of Z-scores extracted from calorimetric measurements of proteins. The energies obtained from these experimental data are compared with those from computer simulations of a lattice model protein. It is suggested that the Z-scores calculated from different knowledge-based potentials are generally too small in comparison with the experimental values.

Models, Chemical↗

Computer simulations of de novo designed helical proteins.

In the context of reduced protein models, Monte Carlo simulations of three de novo designed helical proteins (four-member helical bundle) were performed. At low temperatures, for all proteins under consideration, protein-like folds having different topologies were obtained from random starting conformations. These simulations are consistent with experimental evidence indicating that these de novo designed proteins have the features of a molten globule state. The results of Monte Carlo simulations suggest that these molecules adopt four-helix bundle topologies. They also give insight into the possible mechanism of folding and association, which occurs in these simulations by on-site assembly of the helices. The low-temperature conformations of all three sequences have the features of a molten globule state.

Amino Acid Sequence↗

What is the probability of a chance prediction of a protein structure with an rmsd of 6 A?

BACKGROUND: The root mean square deviation (rmsd) between corresponding atoms of two protein chains is a commonly used measure of similarity between two protein structures. The smaller the rmsd is between two structures, the more similar are these two structures. In protein structure prediction, one needs the rmsd between predicted and experimental structures for which a prediction can be considered to be successful. Success is obvious only when the rmsd is as small as that for closely homologous proteins (< 3 A). To estimate the quality of the prediction in the more general case, one has to compare the native structure not only with the predicted one but also with randomly chosen protein-like folds. One can ask: how many such structures must be considered to find a structure with a given rmsd from the native structure? RESULTS: We calculated the rmsd values between native structures of 142 proteins and all compact structures obtained in the threading of these protein chains over 364 non-homologous structures. The rmsd distributions have a Gaussian form, with the average rmsd approximately proportional to the radius of gyration. CONCLUSIONS: We estimated the number of protein-like structures required to obtain a structure within an rmsd of 6 A to be 10(4)-10(5) for chains of 60-80 residues and 10(11)-10(12) structures for chains of 160-200 residues. The probability of obtaining a 6 A rmsd by chance is so remote that when such structures are obtained from a prediction algorithm, it should be considered quite successful.

Databases, Factual↗

Functional analysis of the Escherichia coli genome for members of the alpha/beta hydrolase family.

BACKGROUND: Database-searching methods based on sequence similarity have become the most commonly used tools for characterizing newly sequenced proteins. Due to the often underestimated functional diversity in protein families and superfamilies, however, it is difficult to make the characterization specific and accurate. In this work, we have extended a method for active-site identification from predicted protein structures. RESULTS: The structural conservation and variation of the active sites of the alpha/beta hydrolases with known structures were studied. The similarities were incorporated into a three-dimensional motif that specifies essential requirements for the enzymatic functions. A threading algorithm was used to align 651 Escherichia coli open reading frames (ORFs) to one of the members of the alpha/beta hydrolase fold family. These ORFs were then screened according to our three-dimensional motif and with an extra requirement that demands conservation of the key active-site residues among the proteins that bear significant sequence similarity to the ORFs. 17 ORFs from E. coli were predicted to have hydrolase activity and their putative active-site residues were identified. Most were in agreement with the experiments and results of other database-searching methods. The study further suggests that YHET_ECOLI, a hypothetical protein classified as a member of the UPF0017 family (an uncharacterized protein family), bears all the hallmarks of the alpha/beta hydrolase family. CONCLUSIONS: The novel feature of our method is that it uses three-dimensional structural information for function prediction. The results demonstrate the importance and necessity of such a method to fill the gap between sequence alignment and function prediction; furthermore, the method provides a way to verify the structure predictions, which enables an expansion of the applicable scope of the threading algorithms.

Amino Acid Sequence↗

Application of an artificial neural network to predict specific class I MHC binding peptide sequences.

Computational methods were used to predict the sequences of peptides that bind to the MHC class I molecule, K(b). The rules for predicting binding sequences, which are limited, are based on preferences for certain amino acids in certain positions of the peptide. It is apparent though, that binding can be influenced by the amino acids in all of the positions of the peptide. An artificial neural network (ANN) has the ability to simultaneously analyze the influence of all of the amino acids of the peptide and thus may improve binding predictions. ANNs were compared to statistically analyzed peptides for their abilities to predict the sequences of K(b) binding peptides. ANN systems were trained on a library of binding and nonbinding peptide sequences from a phage display library. Statistical and ANN methods identified strong binding peptides with preferred amino acids. ANNs detected more subtle binding preferences, enabling them to predict medium binding peptides. The ability to predict class I MHC molecule binding peptides is useful for immunolological therapies involving cytotoxic-T cells.

Amino Acids↗

Reduced protein models and their application to the protein folding problem.

One of the most important unsolved problems of computational biology is prediction of the three-dimensional structure of a protein from its amino acid sequence. In practice, the solution to the protein folding problem demands that two interrelated problems be simultaneously addressed. Potentials that recognize the native state from the myriad of misfolded conformations are required, and the multiple minima conformational search problem must be solved. A means of partly surmounting both problems is to use reduced protein models and knowledge-based potentials. Such models have been employed to elucidate a number of general features of protein folding, including the nature of the energy landscape, the factors responsible for the uniqueness of the native state and the origin of the two-state thermodynamic behavior of globular proteins. Reduced models have also been used to predict protein tertiary and quaternary structure. When combined with a limited amount of experimental information about secondary and tertiary structure, molecules of substantial complexity can be assembled. If predicted secondary structure and tertiary restraints are employed, low resolution models of single domain proteins can be successfully predicted. Thus, simplified protein models have played an important role in furthering the understanding of the physical properties of proteins.

CREB-Binding Protein↗

Optimization of protein structure on lattices using a self-consistent field approach.

Lattice modeling of proteins is commonly used to study the protein folding problem. The reduced number of possible conformations of lattice models enormously facilitates exploration of the conformational space. In this work, we suggest a method to search for the optimal lattice models that reproduced the off-lattice structures with minimal errors in geometry and energetics. The method is based on the self-consistent field optimization of a combined pseudoenergy function that includes two force fields: an "interaction field," that drives the residues to optimize the chain energy, and a "geometrical field," that attracts the residues towards their native positions. By varying the contributions of these force fields in the combined pseudoenergy, one can also test the accuracy of potentials: the better the potentials, i.e., the more accurate the "interaction field," and the smaller the contribution of the "geometrical field" required for building accurate lattice models.

Models, Chemical↗

Combined multiple sequence reduced protein model approach to predict the tertiary structure of small proteins.

By incorporating predicted secondary and tertiary restraints into ab initio folding simulations, low resolution tertiary structures of a test set of 20 nonhomologous proteins have been predicted. These proteins, which represent all secondary structural classes, contain from 37 to 100 residues. Secondary structural restraints are provided by the PHD secondary structure prediction algorithm that incorporates multiple sequence information. Predicted tertiary restraints are obtained from multiple sequence alignments via a two-step process: First, "seed" side chain contacts are identified from a correlated mutation analysis, and then, the seed contacts are "expanded" by an inverse folding algorithm. These predicted restraints are then incorporated into a lattice based, reduced protein model. Depending upon fold complexity, the resulting nativelike topologies exhibit a coordinate root-mean-square deviation, cRMSD, from native between 3.1 and 6.7 A. Overall, this study suggests that the use of restraints derived from multiple sequence alignments combined with a fold assembly algorithm is a promising approach to the prediction of the global topology of small proteins.

Algorithms↗

MONSSTER: a method for folding globular proteins with a small number of distance restraints.

The MONSSTER (MOdeling of New Structures from Secondary and TEritary Restraints) method for folding of proteins using a small number of long-distance restraints (which can be up to seven times less than the total number of residues) and some knowledge of the secondary structure of regular fragments is described. The method employs a high-coordination lattice representation of the protein chain that incorporates a variety of potentials designed to produce protein-like behaviour. These include statistical preferences for secondary structure, side-chain burial interactions, and a hydrogen-bond potential. Using this algorithm, several globular proteins (1ctf, 2gbl, 2trx, 3fxn, 1mba, 1pcy and 6pti) have been folded to moderate-resolution, native-like compact states. For example, the 68 residue 1ctf molecule having ten loosely defined, long-range restraints was reproducibly obtained with a C alpha-backbone root-mean-square deviation (RMSD) from native of about 4. A. Flavodoxin with 35 restraints has been folded to structures whose average RMSD is 4.28 A. Furthermore, using just 20 restraints, myoglobin, which is a 146 residue helical protein, has been folded to structures whose average RMSD from native is 5.65 A. Plastocyanin with 25 long-range restraints adopts conformations whose average RMSD is 5.44 A. Possible applications of the proposed approach to the refinement of structures from NMR data, homology model-building and the determination of tertiary structure when the secondary structure and a small number of restraints are predicted are briefly discussed.

Algorithms↗

Improved method for prediction of protein backbone U-turn positions and major secondary structural elements between U-turns.

A new and more accurate method has been developed for predicting the backbone U-turn positions (where the chain reverses global direction) and the dominant secondary structure elements between U-turns in globular proteins. The current approach uses sequence-specific secondary structure propensities and multiple sequence information. The latter plays an important role in the enhanced success of this approach. Application to two sets (total 108) of small to medium-sized, single-domain proteins indicates that approximately 94% of the U-turn locations are correctly predicted within three residues, as are 88% of dominant secondary structure elements. These results are significantly better than our previous method (Kolinski et al., Proteins 27:290-308, 1997). The current study strongly suggests that the U-turn locations are primarily determined by local interactions. Furthermore, both global length constraints and local interactions contribute significantly to the determination of the secondary structure types between U-turns. Accurate U-turn predictions are crucial for accurate secondary structure predictions in the current method. Protein structure modeling, tertiary structure predictions, and possibly, fold recognition should benefit from the predicted structural data provided by this new method.

Amino Acid Sequence↗

Derivation and testing of pair potentials for protein folding. When is the quasichemical approximation correct?

Many existing derivations of knowledge-based statistical pair potentials invoke the quasichemical approximation to estimate the expected side-chain contact frequency if there were no amino acid pair-specific interactions. At first glance, the quasichemical approximation that treats the residues in a protein as being disconnected and expresses the side-chain contact probability as being proportional to the product of the mole fractions of the pair of residues would appear to be rather severe. To investigate the validity of this approximation, we introduce two new reference states in which no specific pair interactions between amino acids are allowed, but in which the connectivity of the protein chain is retained. The first estimates the expected number of side-chain contracts by treating the protein as a Gaussian random coil polymer. The second, more realistic reference state includes the effects of chain connectivity, secondary structure, and chain compactness by estimating the expected side-chain contrast probability by placing the sequence of interest in each member of a library of structures of comparable compactness to the native conformation. The side-chain contact maps are not allowed to readjust to the sequence of interest, i.e., the side chains cannot repack. This situation would hold rigorously if all amino acids were the same size. Both reference states effectively permit the factorization of the side-chain contact probability into sequence-dependent and structure-dependent terms. Then, because the sequence distribution of amino acids in proteins is random, the quasichemical approximation to each of these reference states is shown to be excellent. Thus, the range of validity of the quasichemical approximation is determined by the magnitude of the side-chain repacking term, which is, at present, unknown. Finally, the performance of these two sets of pair interaction potentials as well as side-chain contact fraction-based interaction scales is assessed by inverse folding tests both without and with allowing for gaps.

Models, Chemical↗

Simultaneous and coupled energy optimization of homologous proteins: a new tool for structure prediction.

BACKGROUND: Homology-based modeling and global optimization of energy are two complementary approaches to prediction of protein structures. A combination of the two approaches is proposed in which a novel component is added to the energy and forces similarity between homologous proteins. RESULTS: The combination was tested for two families: pancreatic hormones and homeodomains. The simulated lowest-energy structure of the pancreatic hormones is a reasonable approximation to the native fold. The lowest-energy structure of the homeodomains has 80% of the native contacts, but the helices are not packed correctly. The fourth lowest energy structure of the homeodomains has the correct helix packing (RMS 5.4 A and 82% of the correct contacts). Optimizations of a single protein of the family yield considerably worse structures. CONCLUSIONS: Use of coupled homologous proteins in the search for the native fold is more successful than the folding of a single protein in the family.

Amino Acid Sequence↗

Recognition of protein structure on coarse lattices with residue-residue energy functions.

We suggest and test potentials for the modeling of protein structure on coarse lattices. The coarser the lattice, the more complete and faster is the exploration of the conformational space of a molecule. However, there are inevitable energy errors in lattice modeling caused by distortions in distances between interacting residues; the coarser the lattice, the larger are the energy errors. It is generally believed that an improvement in the accuracy of lattice modelling can be achieved only by reducing the lattice spacing. We reduce the errors on coarse lattices with lattice-adapted potentials. Two methods are used: in the first approach, 'lattice-derived' potentials are obtained directly from a database of lattice models of protein structure; in the second approach, we derive 'lattice-adjusted' potentials using our previously developed method of statistical adjustment of the 'off-lattice' energy functions for lattices. The derivation of off-lattice Calpha atom-based distance-dependent pairwise potentials has been reported previously. The accuracy of 'lattice-derived', 'lattice-adjusted' and 'off-lattice' potentials is estimated in threading tests. It is shown that 'lattice-derived' and 'lattice-adjusted' potentials give virtually the same accuracy and ensure reasonable protein fold recognition on the coarsest considered lattice (spacing 3.8 A), however, the 'off-lattice' potentials, which efficiently recognize off-lattice folds, do not work on this lattice, mainly because of the errors in short-range interactions between neighboring residues.

Models, Chemical↗

Sequence-structure specificity--how does an inverse folding approach work?

The inverse folding approach is a powerful tool in protein structure prediction when the native state of a sequence adopts one of the known protein folds. This is because some proteins show strong sequence-structure specificity in inverse folding experiments that allow gaps and insertions in the sequence-structure alignment. In those cases when structures similar to their native folds are included in the structure database, the z-scores (which measure the sequence-structure specificity) of these folds are well separated from those of other alternative structures. In this paper, we seek to understand the origin of this sequence-structure specificity and to identify how the specificity arises on passing from a short peptide chain to the entire protein sequence. To accomplish this objective, a simplified version of inverse folding, gapless inverse folding, is performed using sequence fragments of different sizes from 53 proteins. The results indicate that usually a significant portion of the entire protein sequence is necessary to show sequence-structure specificity, but there are regions in the sequence that begin to show this specificity at relatively short fragment size (15-20 residues). An island picture, in which the regions in the sequence that recognize their own native structure grow from some seed fragments, is observed as the fragment size increases. Usually, more similar structures to the native states are found in the top-scoring structural fragments in these high-specificity regions.

Amino Acid Sequence↗

A method for the prediction of surface "U"-turns and transglobular connections in small proteins.

A simple method for predicting the location of surface loops/turns that change the overall direction of the chain that is, "U" turns, and assigning the dominant secondary structure of the intervening transglobular blocks in small, single-domain globular proteins has been developed. Since the emphasis of the method is on the prediction of the major topological elements that comprise the global structure of the protein rather than on a detailed local secondary structure description, this approach is complementary to standard secondary structure prediction schemes. Consequently, it may be useful in the early stages of tertiary structure prediction when establishment of the structural class and possible folding topologies is of interest. Application to a set of small proteins of known structure indicates a high level of accuracy. The prediction of the approximate location of the surface turns/loops that are responsible for the change in overall chain direction is correct in more than 95% of the cases. The accuracy for the dominant secondary structure assignment for the linear blocks between such surface turns/loops is in the range of 82%.

Algorithms↗

Method for low resolution prediction of small protein tertiary structure.

A new method for the de novo prediction of protein structures at low resolution has been developed. Starting from a multiple sequence alignment, protein secondary structure is predicted, and only those topological elements with high reliability are selected. Then, the multiple sequence alignment and the secondary structure prediction are combined to predict side chain contacts. Such contact map prediction is carried out in two stages. First, an analysis of correlated mutations is carried out to identify pairs of topological elements of secondary structure which are in contact. Then, inverse folding is used to select compatible fragments in contact, thereby enriching the number and identity of predicted side chain contacts. The final outcome of the procedure is a set of noisy secondary and tertiary restraints. These are used as a restrained potential in a Monte Carlo simulation of simplified protein models driven by statistical potentials. Low energy structures are then searched for by using simulated annealing techniques. Implementation of the restraints is carried out so as to take into account of their low resolution. Using this procedure, it has been possible to predict de novo the structure of three very different protein topologies: an alpha/beta protein, the bovine pancreatic trypsin inhibitor (6pti), an alpha-helical protein, calbindin (3icb), and an all beta- protein, the SH3 domain of spectrin (1shg). In all cases, low resolution folds have been obtained with a root mean square deviation (RMSD) of 4.5-5.5 A with respect to the native structure. Some misfolded topologies appear in the simulations, but it is possible to select the native one on energetic grounds. Thus, it is demonstrated that the methodology is general for all protein motifs. Work is in progress in order to test the methodology on a larger set of protein structures.

Amino Acid Sequence↗

High coordination lattice models of protein structure, dynamics and thermodynamics.

A high coordination lattice discretization of protein conformational space is described. The model allows discrete representation of polypeptide chains of globular proteins and small macromolecular assemblies with an accuracy comparable to the accuracy of crystallographic structures. Knowledge based force field, that consists of sequence specific short range interactions, cooperative model of hydrogen bond network and tertiary one body, two body and multibody interactions, is outlined and discussed. A model of stochastic dynamics for these protein models is also described. The proposed method enables moderate resolution tertiary structure prediction of simple and small globular proteins. Its applicability in structure prediction increases significantly when evolutionary information is exploited or/and when sparse experimental data are available. The model responds correctly to sequence mutations and could be used at early stages of a computer aided protein design and protein redesign. Computational speed, associated with the discrete structure of the model, enables studies of the long time dynamics of polypeptides and proteins and quite detailed theoretical studies of thermodynamics of nontrivial protein models.

Amino Acid Sequence↗

Method for predicting the state of association of discretized protein models. Application to leucine zippers.

A method that employs a transfer matrix treatment combined with Monte Carlo sampling has been used to calculate the configurational free energies of folded and unfolded states of lattice models of proteins. The method is successfully applied to study the monomer-dimer equilibria in various coiled coils. For the short coiled coils, GCN4 leucine zipper, and its fragments, Fos and Jun, very good agreement is found with experiment. Experimentally, some subdomains of the GCN4 leucine zipper form stable dimeric structures, suggesting the regions of differential stability in the parent structure. Our calculations suggest that the stabilities of the subdomains are in general different from the values expected simply from the stability of the corresponding fragment in the wild type molecule. Furthermore, parts of the fragments structurally rearrange in some regions with respect to their corresponding wild type positions. Our results suggest for an Asn in the dimerization interface at least a pair of hydrophobic interacting helical turns at each side is required to stabilize the stable coiled coil. Finally, the specificity of heterodimer formation in the Fos-Jun system comes from the relative instability of Fos homodimers, resulting from unfavorable intra- and interhelical interactions in the interfacial coiled coil region.

Amino Acid Sequence↗