Search PubMed⌕ Search

Biomedical subjects

Jeffrey Skolnick

Publications and source records attributed to Jeffrey Skolnick.

44 records · Page 3Linked to original sources

Protein fragment reconstruction using various modeling techniques.

Recently developed reduced models of proteins with knowledge-based force fields have been applied to a specific case of comparative modeling. From twenty high resolution protein structures of various structural classes, significant fragments of their chains have been removed and treated as unknown. The remaining portions of the structures were treated as fixed - i.e., as templates with an exact alignment. Then, the missed fragments were reconstructed using several modeling tools. These included three reduced types of protein models: the lattice SICHO (Side Chain Only) model, the lattice CABS (Calpha + Cbeta + Side group) model and an off-lattice model similar to the CABS model and called REFINER. The obtained reduced models were compared with more standard comparative modeling tools such as MODELLER and the SWISS-MODEL server. The reduced model results are qualitatively better for the higher resolution lattice models, clearly suggesting that these are now mature, competitive and complementary (in the range of sparse alignments) to the classical tools of comparative modeling. Comparison between the various reduced models strongly suggests that the essential ingredient for the sucessful and accurate modeling of protein structures is not the representation of conformational space (lattice, off-lattice, all-atom) but, rather, the specificity of the force fields used and, perhaps, the sampling techniques employed. These conclusions are encouraging for the future application of the fast reduced models in comparative modeling on a genomic scale.

Amino Acid Sequence↗

Multimeric threading-based prediction of protein-protein interactions on a genomic scale: application to the Saccharomyces cerevisiae proteome.

MULTIPROSPECTOR, a multimeric threading algorithm for the prediction of protein-protein interactions, is applied to the genome of Saccharomyces cerevisiae. Each possible pairwise interaction among more than 6000 encoded proteins is evaluated against a dimer database of 768 complex structures by using a confidence estimate of the fold assignment and the magnitude of the statistical interfacial potentials. In total, 7321 interactions between pairs of different proteins are predicted, based on 304 complex structures. Quality estimation based on the coincidence of subcellular localizations and biological functions of the predicted interactors shows that our approach ranks third when compared with all other large-scale methods. Unlike other in silico methods, MULTIPROSPECTOR is able to identify the residues that participate directly in the interaction. Three hundred seventy-four of our predictions can be found by at least one of the other studies, which is compatible with the overlap between two different other methods. From the analysis of the mRNA abundance data, our method does not bias towards proteins with high abundance. Finally, several relevant predictions involved in various functions are presented. In summary, we provide a novel approach to predict protein-protein interactions on a genomic scale that is a useful complement to experimental methods.

DNA, Fungal↗

MULTIPROSPECTOR: an algorithm for the prediction of protein-protein interactions by multimeric threading.

In this postgenomic era, the ability to identify protein-protein interactions on a genomic scale is very important to assist in the assignment of physiological function. Because of the increasing number of solved structures involving protein complexes, the time is ripe to extend threading to the prediction of quaternary structure. In this spirit, a multimeric threading approach has been developed. The approach is comprised of two phases. In the first phase, traditional threading on a single chain is applied to generate a set of potential structures for the query sequences. In particular, we use our recently developed threading algorithm, PROSPECTOR. Then, for those proteins whose template structures are part of a known complex, we rethread on both partners in the complex and now include a protein-protein interfacial energy. To perform this analysis, a database of multimeric protein structures has been constructed, the necessary interfacial pairwise potentials have been derived, and a set of empirical indicators to identify true multimers based on the threading Z-score and the magnitude of the interfacial energy have been established. The algorithm has been tested on a benchmark set comprised of 40 homodimers, 15 heterodimers, and 69 monomers that were scanned against a protein library of 2478 structures that comprise a representative set of structures in the Protein Data Bank. Of these, the method correctly recognized and assigned 36 homodimers, 15 heterodimers, and 65 monomers. This protocol was applied to identify partners and assign quaternary structures of proteins found in the yeast database of interacting proteins. Our multimeric threading algorithm correctly predicts 144 interacting proteins, compared to the 56 (26) cases assigned by PSI-BLAST using a (less) permissive E-value of 1 (0.01). Next, all possible pairs of yeast proteins have been examined. Predictions (n = 2865) of protein-protein interactions are made; 1138 of these 2865 interactions have counterparts in the Database of Interacting Proteins. In contrast, PSI-BLAST made 1781 predictions, and 1215 have counterparts in DIP. An estimation of the false-negative rate for yeast-predicted interactions has also been provided. Thus, a promising approach to help assist in the assignment of protein-protein interactions on a genomic scale has been developed.

Algorithms↗

Local energy landscape flattening: parallel hyperbolic Monte Carlo sampling of protein folding.

Among the major difficulties in protein structure prediction is the roughness of the energy landscape that must be searched for the global energy minimum. To address this issue, we have developed a novel Monte Carlo algorithm called parallel hyperbolic sampling (PHS) that logarithmically flattens local high-energy barriers and, therefore, allows the simulation to tunnel more efficiently through energetically inaccessible regions to low-energy valleys. Here, we show the utility of this approach by applying it to the SICHO (SIde-CHain-Only) protein model. For the same CPU time, the parallel hyperbolic sampling method can identify much lower energy states and explore a larger region phase space than the commonly used replica sampling (RS) Monte Carlo method. By clustering the simulated structures obtained in the PHS implementation of the SICHO model, we can successfully predict, among a representative benchmark 65 proteins set, 50 cases in which one of the top 5 clusters have a root-mean-square deviation (RMSD) from the native structure below 6.5 A. Compared with our previous calculations that used RS as the conformational search procedure, the number of successful predictions increased by four and the CPU cost is reduced. By comparing the structure clusters produced by both PHS and RS, we find a strong correlation between the quality of predicted structures and the minimum relative RMSD (mrRMSD) of structures clusters identified by using different search engines. This mrRMSD correlation may be useful in blind prediction as an indicator of the likelihood of successful folds.

Models, Molecular↗

Ab initio protein structure prediction on a genomic scale: application to the Mycoplasma genitalium genome.

An ab initio protein structure prediction procedure, TOUCHSTONE, was applied to all 85 small proteins of the Mycoplasma genitalium genome. TOUCHSTONE is based on a Monte Carlo refinement of a lattice model of proteins, which uses threading-based tertiary restraints. Such restraints are derived by extracting consensus contacts and local secondary structure from at least weakly scoring structures that, in some cases, can lack any global similarity to the sequence of interest. Selection of the native fold was done by using the convergence of the simulation from two different conformational search schemes and the lowest energy structure by a knowledge-based atomic-detailed potential. Among the 85 proteins, for 34 proteins with significant threading hits, the template structures were reasonably well reproduced. Of the remaining 51 proteins, 29 proteins converged to five or fewer clusters. In the test set, 84.8% of the proteins that converged to five or fewer clusters had a correct fold among the clusters. If this statistic is simply applied, 24 proteins (84.8% of the 29 proteins) may have correct folds. Thus, the topology of a total of 58 proteins probably has been correctly predicted. Based on these results, ab initio protein structure prediction is becoming a practical approach.

Algorithms↗

Docking of small ligands to low-resolution and theoretically predicted receptor structures.

We have developed a simple docking procedure that is able to utilize low-resolution models of proteins created by structure prediction algorithms such as threading or ab initio folding to predict the conformation of receptor-small ligand complexes. In our approach, using only approximate, discretized models of both molecules, we search for the steric and quasi-chemical complementarity between a ligand and the receptor molecules. This averaging procedure allows for the compensation of numerous structural inaccuracies resulting from the theoretical predictions of the receptor structure. The best relative orientation of these two models is obtained by an exhaustive scan over the rigid body's six-dimensional translational and rotational degrees of freedom. The search method is based on a real space grid-searching algorithm, unlike docking methods based on the fast Fourier Transform algorithm. We have applied this algorithm to rebuild structures of several complexes available in the Protein Data Bank. The structures of the receptors are produced by means of our threading algorithm PROSPECTOR, subsequently refined, and then utilized in the docking experiment. In many cases, not only is the localization of the binding site on the receptor surface correctly identified, but the proper orientation of the bounded ligand is also reasonably well reproduced within the level of accuracy of the modeled receptor itself.

Algorithms↗

Numerical study of the entropy loss of dimerization and the folding thermodynamics of the GCN4 leucine zipper.

A lattice-based model of a protein and the Monte Carlo simulation method are used to calculate the entropy loss of dimerization of the GCN4 leucine zipper. In the representation used, a protein is a sequence of interaction centers arranged on a cubic lattice, with effective interaction potentials that are both of physical and statistical nature. The Monte Carlo simulation method is then used to sample the partition functions of both the monomer and dimer forms as a function of temperature. A method is described to estimate the entropy loss upon dimerization, a quantity that enters the free energy difference between monomer and dimer, and the corresponding dimerization reaction constant. As expected, but contrary to previous numerical studies, we find that the entropy loss of dimerization is a strong function of energy (or temperature), except in the limit of large energies in which the motion of the two dimer chains becomes largely uncorrelated. At the monomer-dimer transition temperature we find that the entropy loss of dimerization is approximately five times smaller than the value that would result from ideal gas statistics, a result that is qualitatively consistent with a recent experimental determination of the entropy loss of dimerization of a synthetic peptide that also forms a two-stranded alpha-helical coiled coil.

Biophysical Phenomena↗

Computer simulations of protein folding with a small number of distance restraints.

A high coordination lattice model was used to represent the protein chain. Lattice points correspond to amino-acid side groups. A complicated force field was designed in order to reproduce a protein-like behavior of the chain. Long-distance tertiary restraints were also introduced into the model. The Replica Exchange Monte Carlo method was applied to find the lowest energy states of the folded chain and to solve the problem of multiple minima. In this method, a set of replicas of the model chain was simulated independently in different temperatures with the exchanges of replicas allowed. The model chains, which consisted of up to 100 residues, were folded to structures whose root-mean-square deviation (RMSD) from their native state was between 2.5 and 5 A. Introduction of restrain based on the positions of the backbone hydrogen atoms led to an improvement in the number of successful simulation runs. A small improvement (about 0.5 A) was also achieved in the RMSD of the folds. The proposed method can be used for the refinement of structures determined experimentally from NMR data.

Algorithms↗