Search PubMed⌕ Search

Biomedical subjects

J Skolnick

Publications and source records attributed to J Skolnick.

At least 19 recordsLinked to original sources

Defrosting the frozen approximation: PROSPECTOR--a new approach to threading.

PROSPECTOR (PROtein Structure Predictor Employing Combined Threading to Optimize Results) is a new threading approach that uses sequence profiles to generate an initial probe-template alignment and then uses this "partly thawed" alignment in the evaluation of pair interactions. Two types of sequence profiles are used: the close set, composed of sequences in which sequence identity lies between 35% and 90%; and the distant set, composed of sequences with a FASTA E-score less than 10. Thus, a total of four scoring functions are used in a hierarchical method: the close (distant) sequence profiles screen a structural database to provide an initial alignment of the probe sequence in each of the templates. The same database is then screened with a scoring function composed of sequence plus secondary structure plus pair interaction profiles. This combined hierarchical threading method is called PROSPECTOR1. For the original Fischer database, 59 of 68 pairs are correctly identified in the top position. Next, the set of the top 20 scoring sequences (four scoring functions times the top five structures) is used to construct a protein-specific pair potential based on consensus side-chain contacts occurring in 25% of the structures. In subsequent threading iterations, this protein-specific pair potential, when combined in a composite manner, is found to be more sensitive in identifying the correct pairs than when the original statistical potential is used, and it increases the number of recognized structures for the combined scoring functions, termed PROSPECTOR2, to a total of 61 Fischer pairs identified in the top position. Application to a second, smaller Fischer database of 27 probe-template pairs places 18 (17) structures in the top position for PROSPECTOR1 (PROSPECTOR2). Overall, these studies show that the use of pair interactions as assessed by the improved Z-score enhances the specificity of probe-template matches. Thus, when the hierarchy of scoring functions is combined, the ability to identify correct probe-template pairs is significantly enhanced. Finally, a web server has been established for use by the academic community (http://bioinformatics.danforthcenter.org/services/threading.html).

Benchmarking↗

Accurate reconstruction of all-atom protein representations from side-chain-based low-resolution models.

A procedure for the reconstruction of all-atom protein structures from side-chain center-based low-resolution models is introduced and applied to a set of test proteins with high-resolution X-ray structures. The accuracy of the rebuilt all-atom models is measured by root mean square deviations to the corresponding X-ray structures and percentages of correct chi(1) and chi(2) side-chain dihedrals. The benefit of including C(alpha) positions in the low-resolution model is examined, and the effect of lattice-based models on the reconstruction accuracy is discussed. Programs and scripts implementing the reconstruction procedure are made available through the NIH research resource for Multiscale Modeling Tools in Structural Biology (http://mmtsb.scripps.edu).

Models, Molecular↗

Derivation of protein-specific pair potentials based on weak sequence fragment similarity.

A method is presented for the derivation of knowledge-based pair potentials that corrects for the various compositions of different proteins. The resulting statistical pair potential is more specific than that derived from previous approaches as assessed by gapless threading results. Additionally, a methodology is presented that interpolates between statistical potentials when no homologous examples to the protein of interest are in the structural database used to derive the potential, to a Go-like potential (in which native interactions are favorable and all nonnative interactions are not) when homologous proteins are present. For cases in which no protein exceeds 30% sequence identity, pairs of weakly homologous interacting fragments are employed to enhance the specificity of the potential. In gapless threading, the mean z score increases from -10.4 for the best statistical pair potential to -12.8 when the local sequence similarity, fragment-based pair potentials are used. Examination of the ab initio structure prediction of four representative globular proteins consistently reveals a qualitative improvement in the yield of structures in the 4 to 6 A rmsd from native range when the fragment-based pair potential is used relative to that when the quasichemical pair potential is employed. This suggests that such protein-specific potentials provide a significant advantage relative to generic quasichemical potentials.

Amino Acids↗

Computer simulations of the properties of the alpha2, alpha2C, and alpha2D de novo designed helical proteins.

Reduced lattice models of the three de novo designed helical proteins alpha2, alpha2C, and alpha2D were studied. Low temperature stable folds were obtained for all three proteins. In all cases, the lowest energy folds were four-helix bundles. The folding pathway is qualitatively the same for all proteins studied. The energies of various topologies are similar, especially for the alpha2 polypeptide. The simulated crossover from molten globule to native-like behavior is very similar to that seen in experimental studies. Simulations on a reduced protein model reproduce most of the experimental properties of the alpha2, alpha2C, and alpha2D proteins. Stable four-helix bundle structures were obtained, with increasing native-like behavior on-going from alpha2 to alpha2D that mimics experiment.

Amino Acid Sequence↗

Sequence evolution and the mechanism of protein folding.

The impact on protein evolution of the physical laws that govern folding remains obscure. Here, by analyzing in silico-evolved sequences subjected to evolutionary pressure for fast folding, it is shown that: First, a subset of residues in the thermodynamic folding nucleus is mainly responsible for modulating the protein folding rate. Second and most important, the protein topology itself is of paramount importance in determining the location of these residues in the structure. Further stabilization of the interactions in this nucleus leads to fast folding sequences. Third, these nucleation points restrict the sequence space available to the protein during evolution. Correlated mutations between positions around these hot spots arise in a statistically significant manner, and most involve contacting residues. When a similar analysis is carried out on real proteins, qualitatively similar results are obtained.

Biophysical Phenomena↗

From genes to protein structure and function: novel applications of computational approaches in the genomic era.

The genome-sequencing projects are providing a detailed 'parts list' of life. A key to comprehending this list is understanding the function of each gene and each protein at various levels. Sequence-based methods for function prediction are inadequate because of the multifunctional nature of proteins. However, just knowing the structure of the protein is also insufficient for prediction of multiple functional sites. Structural descriptors for protein functional sites are crucial for unlocking the secrets in both the sequence and structural-genomics projects.

Genes↗

Structural genomics and its importance for gene function analysis.

Structural genomics projects aim to solve the experimental structures of all possible protein folds. Such projects entail a conceptual shift from traditional structural biology in which structural information is obtained on known proteins to one in which the structure of a protein is determined first and the function assigned only later. Whereas the goal of converting protein structure into function can be accomplished by traditional sequence motif-based approaches, recent studies have shown that assignment of a protein's biochemical function can also be achieved by scanning its structure for a match to the geometry and chemical identity of a known active site. Importantly, this approach can use low-resolution structures provided by contemporary structure prediction methods. When applied to genomes, structural information (either experimental or predicted) is likely to play an important role in high-throughput function assignment.

Animals↗

A method for the improvement of threading-based protein models.

A new method for the homology-based modeling of protein three-dimensional structures is proposed and evaluated. The alignment of a query sequence to a structural template produced by threading algorithms usually produces low-resolution molecular models. The proposed method attempts to improve these models. In the first stage, a high-coordination lattice approximation of the query protein fold is built by suitable tracking of the incomplete alignment of the structural template and connection of the alignment gaps. These initial lattice folds are very similar to the structures resulting from standard molecular modeling protocols. Then, a Monte Carlo simulated annealing procedure is used to refine the initial structure. The process is controlled by the model's internal force field and a set of loosely defined restraints that keep the lattice chain in the vicinity of the template conformation. The internal force field consists of several knowledge-based statistical potentials that are enhanced by a proper analysis of multiple sequence alignments. The template restraints are implemented such that the model chain can slide along the template structure or even ignore a substantial fraction of the initial alignment. The resulting lattice models are, in most cases, closer (sometimes much closer) to the target structure than the initial threading-based models. All atom models could easily be built from the lattice chains. The method is illustrated on 12 examples of target/template pairs whose initial threading alignments are of varying quality. Possible applications of the proposed method for use in protein function annotation are briefly discussed.

Amino Acid Sequence↗

Correlation between knowledge-based and detailed atomic potentials: application to the unfolding of the GCN4 leucine zipper.

The relationship between the unfolding pseudo free energies of reduced and detailed atomic models of the GCN4 leucine zipper is examined. Starting from the native crystal structure, a large number of conformations ranging from folded to unfolded were generated by all-atom molecular dynamics unfolding simulations in an aqueous environment at elevated temperatures. For the detailed atomic model, the pseudo free energies are obtained by combining the CHARMM all-atom potential with a solvation component from the generalized Born, surface accessibility, GB/SA, model. Reduced model energies were evaluated using a knowledge-based potential. Both energies are highly correlated. In addition, both show a good correlation with the root mean square deviation, RMSD, of the backbone from native. These results suggest that knowledge-based potentials are capable of describing at least some of the properties of the folded as well as the unfolded states of proteins, even though they are derived from a database of native protein structures. Since only conformations generated from an unfolding simulation are used, we cannot assess whether these potentials can discriminate the native conformation from the manifold of alternative, low-energy misfolded states. Nevertheless, these results also have significant implications for the development of a methodology for multiscale modeling of proteins that combines reduced and detailed atomic models.

DNA-Binding Proteins↗

Averaging interaction energies over homologs improves protein fold recognition in gapless threading.

Protein structure prediction is limited by the inaccuracy of the simplified energy functions necessary for efficient sorting over many conformations. It was recently suggested (Finkelstein, Phys Rev Lett 1998;80:4823-4825) that these errors can be reduced by energy averaging over a set of homologous sequences. This conclusion is confirmed in this study by testing protein structure recognition in gapless threading. The accuracy of recognition was estimated by the Z-score values obtained in gapless threading tests. For threading, we used 20 target proteins, each having from 20 to 70 homologs taken from the HSSP sequence base. The energy of the native structures was compared with the energy from 34 to 75 thousand of alternative structures generated by threading. The energy calculations were done with our recently developed Calpha atom-based phenomenological potentials. We show that averaging of protein energies over homologs reduces the Z-score from approximately -6.1 (average Z-score for individual chains) to approximately -8.1. This means that a correct fold can be found among 3 x 10(9) random folds in the first case and among 3 x 10(15) in the second. Such increase in selectivity is important for recognition of protein folds.

Cytochrome c Group↗

Ab initio folding of proteins using restraints derived from evolutionary information.

We present our predictions in the ab initio structure prediction category of CASP3. Eleven targets were folded, using a method based on a Monte Carlo search driven by secondary and tertiary restraints derived from multiple sequence alignments. Our results can be qualitatively summarized as follows: The global fold can be considered "correct" for targets 65 and 74, "almost correct" for targets 64, 75, and 77, "half-correct" for target 79, and "wrong" for targets 52, 56, 59, and 63. Target 72 has not yet been solved experimentally. On average, for small helical and alpha/beta proteins (on the order of 110 residues or smaller), the method predicted low resolution structures with a reasonably good prediction of the global topology. Most encouraging is that in some situations, such as with target 75 and, particularly, target 77, the method can predict a substantial portion of a rare or even a novel fold. However, the current method still fails on some beta proteins, proteins over the 110-residue threshold, and sequences in which only a poor multiple sequence alignment can be built. On the other hand, for small proteins, the method gives results of quality at least similar to that of threading, with the advantage of not being restricted to known folds in the protein database. Overall, these results indicate that some progress has been made on the ab initio protein folding problem. Detailed information about our results can be obtained by connecting to http:/(/)www.bioinformatics.danforthcenter.org/+ ++CASP3.

Algorithms↗

De novo simulations of the folding thermodynamics of the GCN4 leucine zipper.

Entropy Sampling Monte Carlo (ESMC) simulations were carried out to study the thermodynamics of the folding transition in the GCN4 leucine zipper (GCN4-lz) in the context of a reduced model. Using the calculated partition functions for the monomer and dimer, and taking into account the equilibrium between the monomer and dimer, the average helix content of the GCN4-lz was computed over a range of temperatures and chain concentrations. The predicted helix contents for the native and denatured states of GCN4-lz agree with the experimental values. Similar to experimental results, our helix content versus temperature curves show a small linear decline in helix content with an increase in temperature in the native region. This is followed by a sharp transition to the denatured state. van't Hoff analysis of the helix content versus temperature curves indicates that the folding transition can be described using a two-state model. This indicates that knowledge-based potentials can be used to describe the properties of the folded and unfolded states of proteins.

Computer Simulation↗

Dynamics and thermodynamics of beta-hairpin assembly: insights from various simulation techniques.

Small peptides that might have some features of globular proteins can provide important insights into the protein folding problem. Two simulation methods, Monte Carlo Dynamics (MCD), based on the Metropolis sampling scheme, and Entropy Sampling Monte Carlo (ESMC), were applied in a study of a high-resolution lattice model of the C-terminal fragment of the B1 domain of protein G. The results provide a detailed description of folding dynamics and thermodynamics and agree with recent experimental findings (. Nature. 390:196-197). In particular, it was found that the folding is cooperative and has features of an all-or-none transition. Hairpin assembly is usually initiated by turn formation; however, hydrophobic collapse, followed by the system rearrangement, was also observed. The denatured state exhibits a substantial amount of fluctuating helical conformations, despite the strong beta-type secondary structure propensities encoded in the sequence.

Amino Acid Sequence↗

Structure-based functional motif identifies a potential disulfide oxidoreductase active site in the serine/threonine protein phosphatase-1 subfamily.

In previous work, 3-dimensional descriptors of protein function ('fuzzy functional forms') were used to identify disulfide oxidoreductase active sites in high-resolution protein structures. During this analysis, a potential disulfide oxidoreductase active site in the serine/threonine protein phosphatase-1 (PP1) crystal structure was discovered. In PP1, the potential redox active site is located in close proximity to the phosphatase active site. This result is interesting in view of literature suggesting that serine/threonine phosphatases could be subject to redox control mechanisms within the cell; however, the actual source of this control is unknown. Additional analysis presented here shows that the putative oxidoreductase active site is highly conserved in the serine/threonine phosphatase-1 subfamily, but not in the serine/threonine phosphatase-2A or -2B subfamilies. These results demonstrate the significant advantages of using structure-based motifs for protein functional site identification. First, a putative disulfide oxidoreductase active site has been identified in serine-threonine phosphatases using a descriptor built from the glutaredoxin/thioredoxin family, proteins that have no apparent evolutionary relationship whatsoever to the PP1 proteins. Second, the proximity of the putative disulfide oxidoreductase active site to the phosphatase active site provides evidence toward a regulatory control mechanism. No sequence-based method could provide either piece of information.

Amino Acid Sequence↗

Family accommodation of obsessive-compulsive symptoms: instrument development and assessment of family behavior.

Relatives frequently accommodate patients' obsessive-compulsive symptoms and clinicians hypothesize that such accommodations adversely affect patient outcome. This study's purpose was to develop a valid and reliable measure, the Family Accommodation Scale for Obsessive-Compulsive Disorder (FAS), and to investigate the family accommodation construct. We administered the FAS and additional family and patient measures to 36 adult obsessive-compulsive patients and their primary caregivers. The FAS demonstrated excellent interrater reliability and good internal consistency and performed well on assessment of its convergent and discriminant validity. Family accommodation was significantly associated with patient symptom severity and functioning, and with relatives' own obsessive-compulsive symptoms. Although most relatives accommodated patient symptoms, many did not believe that such accommodations improved the patient's clinical status. The FAS will provide researchers and clinicians with a useful tool for assessing family accommodation and for identifying families who may benefit from interventions aimed at developing more adaptive coping strategies.

Adaptation, Psychological↗

From fold predictions to function predictions: automation of functional site conservation analysis for functional genome predictions.

A database of functional sites for proteins with known structures, SITE, is constructed and used in conjunction with a simple pattern matching program SiteMatch to evaluate possible function conservation in a recently constructed database of fold predictions for Escherichia coli proteins (Rychlewski L et al., 1999, Protein Sci 8:614-624). In this and other prediction databases, fold predictions are based on algorithms that can recognize weak sequence similarities and putatively assign new proteins into already characterized protein families. It is not clear whether such sequence similarities arise from distant homologies or general similarity of physicochemical features along the sequence. Leaving aside the important question of nature of relations within fold superfamilies, it is possible to assess possible function conservation by looking at the pattern of conservation of crucial functional residues. SITE consists of a multilevel function description based on structure annotations and structure analyses. In particular, active site residues, ligand binding residues, and patterns of hydrophobic residues on the protein surface are used to describe different functional features. SiteMatch, a simple pattern matching program, is designed to check the conservation of residues involved in protein activity in alignments generated by any alignment method. Here, this procedure is used to study conservation of functional features in alignments between protein sequences from the E. coli genome and their optimal structural templates. The optimal templates were identified and alignments taken from the database of genomic structural predictions was described in a previous publication (Rychlewski L et al., 1999, Protein Sci 8:614-624). An automated assessment of function conservation is used to analyze the relation between fold and function similarity for a large number of fold predictions. For instance, it is shown that identifying low significance predictions with a high level of functional residue conservations can be used to extend the prediction sensitivity for fold prediction methods. Over 100 new fold/function predictions in this class were obtained in the E. coli genome. At the same time, about 30% of our previous fold predictions are not confirmed as function predictions, further highlighting the problem of function divergence in fold superfamilies.

Algorithms↗