Search PubMed⌕ Search

Biomedical subjects

K T Simons

Publications and source records attributed to K T Simons.

10 recordsLinked to original sources

Prospects for ab initio protein structural genomics.

We present the results of a large-scale testing of the ROSETTA method for ab initio protein structure prediction. Models were generated for two independently generated lists of small proteins (up to 150 amino acid residues), and the results were evaluated using traditional rmsd based measures and a novel measure based on the structure-based comparison of the models to the structures in the PDB using DALI. For 111 of 136 all alpha and alpha/beta proteins 50 to 150 residues in length, the method produced at least one model within 7 A rmsd of the native structure in 1000 attempts. For 60 of these proteins, the closest structure match in the PDB to at least one of the ten most frequently generated conformations was found to be structurally related (four standard deviations above background) to the native protein. These results suggest that ab initio structure prediction approaches may soon be useful for generating low resolution models and identifying distantly related proteins with similar structures and perhaps functions for these classes of proteins on the genome scale.

Computer Simulation↗

Topology, stability, sequence, and length: defining the determinants of two-state protein folding kinetics.

The fastest simple, single domain proteins fold a million times more rapidly than the slowest. Ultimately this broad kinetic spectrum is determined by the amino acid sequences that define these proteins, suggesting that the mechanisms that underlie folding may be almost as complex as the sequences that encode them. Here, however, we summarize recent experimental results which suggest that (1) despite a vast diversity of structures and functions, there are fundamental similarities in the folding mechanisms of single domain proteins and (2) rather than being highly sensitive to the finest details of sequence, their folding kinetics are determined primarily by the large-scale, redundant features of sequence that determine a protein's gross structural properties. That folding kinetics can be predicted using simple, empirical, structure-based rules suggests that the fundamental physics underlying folding may be quite straightforward and that a general and quantitative theory of protein folding rates and mechanisms (as opposed to unfolding rates and thus protein stability) may be near on the horizon.

Amino Acid Sequence↗

Improved recognition of native-like protein structures using a combination of sequence-dependent and sequence-independent features of proteins.

We describe the development of a scoring function based on the decomposition P(structure/sequence) proportional to P(sequence/structure) *P(structure), which outperforms previous scoring functions in correctly identifying native-like protein structures in large ensembles of compact decoys. The first term captures sequence-dependent features of protein structures, such as the burial of hydrophobic residues in the core, the second term, universal sequence-independent features, such as the assembly of beta-strands into beta-sheets. The efficacies of a wide variety of sequence-dependent and sequence-independent features of protein structures for recognizing native-like structures were systematically evaluated using ensembles of approximately 30,000 compact conformations with fixed secondary structure for each of 17 small protein domains. The best results were obtained using a core scoring function with P(sequence/structure) parameterized similarly to our previous work (Simons et al., J Mol Biol 1997;268:209-225] and P(structure) focused on secondary structure packing preferences; while several additional features had some discriminatory power on their own, they did not provide any additional discriminatory power when combined with the core scoring function. Our results, on both the training set and the independent decoy set of Park and Levitt (J Mol Biol 1996;258:367-392), suggest that this scoring function should contribute to the prediction of tertiary structure from knowledge of sequence and secondary structure.

Models, Statistical↗

Ab initio protein structure prediction of CASP III targets using ROSETTA.

To generate structures consistent with both the local and nonlocal interactions responsible for protein stability, 3 and 9 residue fragments of known structures with local sequences similar to the target sequence were assembled into complete tertiary structures using a Monte Carlo simulated annealing procedure (Simons et al., J Mol Biol 1997; 268:209-225). The scoring function used in the simulated annealing procedure consists of sequence-dependent terms representing hydrophobic burial and specific pair interactions such as electrostatics and disulfide bonding and sequence-independent terms representing hard sphere packing, alpha-helix and beta-strand packing, and the collection of beta-strands in beta-sheets (Simons et al., Proteins 1999;34:82-95). For each of 21 small, ab initio targets, 1,200 final structures were constructed, each the result of 100,000 attempted fragment substitutions. The five structures submitted for the CASP III experiment were chosen from the approximately 25 structures with the lowest scores in the broadest minima (assessed through the number of structural neighbors; Shortle et al., Proc Natl Acad Sci USA 1998;95:1158-1162). The results were encouraging: highlights of the predictions include a 99-residue segment for MarA with an rmsd of 6.4 A to the native structure, a 95-residue (full length) prediction for the EH2 domain of EPS15 with an rmsd of 6.0 A, a 75-residue segment of DNAB helicase with an rmsd of 4.7 A, and a 67-residue segment of ribosomal protein L30 with an rmsd of 3.8 A. These results suggest that ab initio methods may soon become useful for low-resolution structure prediction for proteins that lack a close homologue of known structure.

Algorithms↗

Robustness of protein folding kinetics to surface hydrophobic substitutions.

We use both combinatorial and site-directed mutagenesis to explore the consequences of surface hydrophobic substitutions for the folding of two small single domain proteins, the src SH3 domain, and the IgG binding domain of Peptostreptococcal protein L. We find that in almost every case, destabilizing surface hydrophobic substitutions have much larger effects on the rate of unfolding than on the rate of folding, suggesting that nonnative hydrophobic interactions do not significantly interfere with the rate of core assembly.

Amino Acid Substitution↗

Clustering of low-energy conformations near the native structures of small proteins.

Recent experimental studies of the denatured state and theoretical analyses of the folding landscape suggest that there are a large multiplicity of low-energy, partially folded conformations near the native state. In this report, we describe a strategy for predicting protein structure based on the working hypothesis that there are a greater number of low-energy conformations surrounding the correct fold than there are surrounding low-energy incorrect folds. To test this idea, 12 ensembles of 500 to 1,000 low-energy structures for 10 small proteins were analyzed by calculating the rms deviation of the Calpha coordinates between each conformation and every other conformation in the ensemble. In all 12 cases, the conformation with the greatest number of conformations within 4-A rms deviation was closer to the native structure than were the majority of conformations in the ensemble, and in most cases it was among the closest 1 to 5%. These results suggest that, to fold efficiently and retain robustness to changes in amino acid sequence, proteins may have evolved a native structure situated within a broad basin of low-energy conformations, a feature which could facilitate the prediction of protein structure at low resolution.

Computer Simulation↗

Contact order, transition state placement and the refolding rates of single domain proteins.

Theoretical studies have suggested relationships between the size, stability and topology of a protein fold and the rate and mechanisms by which it is achieved. The recent characterization of the refolding of a number of simple, single domain proteins has provided a means of testing these assertions. Our investigations have revealed statistically significant correlations between the average sequence separation between contacting residues in the native state and the rate and transition state placement of folding for a non-homologous set of simple, single domain proteins. These indicate that proteins featuring primarily sequence-local contacts tend to fold more rapidly and exhibit less compact folding transition states than those characterized by more non-local interactions. No significant relationship is apparent between protein length and folding rates, but a weak correlation is observed between length and the fraction of solvent-exposed surface area buried in the transition state. Anticipated strong relationships between equilibrium folding free energy and folding kinetics, or between chemical denaturant and temperature dependence-derived measures of transition state placement, are not apparent. The observed correlations are consistent with a model of protein folding in which the size and stability of the polypeptide segments organized in the transition state are largely independent of protein length, but are related to the topological complexity of the native state. The correlation between topological complexity and folding rates may reflect chain entropy contributions to the folding barrier.

Animals↗

Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions.

We explore the ability of a simple simulated annealing procedure to assemble native-like structures from fragments of unrelated protein structures with similar local sequences using Bayesian scoring functions. Environment and residue pair specific contributions to the scoring functions appear as the first two terms in a series expansion for the residue probability distributions in the protein database; the decoupling of the distance and environment dependencies of the distributions resolves the major problems with current database-derived scoring functions noted by Thomas and Dill. The simulated annealing procedure rapidly and frequently generates native-like structures for small helical proteins and better than random structures for small beta sheet containing proteins. Most of the simulated structures have native-like solvent accessibility and secondary structure patterns, and thus ensembles of these structures provide a particularly challenging set of decoys for evaluating scoring functions. We investigate the effects of multiple sequence information and different types of conformational constraints on the overall performance of the method, and the ability of a variety of recently developed scoring functions to recognize the native-like conformations in the ensembles of simulated structures.

Bayes Theorem↗

Characterization of the free energy spectrum of peptostreptococcal protein L.

BACKGROUND: Native state hydrogen/deuterium exchange studies on cytochrome c and RNase H revealed the presence of excited states with partially formed native structure. We set out to determine whether such excited states are populated for a very small and simple protein, the IgG-binding domain of peptostreptococcal protein L. RESULTS: Hydrogen/deuterium exchange data on protein L in 0-1.2 M guanidine fit well to a simple model in which the only contributions to exchange are denaturant-independent local fluctuations and global unfolding. A substantial discrepancy emerged between unfolding free energy estimates from hydrogen/deuterium exchange and linear extrapolation of earlier guanidine denaturation experiments. A better determined estimate of the free energy of unfolding obtained by global analysis of a series of thermal denaturation experiments in the presence of 0-3 M guanidine was in good agreement with the estimate from hydrogen/deuterium exchange. CONCLUSIONS: For protein L under native conditions, there do not appear to be partially folded states with free energies intermediate between that of the folded and unfolded states. The linear extrapolation method significantly underestimates the free energy of folding of protein L due to deviations from linearity in the dependence of the free energy on the denaturant concentration.

Bacterial Proteins↗

Local sequence-structure correlations in proteins.

Considerable progress has been made in understanding the relationship between local amino acid sequence and local protein structure. Recent highlights include numerous studies of the structures adopted by short peptides, new approaches to correlating sequence patterns with structure patterns, and folding simulations using simple potentials.

Amino Acid Sequence↗