Search PubMedSearch

Biomedical subjects

J Moult

Publications and source records attributed to J Moult.

At least 19 recordsLinked to original sources

Processing and analysis of CASP3 protein structure predictions.

Livermore Prediction Center provides basic infrastructure for the CASP (Critical Assessment of Structure Prediction) experiments, including prediction processing and verification servers, a system of prediction evaluation tools, and interactive numerical and graphical displays. Here we outline the essentials of our approach, with discussion of the superposition procedures, definitions of basic measures, and descriptions of new methods developed to analyze predictions. Our primary focus is on the evaluation of three-dimensional models and secondary structure predictions. To put the results of the three prediction experiments held to date on the same footing, the latest CASP3 evaluation criteria were retrospectively applied to both CASP1 and CASP2 predictions. Finally, we give an overview of our website (http:/(/)PredictionCenter.llnl.gov), which makes the target structures, predictions, and the evaluation system accessible to the community.

Amino Acid Sequence

Some measures of comparative performance in the three CASPs.

Performance in the three Critical Assessment of protein Structure Prediction (CASP) experiments has been compared in the areas of alignment accuracy for models based on homology and three-dimensional accuracy for models produced by using ab initio prediction methods. The homologous models span the comparative modeling and fold-recognition regimes. Each CASP target is assigned a relative difficulty based on the extent of sequence identity and the degree of structural overlap with the best available template. There is a clear improvement in alignment accuracy between CASP1 and CASPs 2 and 3 over much of the difficulty scale but no apparent improvement between CASP2 and CASP3. Encouragingly, the best ab initio models of small targets are clearly more accurate in CASP3 than in CASPs 1 and 2.

Algorithms

Predicting protein three-dimensional structure.

The current state of the art in modeling protein structure has been assessed, based on the results of the CASP (Critical Assessment of protein Structure Prediction) experiments. In comparative modeling, improvements have been made in sequence alignment, sidechain orientation and loop building. Refinement of the models remains a serious challenge. Improved sequence profile methods have had a large impact in fold recognition. Although there has been some progress in alignment quality, this factor still limits model usefulness. In ab initio structure prediction, there has been notable progress in building approximately correct structures of 40-60 residue-long protein fragments. There is still a long way to go before the general ab initio prediction problem is solved. Overall, the field is maturing into a practical technology, able to deliver useful models for a large number of sequences.

Models, Molecular

Local electrostatic optimization in proteins.

A simple electrostatic model has been used to investigate the extent to which the structure of protein molecules is organized to optimize the internal electrostatic interactions. We find that the model provides a favorable total intra-protein electrostatic energy for almost all polar and charged groups of atoms, suggesting a high degree of structural optimization. By contrast, a significant fraction of individual group-group interactions are found to be unfavorable. An analysis as a function of the range of interactions included shows the electrostatic organization is generally relatively short range (up to 6 or 7 A between group centers). Although the model is very simple, it is useful for assessing the overall quality of protein experimental structures, for pin-pointing some types of errors and as a guide to improving protein design.

Crystallography, X-Ray

A graph-theoretic algorithm for comparative modeling of protein structure.

The interconnected nature of interactions in protein structures appears to be the major hurdle in preventing the construction of accurate comparative models. We present an algorithm that uses graph theory to handle this problem. Each possible conformation of a residue in an amino acid sequence is represented using the notion of a node in a graph. Each node is given a weight based on the degree of the interaction between its side-chain atoms and the local main-chain atoms. Edges are then drawn between pairs of residue conformations/nodes that are consistent with each other (i.e. clash-free and satisfying geometrical constraints). The edges are weighted based on the interactions between the atoms of the two nodes. Once the entire graph is constructed, all the maximal sets of completely connected nodes (cliques) are found using a clique-finding algorithm. The cliques with the best weights represent the optimal combinations of the various main-chain and side-chain possibilities, taking the respective environments into account. The algorithm is used in a comparative modeling scenario to build side-chains, regions of main chain, and mix and match between different homologs in a context-sensitive manner. The predictive power of this method is assessed by applying it to cases where the experimental structure is not known in advance.

Algorithms

Conformation of the sebacyl beta1Lys82-beta2Lys82 crosslink in T-state human hemoglobin.

The crystal structure of human T state hemoglobin crosslinked with bis(3,5-dibromo-salicyl) sebacate has been determined at 1.9 A resolution. The final crystallographic R factor is 0.168 with root-mean-square deviations (RMSD) from ideal bond distance of 0.018 A. The 10-carbon sebacyl residue found in the beta cleft covalently links the two betaLys82 residues. The sebacyl residue assumes a zigzag conformation with cis amide bonds formed by the NZ atoms of betaLys82's and the sebacyl carbonyl oxygens. The atoms of the crosslink have an occupancy factor of 1.0 with an average temperature factor for all atoms of 34 A2. An RMSD of 0.27 for all CA's of the tetramer is observed when the crosslinked deoxyhemoglobin is compared with deoxyhemoglobin refined by using a similar protocol, 2HHD [Fronticelli et al. J. Biol. Chem. 269: 23965-23969, 1994]. Thus, no significant perturbations in the tertiary or quaternary structure are introduced by the presence of the sebacyl residue. However, the sebacyl residue does displace seven water molecules in the beta cleft and the conformations of the beta1Lys82 and beta2Lys82 are altered because of the crosslinking. The carbonyl oxygen that is part of the amide bond formed with the NZ of beta2Lys82 forms a hydrogen bond with side chain of beta2Asn139 that is in turn hydrogen-bonded to the side chain of beta2Arg104. A comparison of the observed conformation with that modeled [Bucci et al. Biochemistry 35:3418-3425, 1996] shows significant differences. The differences in the structures can be rationalized in terms of compensating changes in the estimated free-energy balance, based on differences in exposed surface areas and the observed shift in the side-chain hydrogen-bonding pattern involving beta2Arg104, beta2Asn139, and the associated sebacyl carbonyl group.

Cross-Linking Reagents

An all-atom distance-dependent conditional probability discriminatory function for protein structure prediction.

We present a formalism to compute the probability of an amino acid sequence conformation being native-like, given a set of pairwise atom-atom distances. The formalism is used to derive three discriminatory functions with different types of representations for the atom-atom contacts observed in a database of protein structures. These functions include two virtual atom representations and one all-heavy atom representation. When applied to six different decoy sets containing a range of correct and incorrect conformations of amino acid sequences, the all-atom distance-dependent discriminatory function is able to identify correct from incorrect more often than the discriminatory functions using approximate representations. We illustrate the importance of using a detailed atomic description for obtaining the most accurate discrimination, and the necessity for testing discriminatory functions against a wide variety of decoys. The discriminatory function is also shown to be capable of capturing the fine details of atom-atom preferences. These results suggest that the all-atom distance-dependent discriminatory function will be useful for protein structure prediction and model refinement.

Models, Theoretical

Determinants of side chain conformational preferences in protein structures.

A discriminatory function based on a statistical analysis of atomic contacts in protein structures is used for selecting side chain rotamers given a peptide main chain. The function allows us to rank different possible side chain conformations on the basis of contacts between side chain atoms and atoms in the environment. We compare the differences in constructing side chain conformations using contacts with only the local main chain, using the entire main chain, and by building pairs of side chains simultaneously with local main chain information. Using only the local main chain allows us to construct side chains with approximately 75% of the chi1 angles within 30 degrees of the experimental value, and an average side chain atom r.m.s.d. of 1.72 A in a set of 10 proteins. The results of constructing side chains for the 10 proteins are compared with the results of other side chain building methods previously published. The comparison shows similar accuracies. An advantage of the present method is that it can be used to select a small number of likely side chain conformations for each residue, thus permitting limited combinatorial searches for building multiple protein side chains simultaneously.

Amino Acid Sequence

Protein folding simulations with genetic algorithms and a detailed molecular description.

We have explored the application of genetic algorithms (GA) to the determination of protein structure from sequence, using a full atom representation. A free energy function with point charge electrostatics and an area based solvation model is used. The method is found to be superior to previously investigated Monte Carlo algorithms. For selected fragments, up to 14 residues long, the lowest free energy structures produced by the GA are similar in conformation to the corresponding experimental structures in most cases. There are three main conclusions from these results. First, the genetic algorithm is an effective method for searching amongst the compact conformations of a polypeptide chain. Second, the free energy function is generally able to select native-like conformations. However, some deficiencies are identified, and further development is proposed. Third, the selection of native-like conformations for some protein fragments establishes that in these cases the conformation observed in the full protein structure is largely context independent. The implications for the nature of protein folding pathways are discussed.

Algorithms

Ab initio protein folding simulations with genetic algorithms: simulations on the complete sequence of small proteins.

Ab-initio folding simulations have been performed on three small proteins using a genetic algorithm- (GA-) based search method which operates on an all atom representation. Simulations were also performed on a number of small peptides expected to be independent folding units. The present genetic algorithm incorporates the results of developments made to the method first tested in CASP1. Additional operators have been introduced into the search in order to allow the simulation of longer sequences and to avoid the simulation of longer sequences and to avoid premature free energy convergence. Secondary structure information derived from a consensus of eight methods and Monte Carlo simulations on sets of homologous sequences has been used to bias the starting populations used in the GA simulations. For the fragment simulations, the results generally have approximately correct local structure, but tend to be too compact, leading to poor RMS error values. One of the three small protein structures has the topology and most of the general organization correct, although many of the details are incorrect.

Algorithms

Handling context-sensitivity in protein structures using graph theory: bona fide prediction.

We constructed five comparative models in a blind manner for the second meeting on the Critical Assessment of protein Structure Prediction methods (CASP2). The method used is based on a novel graph-theoretic clique-finding approach, and attempts to address the problem of interconnected structural changes in the comparative modeling of protein structures. We discuss briefly how the method is used for protein structure prediction, and detail how it performs in the blind tests. We find that compared to CASP1, significant improvements in building insertions and deletions and sidechain conformations have been achieved.

Amino Acid Sequence

Criteria for evaluating protein structures derived from comparative modeling.

Following the first experiment for the Critical Assessment of methods for protein Structure Prediction (CASP1), numerical criteria were devised to analyze the performance of prediction methods. We report here the criteria for comparative modeling, and how effective they were in CASP2. These criteria are intended to evolve into a set of numerical measures that provide a comprehensive assessment of the quality of a structure produced by comparative modeling, and provide a means of investigating which modeling methods are most effective, so as to establish where future effort may be most productively applied.

Evaluation Studies as Topic

Comparison of database potentials and molecular mechanics force fields.

The advantages and disadvantages of database and molecular mechanics force fields for the study of macromolecules are compared, with emphasis on the ability to distinguish between correct and incorrect structures. Molecular mechanics force fields have the advantage of resting on a clear theoretical basis, permitting an in-depth analysis of different contributions. On the other hand, large simplifications are necessary for tractable computing, and there has so far been little effective testing at the macromolecular level. Database potentials allow greater freedom of functional form and have been shown to be effective at discriminating between correct and incorrect complete structures. The principal negative is a controversial relationship to free energy. More testing and comparison of both sorts of potential are needed.

Databases, Factual

Chaos in protein dynamics.

MD simulations, currently the most detailed description of the dynamic evolution of proteins, are based on the repeated solution of a set of differential equations implementing Newton's second law. Many such systems are known to exhibit chaotic behavior, i.e., very small changes in initial conditions are amplified exponentially and lead to vastly different, inherently unpredictable behavior. We have investigated the response of a protein fragment in an explicit solvent environment to very small perturbations of the atomic positions (10(-3)-10(-9) A). Independent of the starting conformation (native-like, compact, extended), perturbed dynamics trajectories deviated rapidly, leading to conformations that differ by approximately 1 A RMSD within 1-2 ps. Furthermore, introducing the perturbation more than 1-2 ps before a significant conformational transition leads to a loss of the transition in the perturbed trajectories. We present evidence that the observed chaotic behavior reflects physical properties of the system rather than numerical instabilities of the calculation and discuss the implications for models of protein folding and the use of MD as a tool to analyze protein folding pathways.

Bacterial Proteins

Elimination of the hydrolytic water molecule in a class A beta-lactamase mutant: crystal structure and kinetics.

Two site-directed mutant enzymes of the class A beta-lactamase from Staphylococcus aureus PC1 were produced with the goal of blocking the site that in the native enzyme is occupied by the proposed hydrolytic water molecule. The crystal structures of these two mutant enzymes, N170Q and N170M, have been determined and refined at 2.2 and 2.0 A, respectively. They reveal that the side chain of Gln 170 displaces the water molecule, whereas that of Met170 does not. In both cases, the catalytic rates with benzylpenicillin are reduced by 10(4) compared with the native enzyme. With nitrocefin, the N170Q mutant enzyme exhibits an approximately 800-fold reduced rate compared with the native enzyme and in addition, a fast initial burst with stoichiometry of 1 mol of degraded nitrocefin/mol of enzyme. Stopped-flow kinetic experiments establish that the rate constant of the burst is 250 s-1, a value comparable with the rate of acylation of the native enzyme. Two structurally based mechanisms that explain the kinetic properties of the N170Q beta-lactamase are proposed, both invoking a deacylation-impaired enzyme due to the elimination of the hydrolytic water molecule. The catalytic rate of the N170M mutant enzyme with nitrocefin is reduced by approximately 50-fold compared with the native enzyme, and the slow progressive inhibition that is revealed indicates that the hydrolysis proceeds via a branched pathway mechanism. This is consistent with the structural data that show that the water site is preserved and that Met170 occupies part of the space that is required for substrate binding. The short contacts between the substrate and the enzyme may lead to structure perturbation and inactivation.

Base Sequence

Local interactions dominate folding in a simple protein model.

Recent computational studies of simple models of protein folding have concluded that a pronounced energy minimum (i.e. large gap in energy between low-energy states of the model) is a necessary and sufficient condition to ensure folding of a sequence to its lowest-energy conformation. Here, we show that this conclusion strongly depends on the particular temperature scheme selected to govern the simulations. On the other hand, we show that there is a dominant factor determining if a sequence is foldable. That is, the strength of possible interactions between residues close in the sequence. We show that sequences with many possible strong local interactions (either favorable or, more surprisingly, a mixture of strong favorable and unfavorable ones) are easy to fold. Progressively increasing the strength of such local interactions makes sequences easier and easier to fold. These results support the idea that initial formation of local substructures is important to the foldability of real proteins.

Mathematical Computing