Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

CASP2 knowledge-based approach to distant homology recognition and fold prediction in CASP4.

In 1996, in CASP2, we presented a semimanual approach to the prediction of protein structure that was aimed at the recognition of probable distant homology, where it existed, between a given target protein and a protein of known structure (Murzin and Bateman, Proteins 1997; Suppl 1:105-112). Central to our method was the knowledge of all known structural and probable evolutionary relationships among proteins of known structure classified in the SCOP database (Murzin et al., J Mol Biol 1995;247:536-540). It was demonstrated that a knowledge-based approach could compete successfully with the best computational methods of the time in the correct recognition of the target protein fold. Four years later, in CASP4, we have applied essentially the same knowledge-based approach to distant homology recognition, concentrating our effort on the improvement of the completeness and alignment accuracy of our models. The manifold increase of available sequence and structure data was to our advantage, as well as was the experience and expertise obtained through the classification of these data. In particular, we were able to model most of our predictions from several distantly related structures rather than from a single parent structure, and we could use more superfamily characteristic features for the refinement of our alignments. Our predictions for each of the attempted distant homology recognition targets ranked among the few top predictions for each of these targets, with the predictions for the hypothetical protein HI0065 (T0104) and the C-terminal domain of the ABC transporter MalK (T0121C) being particularly successful. We also have attempted the prediction of protein folds of some of the targets tentatively assigned to new superfamilies. The average quality of our fold predictions was far less than the quality of our distant homology recognition models, but for the two targets, chorismate lyase (T0086) and Appr>p cyclic phosphodiesterase (T0094), our predictions achieved the top ranking.

ATP-Binding Cassette Transporters↗

Processing and evaluation of predictions in CASP4.

The Livermore Prediction Center conducted the target collection and prediction submission processes for Critical Assessment of Protein Structure Prediction (CASP4) and Critical Assessment of Fully Automated Structure Prediction Methods (CAFASP2). We have also evaluated all the submitted predictions using criteria and methods developed during the course of three previous CASP experiments and preparation for CASP4. We present an overview of the implemented system. Particular attention is paid to newly developed evaluation techniques and data presentation schemes. With the rapid increase in CASP participation and in the number of submitted predictions, special emphasis is placed on methods allowing reliable pre-classification of submissions and on techniques useful in automated evaluation of predictions. We also present an overview of our website, including target structures, predictions, and their evaluations ( http://predictioncenter.llnl.gov).

Automation↗

Assessment of blind predictions of protein-protein interactions: current status of docking methods.

The current status of docking procedures for predicting protein-protein interactions starting from their three-dimensional structure is assessed from a first major evaluation of blind predictions. This evaluation was performed as part of a communitywide experiment on Critical Assessment of PRedicted Interactions (CAPRI). Seven newly determined structures of protein-protein complexes were available as targets for this experiment. These were the complexes between a kinase and its protein substrate, between a T-cell receptor beta-chain and a superantigen, and five antigen-antibody complexes. For each target, the predictors were given the experimental structures of the free components, or of one free and one bound component in a random orientation. The structure of the complex was revealed only at the time of the evaluation. A total of 465 predictions submitted by 19 groups were evaluated. These groups used a wide range of algorithms and scoring functions, some of which were completely novel. The quality of the predicted interactions was evaluated by comparing residue-residue contacts and interface residues to those in the X-ray structures and by analyzing the fit of the ligand molecules (the smaller of the two proteins in the complex) or of interface residues only, in the predicted versus target complexes. A total of 14 groups produced predictions, ranking from acceptable to highly accurate for five of the seven targets. The use of available biochemical and biological information, and in one instance structural information, played a key role in achieving this result. It was essential for identifying the native binding modes for the five correctly predicted targets, including the kinase-substrate complex where the enzyme changes conformation on association. But it was also the cause for missing the correct solution for the two remaining unpredicted targets, which involve unexpected antigen-antibody binding modes. Overall, this analysis reveals genuine progress in docking procedures but also illustrates the remaining serious limitations and points out the need for better scoring functions and more effective ways for handling conformational flexibility.

Algorithms↗

Predicting peptide binding to MHC pockets via molecular modeling, implicit solvation, and global optimization.

Development of a computational prediction method based on molecular modeling, global optimization, and implicit solvation has produced accurate structure and relative binding affinity predictions for peptide amino acids binding to five pockets of the MHC molecule HLA-DRB1*0101. Because peptide binding to MHC molecules is essential to many immune responses, development of such a method for understanding and predicting the forces that drive binding is crucial for pharmaceutical design and disease treatment. Underlying the development of this prediction method are two hypotheses. The first is that pockets formed by the peptide binding groove of MHC molecules are independent, separating the prediction of peptide amino acids that bind within individual pockets from those that bind between pockets. The second hypothesis is that the native state of a system composed of an amino acid bound to a protein pocket corresponds to the system's lowest free energy. The prediction method developed from these hypotheses uses atomistic-level modeling, deterministic global optimization, and three methods of implicit solvation: solvent-accessible area, solvent-accessible volume, and Poisson-Boltzmann electrostatics. The method predicts relative binding affinities of peptide amino acids for pockets of HLA-DRB1*0101 by determining computationally an amino acid's global minimum energy conformation. Prediction results from the method are in agreement with X-ray crystallography data and experimental binding assays.

Amino Acids↗

Prediction of protein B-factor profiles.

The polypeptide backbones and side chains of proteins are constantly moving due to thermal motion and the kinetic energy of the atoms. The B-factors of protein crystal structures reflect the fluctuation of atoms about their average positions and provide important information about protein dynamics. Computational approaches to predict thermal motion are useful for analyzing the dynamic properties of proteins with unknown structures. In this article, we utilize a novel support vector regression (SVR) approach to predict the B-factor distribution (B-factor profile) of a protein from its sequence. We explore schemes for encoding sequences and various settings for the parameters used in SVR. Based on a large dataset of high-resolution proteins, our method predicts the B-factor distribution with a Pearson correlation coefficient (CC) of 0.53. In addition, our method predicts the B-factor profile with a CC of at least 0.56 for more than half of the proteins. Our method also performs well for classifying residues (rigid vs. flexible). For almost all predicted B-factor thresholds, prediction accuracies (percent of correctly predicted residues) are greater than 70%. These results exceed the best results of other sequence-based prediction methods.

Algorithms↗

Assessment of predictions submitted for the CASP6 comparative modeling category.

Here we present a full overview of the Critical Assessment of Protein Structure Prediction (CASP6) comparative modeling category. Prediction accuracy for the 43 comparative modeling targets was assessed through detailed numerical comparisons between predicted and experimental structures. Assessments using standard measures for model backbone quality and structural alignment accuracy highlighted a small number of groups with stand out predictions and these findings were backed up by statistical comparisons. We were able to carry out evaluations of side-chain contacts predictions and side-chain rotamer accuracy, for which one group turned out to have statistically better predictions. We also assessed the prediction quality of structurally divergent regions and biologically important sites. Interestingly we were able to show that predictors were not predicting these important functional regions with any greater accuracy than the rest of the structure. In addition we investigated the ability of predictors to build models that improve on the structural template and reached some tentative conclusions from comparisons with the previous CASP experiment.

Algorithms↗

Structural prediction of peptides binding to MHC class I molecules.

Peptide binding to class I major histocompatibility complex (MHCI) molecules is a key step in the immune response and the structural details of this interaction are of importance in the design of peptide vaccines. Algorithms based on primary sequence have had success in predicting potential antigenic peptides for MHCI, but such algorithms have limited accuracy and provide no structural information. Here, we present an algorithm, PePSSI (peptide-MHC prediction of structure through solvated interfaces), for the prediction of peptide structure when bound to the MHCI molecule, HLA-A2. The algorithm combines sampling of peptide backbone conformations and flexible movement of MHC side chains and is unique among other prediction algorithms in its incorporation of explicit water molecules at the peptide-MHC interface. In an initial test of the algorithm, PePSSI was used to predict the conformation of eight peptides bound to HLA-A2, for which X-ray data are available. Comparison of the predicted and X-ray conformations of these peptides gave RMSD values between 1.301 and 2.475 A. Binding conformations of 266 peptides with known binding affinities for HLA-A2 were then predicted using PePSSI. Structural analyses of these peptide-HLA-A2 conformations showed that peptide binding affinity is positively correlated with the number of peptide-MHC contacts and negatively correlated with the number of interfacial water molecules. These results are consistent with the relatively hydrophobic binding nature of the HLA-A2 peptide binding interface. In summary, PePSSI is capable of rapid and accurate prediction of peptide-MHC binding conformations, which may in turn allow estimation of MHCI-peptide binding affinity.

Algorithms↗

Predicting G-protein coupled receptors-G-protein coupling specificity based on autocross-covariance transform.

Determining G-protein coupled receptors (GPCRs) coupling specificity is very important for further understanding the functions of receptors. A successful method in this area will benefit both basic research and drug discovery practice. Previously published methods rely on the transmembrane topology prediction at training step, even at prediction step. However, the transmembrane topology predicted by even the best algorithm is not of high accuracy. In this study, we developed a new method, autocross-covariance (ACC) transform based support vector machine (SVM), to predict coupling specificity between GPCRs and G-proteins. The primary amino acid sequences are translated into vectors based on the principal physicochemical properties of the amino acids and the data are transformed into a uniform matrix by applying ACC transform. SVMs for nonpromiscuous coupled GPCRs and promiscuous coupled GPCRs were trained and validated by jackknife test and the results thus obtained are very promising. All classifiers were also evaluated by the test datasets with good performance. Besides the high prediction accuracy, the most important feature of this method is that it does not require any transmembrane topology prediction at either training or prediction step but only the primary sequences of proteins. The results indicate that this relatively simple method is applicable. Academic users can freely download the prediction program at http://www.scucic.net/group/database/Service.asp.

Algorithms↗

Structural analysis and prediction of protein mutant stability using distance and torsion potentials: role of secondary structure and solvent accessibility.

Analyzing the factors behind protein stability is a key research topic in molecular biology, and has direct implications on protein structure prediction and protein-protein interactions. We have analyzed protein stability upon point mutations using a distance-dependant pair potential representing mainly through-space interactions, and torsion angle potential representing mainly neighboring effects as a basic statistical mechanical setup for the analysis. The synergetic effect of accessible surface area and secondary structure preferences was used as a classifier for the potentials. In addition, short-, medium-, and long-range interactions of the protein environment were also analyzed. Two datasets of point mutations were taken for the comparison of theoretically predicted stabilizing energy values with experimental DeltaDeltaG and DeltaDeltaGH(2)O from thermal and chemical denaturation experiments. These include 1538 and 1603 mutations, respectively, and contain 101 proteins that share a wide range of sequence identity. The resulting force fields were carefully evaluated with different statistical tests. Results show a maximum correlation of 0.87 with a standard error of 0.71 kcal/mol between predicted and measured DeltaDeltaG values and a prediction accuracy of 85.3% (stabilizing or destabilizing) for all mutations together. A correlation of 0.77 (more than 80% prediction accuracy with a standard error of 0.95 kcal/mol) each for the test dataset of split-sample validation and fivefold crossvalidation was obtained and a correlation of 0.70 (77.4% prediction accuracy with a standard error of 1.17 kcal/mol) was shown by the jackknife test. The same model was implemented, and the results were analyzed for mutations with DeltaDeltaGH(2)O. A correlation of 0.78 (standard error 0.96 kcal/mol) was observed with a prediction efficiency of 84.65%. This model can be used for the future prediction of protein structural stability together with various experimental techniques.

Databases, Protein↗

Measures for the assessment of fuzzy predictions of protein secondary structure.

Many of the recent secondary structure prediction methods incorporate the idea of fuzzy set theory, where instead of assigning a definite secondary structure to a query residue, probability for the residue being in each of the conformational states is estimated. Moreover, continuous assignment of conformational states to the experimentally observed protein structures can be performed in order to reflect inherent flexibility. Although various measures have been developed for evaluating performances of secondary structure prediction methods, they depend only on the most probable secondary structures. They do not assess the accuracy of the probabilities produced by fuzzy prediction methods, and they cannot incorporate information contained in continuous assignments of conformational states to observed structures. Three important measures for evaluating performance of a secondary structure prediction algorithm, Q score, Segment OVerlap (SOV) measure, and the k-state correlation coefficient (Corr), are deformed into fuzzy measures F score, Fuzzy OVerlap (FOV) measure, and the fuzzy correlation coefficient (Forr), so that the new measures not only assess probabilistic outputs of fuzzy prediction methods, but also incorporate information from continuous assignments of secondary structure. As an example of application, prediction results of four fuzzy secondary structure prediction methods, PSIPRED, PROFking, SABLE, and PREDICT, are assessed using the new fuzzy measures.

Computer Simulation↗

Estimating the predictive quality of dose-response after model selection.

Prediction of dose-response is important in dose selection in drug development. As the true dose-response shape is generally unknown, model selection is frequently used, and predictions based on the final selected model. Correctly assessing the quality of the predictions requires accounting for the uncertainties caused by the model selection process, which has been difficult. Recently, a new approach called data perturbation has emerged. It allows important predictive characteristics be computed while taking model selection into consideration. We study, through simulation, the performance of data perturbation in estimating standard error of parameter estimates and prediction errors. Data perturbation was found to give excellent prediction error estimates, although at times large Monte Carlo sizes were needed to obtain good standard error estimates. Overall, it is a useful tool to characterize uncertainties in dose-response predictions, with the potential of allowing more accurate dose selection in drug development. We also look at the influence of model selection on estimation bias. This leads to insights into candidate model choices that enable good dose-response prediction.

Bias↗

Cross-validation performance of mortality prediction models.

Mortality prediction models hold substantial promise as tools for patient management, quality assessment, and, perhaps, health care resource allocation planning. Yet relatively little is known about the predictive validity of these models. We report here a comparison of the cross-validation performance of seven statistical models of patient mortality: (1) ordinary-least-squares (OLS) regression predicting 0/1 death status six months after admission; (2) logistic regression; (3) Cox regression; (4-6) three unit-weight models derived from the logistic regression, and (7) a recursive partitioning classification technique (CART). We calculated the following performance statistics for each model in both a learning and test sample of patients, all of whom were drawn from a nationally representative sample of 2558 Medicare patients with acute myocardial infarction: overall accuracy in predicting six-month mortality, sensitivity and specificity rates, positive and negative predictive values, and per cent improvement in accuracy rates and error rates over model-free predictions (i.e., predictions that make no use of available independent variables). We developed ROC curves based on logistic regression, the best unit-weight model, the single best predictor variable, and a series of CART models generated by varying the misclassification cost specifications. In our sample, the models reduced model-free error rates at the patient level by 8-22 per cent in the test sample. We found that the performance of the logistic regression models was marginally superior to that of other models. The areas under the ROC curves for the best models ranged from 0.61 to 0.63. Overall predictive accuracy for the best models may be adequate to support activities such as quality assessment that involve aggregating over large groups of patients, but the extent to which these models may be appropriately applied to patient-level resource allocation planning is less clear.

Discriminant Analysis↗

An analysis of the Candida albicans genome database for soluble secreted proteins using computer-based prediction algorithms.

We sought to identify all genes in the Candida albicans genome database whose deduced proteins would likely be soluble secreted proteins (the secretome). While certain C. albicans secretory proteins have been studied in detail, more data on the entire secretome is needed. One approach to rapidly predict the functions of an entire proteome is to utilize genomic database information and prediction algorithms. Thus, we used a set of prediction algorithms to computationally define a potential C. albicans secretome. We first assembled a validation set of 47 C. albicans proteins that are known to be secreted and 47 that are known not to be secreted. The presence or absence of an N-terminal signal peptide was correctly predicted by SignalP version 2.0 in 47 of 47 known secreted proteins and in 47 of 47 known non-secreted proteins. When all 6165 C. albicans ORFs from CandidaDB were analysed with SignalP, 495 ORFs were predicted to encode proteins with N-terminal signal peptides. In the set of 495 deduced proteins with N-terminal signal peptides, 350 were predicted to have no transmembrane domains (or a single transmembrane domain at the extreme N-terminus) and 300 of these were predicted not to be GPI-anchored. TargetP was used to eliminate proteins with mitochondrial targeting signals, and the final computationally-predicted C. albicans secretome was estimated to consist of up to 283 ORFs. The C. albicans secretome database is available at http://info.med.yale.edu/intmed/infdis/candida/

Algorithms↗

Enhanced prediction accuracy of protein secondary structure using hydrogen exchange Fourier transform infrared spectroscopy.

A novel equilibrium hydrogen exchange Fourier transform IR (HX-FTIR) spectroscopy method for predicting secondary structure content was employed using spectra obtained for a training set of 23 globular proteins. The IR bandshape and frequency changes resulting from controlled levels of H-D exchange were observed to be protein-dependent. Their analysis revealed these variations to be partly correlated to secondary structure. For each protein, a set of 6 spectra was measured with a systematic variation of the solvent H-D ratio and was subjected to factor analysis. The most significant component spectra for each protein, representing independent aspects of the spectral response to deuteration, were each subjected to a second factor analysis over the entire training set. Restricted multiple regression (RMR) analysis using the loadings of the principal components from 19 of these H-D analyses revealed an improvement in prediction accuracy compared with conventional bandshape-based analyses of FTIR data. Nearly a factor of 2 reduction in error for prediction of helix fractions was found using s1, the average spectral response for the H-D set. In some cases, significant error reduction for prediction of minor components was found using higher factors. Using the same analytical methods, prediction errors with this new deuteration-response-FTIR method were shown to be even better than those obtained by use of electronic circular dichroism (ECD) data for helix predictions and to be significantly lower for ECD-based sheet prediction, making these the best secondary structure predictions obtained with the RMR method. Tests of a limited variable selection scheme showed further improvements, consistent with previous results of this approach using ECD data.

Deuterium Oxide↗

Prediction of protein secondary structure at better than 70% accuracy.

We have trained a two-layered feed-forward neural network on a non-redundant data base of 130 protein chains to predict the secondary structure of water-soluble proteins. A new key aspect is the use of evolutionary information in the form of multiple sequence alignments that are used as input in place of single sequences. The inclusion of protein family information in this form increases the prediction accuracy by six to eight percentage points. A combination of three levels of networks results in an overall three-state accuracy of 70.8% for globular proteins (sustained performance). If four membrane protein chains are included in the evaluation, the overall accuracy drops to 70.2%. The prediction is well balanced between alpha-helix, beta-strand and loop: 65% of the observed strand residues are predicted correctly. The accuracy in predicting the content of three secondary structure types is comparable to that of circular dichroism spectroscopy. The performance accuracy is verified by a sevenfold cross-validation test, and an additional test on 26 recently solved proteins. Of particular practical importance is the definition of a position-specific reliability index. For half of the residues predicted with a high level of reliability the overall accuracy increases to better than 82%. A further strength of the method is the more realistic prediction of segment length. The protein family prediction method is available for testing by academic researchers via an electronic mail server.

Mathematical Computing↗

Prediction of local structure in proteins using a library of sequence-structure motifs.

We describe a new method for local protein structure prediction based on a library of short sequence pattern that correlate strongly with protein three-dimensional structural elements. The library was generated using an automated method for finding correlations between protein sequence and local structure, and contains most previously described local sequence-structure correlations as well as new relationships, including a diverging type-II beta-turn, a frayed helix, and a proline-terminated helix. The query sequence is scanned for segments 7 to 19 residues in length that strongly match one of the 82 patterns in the library. Matching segments are assigned the three-dimensional structure characteristic of the corresponding sequence pattern, and backbone torsion angles for the entire query sequence are then predicted by piecing together mutually compatible segment predictions. In predictions of local structure in a test set of 55 proteins, about 50% of all residues, and 76% of residues covered by high-confidence predictions, were found in eight-residue segments within 1.4 A of their true structures. The predictions are complementary to traditional secondary structure predictions because they are considerably more specific in turn regions, and may contribute to ab initio tertiary structure prediction and fold recognition.

Algorithms↗

Extending the accuracy limits of prediction for side-chain conformations.

Current techniques for the prediction of side-chain conformations on a fixed backbone have an accuracy limit of about 1.0-1.5 A rmsd for core residues. We have carried out a detailed and systematic analysis of the factors that influence the prediction of side-chain conformation and, on this basis, have succeeded in extending the limits of side-chain prediction for core residues to about 0.7 A rmsd from native, and 94 % and 89 % of chi(1) and chi(1+2 ) dihedral angles correctly predicted to within 20 degrees of native, respectively. These results are obtained using a force-field that accounts for only van der Waals interactions and torsional potentials. Prediction accuracy is strongly dependent on the rotamer library used. That is, a complete and detailed rotamer library is essential. The greatest accuracy was obtained with an extensive rotamer library, containing over 7560 members, in which bond lengths and bond angles were taken from the database rather than simply assuming idealized values. Perhaps the most surprising finding is that the combinatorial problem normally associated with the prediction of the side-chain conformation does not appear to be important. This conclusion is based on the fact that the prediction of the conformation of a single side-chain with all others fixed in their native conformations is only slightly more accurate than the simultaneous prediction of all side-chain dihedral angles.

Computational Biology↗

Comparison of inhaled formaldehyde dosimetry predictions with DNA-protein cross-link measurements in the rat nasal passages.

Kimbell and coworkers (Toxicol, Appl. Pharmacol, 121, 253-263, 1993) developed a computational fluid dynamics (CFD) model of a F344 rat nasal passage to quantify local wall mass flux (uptake rate) of inhaled chemical. To simulate formaldehyde uptake, Kimbell et al. assumed that mass transfer of formaldehyde from the air into the nasal lining was fast and complete. This was approximated in the CFD model by setting the formaldehyde concentration at the airway walls to zero. Experimental confirmation of formaldehyde mass-flux predictions is desirable if the CFD model is to be used for predicting formaldehyde dosimetry. The purpose of this study was to see if the CFD model predictions of formaldehyde mass flux are consistent with laboratory data on formaldehyde dosimetry. In this study, a mathematical model of the nasal lining was modified to link CFD dosimetry predictions for inhaled formaldehyde with measured tissue disposition of inhaled gas. This model treats the nasal lining as a single, well-stirred compartment, accounts for formaldehyde reaction via saturable and first-order pathways, and allows comparison of model-predicted DNA-protein cross-links (DPX) with regional DPX measured in formaldehyde-exposed rats. Effective Michaelis-Menten kinetic parameters (Vmax = 3040 microM/min and Km = 59 microM) and a pseudo-first-order rate constant for elimination of formaldehyde by nonsaturable pathways (kf = 6 min-1) were estimated (fit) using an average mass flux derived from experimentally measured uptake of formaldehyde. DPX predictions obtained using the estimated kinetic parameters and linking the CFD model to the nasal-lining model compared well with experimentally measured DPX. The close correlation between predicted and measured DPX in the rat nasal passage supports the CFD model predictions of formaldehyde mass flux at the level of resolution provided by the experimental data.

Administration, Inhalation↗