Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Accuracy of protein flexibility predictions.

Protein structural flexibility is important for catalysis, binding, and allostery. Flexibility has been predicted from amino acid sequence with a sliding window averaging technique and applied primarily to epitope search. New prediction parameters were derived from 92 refined protein structures in an unbiased selection of the Protein Data Bank by developing further the method of Karplus and Schulz (Naturwissenschaften 72:212-213, 1985). The accuracy of four flexibility prediction techniques was studied by comparing atomic temperature factors of known three-dimensional protein structures to predictions by using correlation coefficients. The size of the prediction window was optimized for each method. Predictions made with our new parameters, using an optimized window size of 9 residues in the prediction window, were giving the best results. The difference from another previously used technique was small, whereas two other methods were much poorer. Applicability of the predictions was also tested by searching for known epitopes from amino acid sequences. The best techniques predicted correctly 20 of 31 continuous epitopes in seven proteins. Flexibility parameters have previously been used for calculating protein average flexibility indices which are inversely correlated to protein stability. Indices with the new parameters showed better correlation to protein stability than those used previously; furthermore they had relationship even when the old parameters failed.

Antigens↗

Protein structure prediction by threading methods: evaluation of current techniques.

This paper evaluates the results of a protein structure prediction contest. The predictions were made using threading procedures, which employ techniques for aligning sequences with 3D structures to select the correct fold of a given sequence from a set of alternatives. Nine different teams submitted 86 predictions, on a total of 21 target proteins with little or no sequence homology to proteins of known structure. The 3D structures of these proteins were newly determined by experimental methods, but not yet published or otherwise available to the predictors. The predictions, made from the amino acid sequence alone, thus represent a genuine test of the current performance of threading methods. Only a subset of all the predictions is evaluated here. It corresponds to the 44 predictions submitted for the 11 target proteins seen to adopt known folds. The predictions for the remaining 10 proteins were not analyzed, although weak similarities with known folds may also exist in these proteins. We find that threading methods are capable of identifying the correct fold in many cases, but not reliably enough as yet. Every team predicts correctly a different set of targets, with virtually all targets predicted correctly by at least one team. Also, common folds such as TIM barrels are recognized more readily than folds with only a few known examples. However, quite surprisingly, the quality of the sequence-structure alignments, corresponding to correctly recognized folds, is generally very poor, as judged by comparison with the corresponding 3D structure alignments. Thus, threading can presently not be relied upon to derive a detailed 3D model from the amino acid sequence. This raises a very intriguing question: how is fold recognition achieved? Our analysis suggests that it may be achieved because threading procedures maximize hydrophobic interactions in the protein core, and are reasonably good at recognizing local secondary structure.

Amino Acid Sequence↗

A comparison of regression trees, logistic regression, generalized additive models, and multivariate adaptive regression splines for predicting AMI mortality.

Clinicians and health service researchers are frequently interested in predicting patient-specific probabilities of adverse events (e.g. death, disease recurrence, post-operative complications, hospital readmission). There is an increasing interest in the use of classification and regression trees (CART) for predicting outcomes in clinical studies. We compared the predictive accuracy of logistic regression with that of regression trees for predicting mortality after hospitalization with an acute myocardial infarction (AMI). We also examined the predictive ability of two other types of data-driven models: generalized additive models (GAMs) and multivariate adaptive regression splines (MARS). We used data on 9484 patients admitted to hospital with an AMI in Ontario. We used repeated split-sample validation: the data were randomly divided into derivation and validation samples. Predictive models were estimated using the derivation sample and the predictive accuracy of the resultant model was assessed using the area under the receiver operating characteristic (ROC) curve in the validation sample. This process was repeated 1000 times-the initial data set was randomly divided into derivation and validation samples 1000 times, and the predictive accuracy of each method was assessed each time. The mean ROC curve area for the regression tree models in the 1000 derivation samples was 0.762, while the mean ROC curve area of a simple logistic regression model was 0.845. The mean ROC curve areas for the other methods ranged from a low of 0.831 to a high of 0.851. Our study shows that regression trees do not perform as well as logistic regression for predicting mortality following AMI. However, the logistic regression model had performance comparable to that of more flexible, data-driven models such as GAMs and MARS.

Data Interpretation, Statistical↗

Frequency, probability, and prediction: easy solutions to cognitive illusions?

Many errors in probabilistic judgment have been attributed to people's inability to think in statistical terms when faced with information about a single case. Prior theoretical analyses and empirical results imply that the errors associated with case-specific reasoning may be reduced when people make frequentistic predictions about a set of cases. In studies of three previously identified cognitive biases, we find that frequency-based predictions are different from-but no better than-case-specific judgments of probability. First, in studies of the "planning fallacy, " we compare the accuracy of aggregate frequency and case-specific probability judgments in predictions of students' real-life projects. When aggregate and single-case predictions are collected from different respondents, there is little difference between the two: Both are overly optimistic and show little predictive validity. However, in within-subject comparisons, the aggregate judgments are significantly more conservative than the single-case predictions, though still optimistically biased. Results from studies of overconfidence in general knowledge and base rate neglect in categorical prediction underline a general conclusion. Frequentistic predictions made for sets of events are no more statistically sophisticated, nor more accurate, than predictions made for individual events using subjective probability.

Bayes Theorem↗

Use of amino acid environment-dependent substitution tables and conformational propensities in structure prediction from aligned sequences of homologous proteins. II. Secondary structures.

A three-step method is presented to predict secondary structures of proteins, by utilizing aligned sequences of homologous proteins. Mean propensities and amino acid substitution patterns at a given site in the aligned sequences are first evaluated for four conformational states (i.e. alpha-helix, beta-strand, buried coil and exposed coil). Capping rules are applied in order to define boundaries of the secondary-structure segments more precisely. In the second step beta-strand is predicted by searching regions predicted as coil for the two patterns characteristic of alternating and fully buried beta-strands. The complete sequences of the solvent-accessibility classes predicted by substitution tables and propensities are also searched using Fourier transform methods for alpha-helical periodicity. After applying capping rules, the alpha-helices and beta-strands predicted in the second step replace, where appropriate, the conformational states predicted in the first step. Finally, in the third step, if one of the four conformational states is assigned to the residues at an equivalent site of aligned sequences in more than a given fraction of the proteins, such a state is reassigned to all the residues at that site. The method is applied to 13 protein families, which contain four folding types, alpha, beta, alpha/beta and alpha + beta. The accuracy of the prediction ranges from 60 to 79% (mean percentage over the 13 families is 69%). For comparison the Garnier-Osguthorpe-Robson (GOR) method is also applied to them. Although the mean prediction accuracy for the GOR method, 58%, can be improved to 63% by applying the second and third steps in this method, there remain four families with less than 55% accuracy. The mean accuracy is relatively higher and poor predictions are reduced in this method.

Amino Acid Sequence↗

Predictions of secondary structure using statistical analyses of electronic and vibrational circular dichroism and Fourier transform infrared spectra of proteins in H2O.

Vibrational circular dichroism (VCD) and Fourier transform IR (FTIR) methods for prediction of protein secondary structure are systematically compared using selective regression analysis. VCD and FTIR spectra over the amide I and II bands of 23 proteins dissolved in H2O were analyzed using the principal component method of factor analysis (PC/FA) and regression fits to fractional components (FC) of secondary structure. Predictive capability was determined by computing structures for proteins sequentially left out of the regression. All possible combinations of PC/FA spectral parameters (coefficients) were used to form a full set of restricted multiple regressions (RMR) of PC/FA coefficients with FC values, both independently for each spectral data set as well as for the VCD and FTIR sets grouped together and with similarly obtained electronic CD (ECD) data. The distribution of predictive error for a set of the best RMR relationships that use a given number of spectral coefficients was used to select the optimal prediction algorithm. Minimum predictive error resulted for a small subset (three to six) of spectral coefficients, which is consistent with our earlier findings using VCD measured for proteins in 2H2O and ECD data. Subtracting the average absorption spectrum from all the training set FTIR spectra before analysis yields more variance in the FTIR band shape and improves the predictive ability of the best PC/FA RMR to near that for the VCD. Both methods (FTIR and VCD) using data for proteins in H2O are somewhat better predictors than amide I' (in 2H2O) VCD alone and, for helix, worse than ECD alone. Combining FTIR and VCD data did not dramatically change the prediction results. Predictions are improved by combining both with ECD data, indicating that the improvement is due to using their very different structural sensitivities. The coupled H2O-based spectral analyses and the mixed amide I' + II VCD plus ECD analysis are comparable for the helix and sheet components, indicating that partial deuteration is not a major source of prediction error.

Circular Dichroism↗

RNA secondary structure prediction based on free energy and phylogenetic analysis.

We describe a computational method for the prediction of RNA secondary structure that uses a combination of free energy and comparative sequence analysis strategies. Using a homology-based sequence alignment as a starting point, all favorable pairings with respect to the Turner energy function are identified. Each potentially paired region within a multiple sequence alignment is scored using a function that combines both predicted free energy and sequence covariation with optimized weightings. High scoring regions are ranked and sequentially incorporated to define a growing secondary structure. Using a single set of optimized parameters, it is possible to accurately predict the foldings of several test RNAs defined previously by extensive phylogenetic and experimental data (including tRNA, 5 S rRNA, SRP RNA, tmRNA, and 16 S rRNA). The algorithm correctly predicts approximately 80% of the secondary structure. A range of parameters have been tested to define the minimal sequence information content required to accurately predict secondary structure and to assess the importance of individual terms in the prediction scheme. This analysis indicates that prediction accuracy most strongly depends upon covariational information and only weakly on the energetic terms. However, relatively few sequences prove sufficient to provide the covariational information required for an accurate prediction. Secondary structures can be accurately defined by alignments with as few as five sequences and predictions improve only moderately with the inclusion of additional sequences.

Algorithms↗

Prediction of the secondary structure content of globular proteins based on structural classes.

The prediction of the secondary structure content (alpha-helix and beta-strand content) of a globular protein may play an important complementary role in the prediction of the protein's structure. We propose a new prediction algorithm based on Chou's database [Chou (1995), Proteins Struct. Funct. Genet. 21, 319]. The new algorithm is an improved multiple linear regression method, taking the nonlinear and coupling terms of the frequencies of different amino acids into account. The prediction is also based on the structural classes of proteins. A resubstitution examination for the algorithm shows that the average errors are 0.040 and 0.033 for the prediction of alpha-helix content and beta-strand content, respectively. The examination of cross-validation, the jackknife analysis, shows that the average errors are 0.051 and 0.044 for the prediction of alpha-helix content and beta-strand content, respectively. Both examinations indicate the self-consistency and the extrapolative effectiveness of the new algorithm. Compared with the other methods available currently, our method has the merits of simplicity and convenience for use, as well as a high prediction accuracy. By incorporating the prediction of the structural classes, the only input of our method is the amino acid composition of the protein to be predicted.

Algorithms↗

The inability of physicians to predict the outcome of in-hospital resuscitation.

OBJECTIVE: To measure the accuracy, reliability, and discrimination of physicians' predictions of the outcome of in-hospital cardiopulmonary resuscitation (CPR), using a large series of detailed clinical vignettes of patients with known outcomes. DESIGN: Faculty and resident physicians at three university-affiliated generalist training programs were given one-page summaries of admission data for patients who later underwent in-hospital CPR. These summaries included all pre-arrest variables known to be related to the outcome of CPR. Physicians were asked to estimate the probability that patients would survive the resuscitation long enough to be stabilized, and the probability of survival to discharge. SETTING: Patient cases were derived from a consecutive series of patients undergoing CPR at two urban teaching hospitals in Detroit, Michigan. PARTICIPANTS: Faculty members and residents at a university-based department of internal medicine and two university-based departments of family medicine were surveyed. INTERVENTIONS: Accuracy of the physician predictions was assessed by comparing the mean predicted probability of survival with the percentage of patients who actually survived. The reliability of probability estimates of survival was evaluated by assessing the numerical proximity of the estimates to the actual outcome of the resuscitative effort. The ability to discriminate between survivors and nonsurvivors was measured by comparing the mean predicted probability of survival for those patients who survived CPR with that for those who did not, and by stratifying physician predictions and measuring the area under a receiver operating characteristic (ROC) curve. MEASUREMENTS AND MAIN RESULTS. Physicians (n = 51) made a total of 713 estimates, and showed poor accuracy, reliability, and discrimination in predicting the outcome of in-hospital CPR. The mean predicted probability of survival to discharge did not differ between patients who actually survived to discharge and those who did not (29.5% vs 26.4%, z = 0.35, p = .73). Similarly, the mean predicted probabilities of surviving resuscitation were the same for patients who actually survived long enough to be stabilized and those who did not (37.8% vs 39.9%, z = 0.55, p = .58). Accounting for type of physician and institution by analysis of variance did not change this finding. The area under the ROC curve for the prediction of arrest survival was 0.476, which is not significantly different from 0.5, and is consistent with an ability to discriminate between survivors and nonsurvivors that is no better than random choice. CONCLUSIONS: Physicians were no better at identifying patients who would survive resuscitation than would be expected by chance alone. Further work is needed to establish which variables are used by physicians in the decision-making process, and to design educational interventions that will make physicians more accurate prognosticators.

Adult↗

Audit of a rural hospital's coronary care unit: comparison of two predictive instruments of acute myocardial infarction mortality.

The purpose of the study was to compare observed mortality of a rural hospital coronary care unit with mortality rates estimated by two predictive instruments of mortality. The mortality rates of 86 consecutive patients with confirmed acute myocardial infarction were compared with those predicted by the presence or absence of eight risk factors for mortality identified by the Thrombolysis in Myocardial Infarction (TIMI) trial, and mortality predicted by a logistic regression equation LRE). Seventeen patients (20 per cent) died within 6 weeks of admission; the number of TIMI risk factors present predicted a mortality of 9.8 per cent, and the instrument of Selker's predicted a mortality of 25.9 per cent. Patients with 3 TIMI risk factors had a significantly higher mortality than predicted (46.2 versus 13.0 per cent, p < 0.01). There were no significant differences between the receiver operating characteristic (ROC) curve of either instrument. The predictions of Selker's instrument, however, showed no significant difference from observed mortality, even when the patients were grouped into quintiles, and the predicted mortality rates were corrected for any presumed benefit from thrombolysis. The predictive instrument of Selker more consistently estimates observed mortality than the presence of risk factors identified by the TIMI trial.

Aged↗

Three-dimensional soft tissue prediction using finite elements. Part II: Clinical application.

BACKGROUND AND AIM: The goal of this study was to analyze the validity and prediction accuracy of a newly-developed procedure for three-dimensional soft tissue prediction based on Finite Element Method, and to compare the results with prediction produced using an existing two-dimensional prediction program (Dentofacial Planner Plus). PATIENTS AND METHODS: In twelve patients who underwent combined surgical-orthodontic treatment, profile prediction was generated using both procedures preoperatively and then compared at predefined measurement points with the patient's actual postoperative soft tissue status. RESULTS: The deviations observed depended on the facial region, whereby the prediction errors for both procedures were much greater in the lower facial third than in the midfacial third. Calculating in all the measurement points, the mean horizontal prediction error was 0.32 mm for the Finite Element Method and 0.75 mm for the Dentofacial Planner Plus. Overall, we were able to demonstrate the new procedure's superior validity and quality of visualization. In addition to profile prediction, the procedure allows a differentiated three-dimensional assessment of esthetically important regions such as the cheeks, nasolabial folds and the nasal wings. Additional X-radiation is not necessary in this risk-free and stress-free procedure. CONCLUSION: Three-dimensional soft tissue prediction employing finite element modeling is a useful aid for implementing esthetically-optimized treatment planning.

Adult↗

Automated generation and evaluation of specific MHC binding predictive tools: ARB matrix applications.

Prediction of which peptides can bind major histocompatibility complex (MHC) molecules is commonly used to assist in the identification of T cell epitopes. However, because of the large numbers of different MHC molecules of interest, each associated with different predictive tools, tool generation and evaluation can be a very resource intensive task. A methodology commonly used to predict MHC binding affinity is the matrix or linear coefficients method. Herein, we described Average Relative Binding (ARB) matrix methods that directly predict IC(50) values allowing combination of searches involving different peptide sizes and alleles into a single global prediction. A computer program was developed to automate the generation and evaluation of ARB predictive tools. Using an in-house MHC binding database, we generated a total of 85 and 13 MHC class I and class II matrices, respectively. Results from the automated evaluation of tool efficiency are presented. We anticipate that this automation framework will be generally applicable to the generation and evaluation of large numbers of MHC predictive methods and tools, and will be of value to centralize and rationalize the process of evaluation of MHC predictions. MHC binding predictions based on ARB matrices were made available at http://epitope.liai.org:8080/matrix web server.

Animals↗

Variation of fractionator estimates and its prediction.

Over the last decade the so-called 'fractionator' has become widespread in anatomical and pathological research for obtaining unbiased estimates of total numbers of particles in biological specimens. Several methods have been proposed for predicting the precision (i.e., the variation) of the estimated total numbers of particles using the fractionator (i.e., for predicting the precision of fractionator estimates). However, the validity of these predicting methods has not been tested so far. As it is impossible to do so with biological experiments, it was carried out here by using a computer simulation. Specimens containing particles, with various particle distributional patterns, were modeled, and the total number of particles in the specimens was estimated repeatedly with various modeled sampling schemes. It could be shown that the empirically estimated precision of the modeled fractionator estimates depend on both the particle distributional pattern in a modeled specimen as well as on the applied sampling scheme. Furthermore, considerable differences between the predicted and the empirically estimated precision of the modeled fractionator estimates were found. This was due partly to an incorrect assumption, which serves as the basis for one of the proposed predicting methods, partly to the fact that for some of the proposed predicting methods important contributions to the variation of fractionator estimates have not been considered, and partly to the fact that the mathematical theory, which serves as the basis for all predicting methods proposed so far, can in principle not be the optimum basis for predicting the precision of fractionator estimates. Based on the results of the computer simulation, a new, simple method is proposed for predicting the precision of fractionator estimates.

Cell Count↗

Prediction of beta-strand packing interactions using the signature product.

The prediction of beta-sheet topology requires the consideration of long-range interactions between beta-strands that are not necessarily consecutive in sequence. Since these interactions are difficult to simulate using ab initio methods, we propose a supplementary method able to assign beta-sheet topology using only sequence information. We envision using the results of our method to reduce the three-dimensional search space of ab initio methods. Our method is based on the signature molecular descriptor, which has been used previously to predict protein-protein interactions successfully, and to develop quantitative structure-activity relationships for small organic drugs and peptide inhibitors. Here, we show how the signature descriptor can be used in a Support Vector Machine to predict whether or not two beta-strands will pack adjacently within a protein. We then show how these predictions can be used to order beta-strands within beta-sheets. Using the entire PDB database with ten-fold cross-validation, we have achieved 74.0% accuracy in packing prediction and 75.6% accuracy in the prediction of edge strands. For the case of beta-strand ordering, we are able to predict the correct ordering accurately for 51.3% of the beta-sheets. Furthermore, using a simple confidence metric, we can determine those sheets for which accurate predictions can be obtained. For the top 25% highest confidence predictions, we are able to achieve 95.7% accuracy in beta-strand ordering. [Figure: see text].

Amino Acid Sequence↗

Lazy structure-activity relationships (lazar) for the prediction of rodent carcinogenicity and Salmonella mutagenicity.

lazar is a new tool for the prediction of toxic properties of chemical structures. It derives predictions for query structures from a database with experimentally determined toxicity data. lazar generates predictions by searching the database for compounds that are similar with respect to a given toxic activity and calculating the prediction from their activities. Apart form the prediction, lazar provides the rationales (structural features and similar compounds) for the prediction and a reliable condence index that indicates, if a query structure falls within the applicability domain of the training database.Leave-one-out (LOO) crossvalidation experiments were carried out for 10 carcinogenicity endpoints ({female/male} {hamster/mouse/rat} carcinogenicity and aggregate endpoints {hamster/mouse/rat} carcinogenicity and rodent carcinogenicity) and Salmonella mutagenicity from the Carcinogenic Potency Database (CPDB). An external validation of Salmonella mutagenicity predictions was performed with a dataset of 3895 structures. Leave-one-out and external validation experiments indicate that Salmonella mutagenicity can be predicted with 85% accuracy for compounds within the applicability domain of the CPDB. The LOO accuracy of lazar predictions of rodent carcinogenicity is 86%, the accuracies for other carcinogenicity endpoints vary between 78 and 95% for structures within the applicability domain.

Algorithms↗

The use of a predicted CPAP equation improves CPAP titration success.

Titration of continuous positive airway pressure (CPAP) is performed to determine the CPAP setting to prescribe for an individual patient. A prediction equation has been published that could be used to improve the success rate of CPAP titrations. The goals of this study were: (1) to test the hypothesis that the use of the prediction equation would achieve a higher rate of successful CPAP titrations; (2) to validate the equation as an accurate predictor of the prescribed CPAP setting and determine the factors that influence the accuracy of the prediction equation. A total of 224 patients underwent CPAP titration prior to using the equation, with a starting pressure of 5 cm H(2)O. A total of 192 patients underwent CPAP titration using the equation-predicted CPAP level as the starting pressure (median starting pressure of 8 cm H(2)O [interquartile range 7, 10 cm H(2)O]). The percentage of successful studies, as defined by a 50% decrease in the apnea-hypopnea index (AHI) and a final AHI < or =10 cm H(2)O, increased from 50% to 68% (p<0.001), while the number of patients who were prescribed a CPAP level that had not been tested decreased from 22% to 5% (p<0.001). The equation was not accurate in predicting the prescribed level of CPAP, with only 30.8% of the patients with a prescribed pressure < or =3 cm H(2)O of the predicted pressure. Female gender was the only predictor of a prescribed pressure < or =3 cm H(2)O from the predicted pressure (odds ratio 3.45, 95% confidence intervals 1.67, 7.13, p<0.001). A CPAP prediction equation modestly increases the rate of successful CPAP titrations by increasing the starting pressure of the titration. The equation does not accurately predict the prescribed CPAP level, reaffirming the need for a titration study to determine the optimal prescribed level in a given patient.

Body Mass Index↗

Prediction in road safety studies: an empirical inquiry.

Studies about the road safety effect of interventions are usually retrospective quasi-experiments. In these, one key task is to predict what would have been the safety of the treated group without the intervention. Such predictions can be made by several methods, one of which is to use a "comparison group." We use 26 yearly counts of reported injury accidents for the Canadian provinces to examine which of several simple methods of prediction would have historically predicted best. We find that the use of more data does not always improve prediction. How well one predicts depends not only on the amount of data used but also on the extent to which the prediction method is in accord with the unknown time trend behind the accident counts. We also find that the use of a comparison group to predict is not always better than predicting that this year's count will be the same as last year's. In addition, the intuitive notion that a good comparison group is that which is thought similar to the treated group is too simple. Both similarity and size (as measured by the number of accidents) are important. Moreover, whatever preconceived notions of similarity we had, were contradicted by the data. If the history of accident counts on the treatment group and on several possible comparison groups is available, a simple method to select the most suitable comparison group is suggested.

Accidents, Traffic↗

The response of single units in the auditory cortex of rhesus monkeys to predicted and to unpredicted sound stimuli.

Rhesus monkeys were trained to predict the nature of short auditory signals of two different types, to which they responded differentially. Prediction was based on visual signals that preceded the auditory ones. The monkey's use of the visual signals as predictors was assessed through two behavioural criteria. (1) performance in trials in which the visual signals was followed by the correct auditory signal (true conditioning) vs performance in trials in which the visual signal was followed by the wrong auditory signal (false conditioning). (2) Reaction time in trials with different types of conditioning. The response of single auditory cortex units to correctly and to incorrectly predicted auditory signals was recorded. The unconditioned response of every unit to each type of auditory stimulus was also obtained. Of 92 units that were analysed in detail, the response of about half was affected by the predictability of the stimulus. Units were affected in two different ways. One involved facilitation of the response to correctly predicted signals and inhibition of the response to incorrectly predicted signals. The other involved facilitation of the response to incorrectly predicted signals and base line response to correctly predicted signals. These findings were discussed in terms of neural mechanisms that relate to prediction.

Animals↗