Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Accounting for intubation status in predicting mortality for victims of motor vehicle crashes.

BACKGROUND: Two of the important predictors of mortality for trauma patients are the Glasgow Coma Scale and the respiratory rate. However, for intubated patients, the verbal response component of the Glasgow Coma Scale and the respiratory rate cannot be accurately obtained. This study extends previous work that attempts to predict mortality accurately for intubated patients without using verbal response and respiratory rate. METHODS: The New York State Trauma Registry was used to identify 1994 and 1995 victims of motor vehicle crashes (MVCs). For the subset of patients who were not intubated, we developed two statistical models to predict mortality: one did not contain verbal response or respiratory rate, and the other contained a predicted verbal response. These were compared with a model that did include verbal response and respiratory rate. We also compared the predictive abilities of the first two models for all MVC patients (intubated and nonintubated) and determined the extent to which intubated patients were at increased risk of dying in the hospital after having adjusted for other predictors of mortality. RESULTS: For nonintubated patients, the statistical model without verbal response and the model with predicted verbal response had slightly better discrimination and worse calibration than the model that included verbal response and respiratory rate. Predicted verbal response did not improve the strength of the model without verbal response. For all MVC patients (intubated and nonintubated), predicted verbal response was not a significant predictor of mortality when used in combination with the other predictors. Intubation status was a significant predictor, with intubated patients having a higher probability of dying in the hospital than patients with otherwise identical risk factors. CONCLUSION: Inpatient mortality for intubated MVC patients can be accurately predicted without respiratory rate or verbal response. There appears to be no need for predicted verbal response to be part of the prediction formula, but intubation status is an important independent predictor of mortality and should be used in statistical models that predict mortality for MVC patients.

Accidents, Traffic↗

Intraoperative electromyography for predicting facial function in vestibular schwannoma surgery.

OBJECTIVE: To assess the validity of intraoperative minimal stimulation threshold (MST) for predicting long-term facial function after vestibular schwannoma surgery. STUDY DESIGN: Prospective blinded study. METHODS: MST after tumor dissection and postoperative clinical facial function, assessed using the House Brackmann grading system (HB), were used to predict long-term clinical facial function, recorded at least 6 months after surgery. RESULTS: Two hundred and nine consecutive patients fulfilled selection criteria and 184 had successful intraoperative electrophysiologic monitoring and were eligible for further study. MST of 0.05 mA had moderate accuracy for predicting good long-term facial function, with 94% sensitivity, 91% positive predictive value (PPV), 60% specificity, and 70% negative predictive value (NPV). A more relevant group of 77 patients with poor postoperative facial function (HB III-VI) were assessed for predicting good long-term function. Applying this criteria, test accuracy fell, with 83% sensitivity, 64% PPV, 60% specificity, and 75% NPV. Postoperative clinical facial function had a greater accuracy for predicting good long-term function, with 83% sensitivity, 79% PPV, 75% specificity, and 79% NPV. A model of predicted probabilities of good outcome (HB I and II) was derived from a logistic regression with two additive predictors (postoperative HB and MST). This demonstrated that for patients with postoperative HB grade V, MST aided prediction. CONCLUSIONS: Intraoperative stimulation thresholds, when assessed against a relevant group of patients with poor postoperative facial function, had poor predictive accuracy. The severity of immediate postoperative clinical facial function was the most accurate predictor of long-term outcome. MST aided long-term prediction in a small but relevant group of patients with postoperative HB grade V facial function.

Cerebellar Neoplasms↗

Threshold prediction using the auditory steady-state response and the tone burst auditory brain stem response: a within-subject comparison.

OBJECTIVE: The purpose of this study was to evaluate the accuracy with which auditory steady-state response (ASSR) and tone burst auditory brain stem response (ABR) thresholds predict behavioral thresholds, using a within-subjects design. Because the spectra of the stimuli used to evoke the ABR and the ASSR differ, it was hypothesized that the predictive accuracy also would differ, particularly in subjects with steeply sloping hearing losses. DESIGN: ASSR and ABR thresholds were recorded in a group of 14 adults with normal hearing, 10 adults with flat, sensorineural hearing losses, and 10 adults with steeply sloping, high-frequency, sensorineural hearing losses. Evoked-potential thresholds were recorded at 1, 1.5, and 2 kHz and were compared with behavioral, pure-tone thresholds. The predictive accuracy of two ABR protocols was evaluated: Blackman-gated tone bursts and linear-gated tone bursts presented in a background of notched noise. Two ASSR stimulation protocols also were evaluated: 100% amplitude-modulated (AM) sinusoids and 100% AM plus 25% frequency-modulated (FM) sinusoids. RESULTS: The results suggested there was no difference in the accuracy with which either ABR protocol predicted behavioral threshold, nor was there any difference in the predictive accuracy of the two ASSR protocols. On average, ABR thresholds were recorded 3 dB closer to behavioral threshold than ASSR thresholds. However, in the subjects with the most steeply sloping hearing losses, ABR thresholds were recorded as much as 25 dB below behavioral threshold, whereas ASSR thresholds were never recorded more than 5 dB below behavioral threshold, which may reflect more spread of excitation for the ABR than for the ASSR. In contrast, the ASSR overestimated behavioral threshold in two subjects with normal hearing, where the ABR provided a more accurate prediction of behavioral threshold. CONCLUSIONS: Both the ABR and the ASSR provided reasonably accurate predictions of behavioral threshold across the three subject groups. There was no evidence that the predictive accuracy of the ABR evoked using Blackman-gated tone bursts differed from the predictive accuracy observed when linear-gated tone bursts were presented in conjunction with notched noise. Similarly, there was no evidence that the predictive accuracy of the AM ASSR differed from the AM/FM ASSR. In general, ABR thresholds were recorded at levels closer to behavioral threshold than the ASSR. For certain individuals with steeply sloping hearing losses, the ASSR may be a more accurate predictor of behavioral thresholds; however, the ABR may be a more appropriate choice when predicting behavioral thresholds in a population where the incidence of normal hearing is expected to be high.

Acoustic Stimulation↗

Contribution of neuropsychological data to the prediction of temporal lobe epilepsy surgery outcome.

PURPOSE: We empirically examined the contribution of neuropsychological data to the prediction of postoperative seizure control relative to base rate information in an existing series of patients undergoing anterior temporal lobectomy (ATL). METHODS: A discriminant function predicting surgery outcome (seizure-free vs. non-seizure-free) was computed separately for samples of patients with left (n = 79) and right (n = 62) temporal lobectomy (LATL, RATL). Predictor variables included 14 measures tapping five neurocognitive domains. The predicted base rates were compared with the actual base rates in the two samples. Finally, overall predictive accuracy was examined in optimal versus suboptimal ATL patients. RESULTS: The base rate of seizure freedom in the LATL group was 74.70%; that in the RATL group was 66.10%. The predictive function for the LATL group achieved a hit rate of 80.00% and a positive predictive power of 92.11%. The function for the RATL group achieved a hit rate of 83.33% and a positive predictive power (PPP) of 89.66%. The overall predictive accuracy for the optimal group was only 55%, but that in the suboptimal group was 72%. CONCLUSIONS: Neuropsychological data used in a multivariate statistical fashion may be able to offer an incremental increase in the prediction of postoperative seizure freedom relative to existing base rates of surgery success in patients with ATL epilepsy. The use of neuropsychological data may be of greatest predictive value in a population of ATL candidates with suboptimal findings with a lower base rate of postoperative seizure freedom, but may actually reduce predictive accuracy in a group of ATL candidates from an optimal population with an already high base rate of surgical success.

Adult↗

Measured diffusion capacity versus prediction equation estimates in blacks without lung disease.

BACKGROUND: Lung volumes in African-Americans are on average 10-15% less than in Caucasians for the same height and are race corrected accordingly. Despite this fact, prediction equation estimates (PEE) of diffusion capacity of CO (DL(CO)) developed in Caucasians are not adjusted for lung volume in the black population. This could result in healthy blacks being labeled as abnormal. OBJECTIVE: To test the hypothesis that healthy black subjects might be labeled as abnormal using three commonly used PEE of DL(CO) which are currently used in the United States. METHODS: Forty-two nonsmoking black subjects with no history of any disease underwent DL(CO) testing. Controls consisted of 12 healthy Caucasian volunteers and the prediction equations themselves. The single breath diffusion capacity was used with a Collins system. The measured diffusing capacity was compared with the Miller, Knudson, and Crapo PEE by entering age, gender, height and weight for each subject into the appropriate equation. Abnormal was defined as a DL(CO) <80% predicted. Methane gas dilution and body plethysmography were used to determine alveolar volume. Values in parentheses in the results section are DLCO adjusted for alveolar volume proportions. RESULTS: The average measured DL(CO) in blacks was 25.85 +/- 6.37 ml/min/mm Hg. This value was significantly different (p < 0.01) compared to the predicted DL(CO) of 29.80 +/- 4.77, 36.45 +/- 6.64, and 35.33 +/- 5.27 for the Miller, Knudson, and Crapo equations, respectively. This resulted in 14/42 (0/42), 33/42 (3/42), and 33/42 (9/42) DL(CO) (DL(CO)/VA) measurements being defined as abnormal using the Miller, Knudson, and Crapo prediction equations, respectively. In Caucasians, the average measured DL(CO) was not different from the Miller PEE. However, the measured DL(CO) was significantly lower than the Knudson and Crapo PEE, although less so than in blacks. This resulted in no Caucasian DL(CO) measurements defined as abnormal with the Miller PEE and some with the Knudson and Crapo PEE, but less so than in blacks. The measured alveolar volumes by methane dilution were slightly but not significantly decreased compared to those determined by plethysmography. Both measured values were significantly different (p < 0.01) compared to the predicted alveolar volumes of 6.19 +/- 0.91, 6.38 +/- 1.07, and 6.05 +/- 0.96 liters for the Miller, Knudson, and Crapo PEE in blacks, with no difference in predicted and measured lung volumes in Caucasians. The difference in predicted versus measured DL(CO) measurements in blacks was 13.2, 29.1, and 26.8%, respectively, for the Miller, Knudson, and Crapo prediction equations. These differences were similar to the reduction in predicted values of 22.5, 24.7, and 20.7% for the above-mentioned prediction equations, respectively, versus the measured alveolar volume by methane (in blacks). A race correction (reduction) of the Miller PEE for diffusion of 12% resulted in only 2/42 DL(CO) measurements being labeled as abnormal. CONCLUSIONS: Current PEE for DLCO when used in healthy blacks can result in an abnormal reading in up to 50% or more of the time. This failure of the PEE is related to a reduction in lung volume in African-Americans that is not accounted for. One approach to overcome this problem, until separate PEE are developed in blacks, is to race correct the Miller PEE for diffusion by 12%. This reduces the DL(CO) error to less than 5% for this population.

Adult↗

Predictive validity of three ActiGraph energy expenditure equations for children.

PURPOSE: This study evaluated the predictive validity of three previously published ActiGraph energy expenditure (EE) prediction equations developed for children and adolescents. METHODS: A total of 45 healthy children and adolescents (mean age: 13.7 +/- 2.6 yr) completed four 5-min activity trials (normal walking, brisk walking, easy running, and fast running) in an indoor exercise facility. During each trial, participants wore an ActiGraph accelerometer on the right hip. EE was monitored breath by breath using the Cosmed K4b portable indirect calorimetry system. Differences and associations between measured and predicted EE were assessed using dependent t-tests and Pearson correlations, respectively. Classification accuracy was assessed using percent agreement, sensitivity, specificity, and area under the receiver operating characteristic (ROC) curve. RESULTS: None of the equations accurately predicted mean energy expenditure during each of the four activity trials. Each equation, however, accurately predicted mean EE in at least one activity trial. The Puyau equation accurately predicted EE during slow walking. The Trost equation accurately predicted EE during slow running. The Freedson equation accurately predicted EE during fast running. None of the three equations accurately predicted EE during brisk walking. The equations exhibited fair to excellent classification accuracy with respect to activity intensity, with the Trost equation exhibiting the highest classification accuracy and the Puyau equation exhibiting the lowest. CONCLUSIONS: These data suggest that the three accelerometer prediction equations do not accurately predict EE on a minute-by-minute basis in children and adolescents during overground walking and running. The equations maybe useful, however, for estimating participation in moderate and vigorous activity.

Acceleration↗

Prediction of methane production from dairy cows using existing mechanistic models and regression equations.

Ruminants may contribute to global warming through the release of methane gas by enteric fermentation. Until now, methane emissions from ruminants were estimated using simple regression equations. The objective of this study was to compare the capacity of dynamic and mechanistic models to that of regression equations to predict methane production from dairy cows. The updated version of the model of Baldwin et al. and a modified version of the model of Dijkstra et al. and the regression equations of Blaxter and Clapperton and Moe and Tyrrell were challenged with 32 experimental diets selected from 13 publications. The predictive capacity of mechanistic models and regression equations was evaluated by comparing predicted and observed methane production using regression analysis. Results of regression showed better prediction of methane production with mechanistic models than with regression equations. The modified model of Dijkstra et al. predicted methane production with the higher R2 (.71) and the smaller error of prediction (19.87% of the observed mean). The model of Baldwin et al. predicted methane production with a similar R2 (.70) but a higher error of prediction (36.93%). However, a large proportion of this error can be eliminated by a correction factor. Predictions using the equations of Moe and Tyrrell and Blaxter and Clapperton were poor (R2 = .42 and .57; error of prediction = 33.72% and 22.93%, respectively). This study demonstrated that from a large variation in diet composition, mechanistic models allow the prediction of methane production more accurately than simple regression equations.

Animal Feed↗

Solitary pulmonary nodules: clinical prediction model versus physicians.

OBJECTIVE: To determine whether a clinical prediction model developed to identify malignant lung nodules based on clinical data and radiologic lung nodule characteristics could predict a malignant lung nodule diagnosis with higher accuracy than physicians. MATERIAL AND METHODS: One hundred cases were obtained by using a stratified random sample from a retrospective cohort of 629 patients with newly discovered 4- to 30-mm radiologically indeterminate solitary pulmonary nodules (SPNs) on chest radiography. A chest radiologist, pulmonologist, thoracic surgeon, and general internist made predictions of a malignant lesion and recommendations for management (thoracotomy, transthoracic needle aspiration biopsy, or observation) on the basis of radiologic and clinical data used to develop the clinical prediction rule. The predictions of a malignant lung nodule were compared with the probability of malignant involvement from a previously validated clinical prediction model to identify malignant nodules on the basis of three clinical characteristics (age, smoking status, and history of cancer greater than or equal to 5 years previously) and three radiologic characteristics (nodule diameter, spiculation, and upper lobe location). RESULTS: Receiver operating characteristic analysis showed no significant difference between the logistic model and the physicians' predictions. Calibration curves revealed that physicians overestimated the probability of a malignant lesion in patients with low risk of malignant disease by the prediction rule; this finding suggests a potential for the decision rule to improve the management of patients with SPNs that are likely to be benign. CONCLUSION: The prediction model was not better than physicians' predictions of malignant SPNs. The prediction rule may have potential to improve the management of patients with SPNs that are likely to be benign.

Diagnosis, Differential↗

[Early prediction of lack of response to treatment with interferon and interferon plus ribavirin using biochemical and virological criteria in patients with chronic hepatitis C].

The objectives of this study were the following: 1) to evaluate the predictive value of the detection of RNA-HVC compared to GPT in the third month of treatment in patients with chronic hepatitis C treated with IFN, and at the first and third month in patients treated with IFN and ribavirin for 6 and 12 months. The study included: A) 80/132 patients treated with IFN (3 MU/3 times a week for 6-12 months), and B) 70/110 patients who had previously not responded to IFN, and who were treated with combination therapy (IFN: standard dose, ribavirin: 1200 mg/day) for 6 months (n = 40) and 12 months (n = 30). In group A, the positive predictive value (the probability of predicting the lack of response if the RNA-HVC was positive or if the GPT was elevated at the third month) was greater for RNA-HVC than for GPT (97.9% vs. 94.4%), although the response was not unequivocal (2.3% vs. 10.5%). The negative predictive value was 48.6% vs. 36.2%, respectively. The prediction level (odds ratio) of RNA-HVC and of GPT was 39.7 vs. 8.78 (p <0.000001 vs. p <0.002). The positive predictive value was 97.6% in patients with genotype 1, 4 and 5, and 100% in those with genotype 2 and 3. In group B, the positive predictive value was also greater for RNA-HVC than for GPT at the first month (100% vs. 94.4%) following six months of therapy, the odds ratio being infinite vs. 7.6. The positive predictive value was greater for RNA-HVC at the third month than at the first (100% vs. 91%), whereas it was similar for GPT (100%) with 12 months of therapy, the odds ratio being greater for GPT than for RNA-HVC at the first month (infinite and 7.27). The following was concluded: 1) detection of RNA-HVC at the third month of treatment with IFN predicts in advance a lack of response in patients, with a minimum risk of error; 2) in patients with six months of combined therapy, the detection of RNA-HVC at the first month is extremely reliable in the prediction of a lack of response, whereas after 12 months of combined therapy, elevated GPT values at the first month and the detection of RNA-HVC at the third are highly predictive of a lack of response.

Adolescent↗

Assessment of a simple artificial neural network for predicting residual neuromuscular block.

BACKGROUND: Postoperative residual curarization (PORC) after surgery is common and its detection has a high error rate. Artificial neural networks are being used increasingly to examine complex data. We hypothesized that a neural network would enhance prediction of PORC. METHODS: In 40 previously reported patients, neuromuscular function, neuromuscular block/antagonist usage and time intervals were recorded throughout anaesthesia until tracheal extubation by an observer uninvolved in patient care. PORC was defined as significant 'fade' (train of four <0.7) at extubation. Neuromuscular function was classified as PORC (value=1) or no PORC (value=0). A back-propagation neural network was trained to assign similar values (0, 1) for prediction of PORC, by examining the impact of (i) the degree of spontaneous recovery at reversal, and (ii) the time since pharmacological reversal, using the jackknife method. Successful prediction was defined as attainment of a predicted value within 0.2 of the target value. RESULTS: Twenty-six patients (65%) had PORC at tracheal extubation. Clinical detection of PORC had a sensitivity of 0 and specificity of 1, with an indeterminate positive predictive value and a negative predictive value of 0.35. Using the artificial neural network, one patient with residual block and one with adequate neuromuscular function were incorrectly classified during the test phase, with no indeterminate predictions, giving an artificial neural network sensitivity of 0.96 (chi(2)=44, P<0.001) and specificity of 0.92 (P=1), with a positive predictive value of 0.96 and a negative predictive value of 0.93 (chi(2)=12, P<0.001). CONCLUSIONS: Neural network-based prediction, using readily available clinical measurements, is significantly better than human judgement in predicting recovery of neuromuscular function.

Adult↗

Multivariate prediction of in-hospital mortality associated with surgical procedures.

BACKGROUND: The aims of this prospective multicenter study were to identify variables associated with in-hospital mortality among patients undergoing surgical procedures, to develop a prediction rule, and to statistically validate its reliability. METHODS: Data from 24,654 consecutive informed patients over 15 years of age were collected from 22 surgical centers between January 1989 and December 1990. Using logistic regression analysis separate models were fit for seven surgical disciplines to predict the risk of 30-day in hospital mortality. Variables used to construct the regression models included age, sex, systolic blood pressure, renal dysfunction, hepatic dysfunction, concomitant diseases, severity of surgery, priority of surgery and duration of anesthesia. The performance of the prediction rule was evaluated by computing sensitivity, specificity and predictive values, analyzing the ROC curve and comparing observed with expected deaths. RESULTS: The significance of the independent variables varied within each model. All models significantly predicted the occurrence of in-hospital mortality. At a 0.5 cuptoint of predicted risk sensitivity of prediction rule was 99.89%, positive predictive value 98.51%, and overall predictive value 98.41%, whereas specificity was 7.92% and negative value slightly higher than 50%. The area under the ROC curve was 0.80 (perfect, 1.0). The correlation between observed and expected deaths was 0.99. CONCLUSION: This prediction rule, developed using multicenter data, is characterized by the following advantages: includes only nine variables; can be utilized by seven different surgical disciplines; is highly accurate, and is easily available to clinicals with access to a microcomputer or programmable calculator. This validated multivariate prediction rule would be useful both to calculate the risk of mortality for an individual surgical patient and to contrast observed and expected mortality rates for an institution or a particular clinician.

Adolescent↗

Nonpalpable stage T1c prostate cancer: prediction of insignificant disease using free/total prostate specific antigen levels and needle biopsy findings.

PURPOSE: Approximately 25% of radical prostatectomies performed for stage T1c disease show potentially insignificant prostate cancer. We previously reported the use of serum prostate specific antigen (PSA) density and needle biopsy findings to predict potentially insignificant cancer. We now evaluate whether using free/total serum PSA levels along with needle biopsy findings can better predict tumor significance. MATERIALS AND METHODS: We studied 163 radical prostatectomy specimens of stage T1c prostate cancer in which free/total serum PSA levels were determined. Free/total serum PSA levels were measured with Tandem-MP PSA assays. Insignificant prostate cancers were organ confined with tumor volumes less than 0.5 cc and Gleason score less than 7. Advanced tumors were either Gleason score 7 or greater, established extraprostatic extension with positive margins, or positive seminal vesicles or lymph nodes. Other cases were considered as moderate tumor. Moderate and advanced tumors were considered significant. RESULTS: Of the tumors 30.7% were insignificant, 49.7% moderate and 19.6% advanced. The best model to predict preoperatively insignificant tumor was a free/total PSA of 0.15 or greater and favorable needle biopsy findings (less than 3 cores involved, none of the cores with greater than 50% tumor involvement and Gleason score less than 7). Using this model of the 18 tumors predicted to be insignificant 17 were insignificant for a positive predictive value of 94.4%. Of the 145 cases that were predicted to be significant 112 were correctly predicted for a negative predictive value of 77.2%. There was only 1 tumor predicted to be insignificant which was classified as moderate. No tumor predicted to be insignificant was advanced. CONCLUSIONS: In conjunction with needle biopsy findings, free/total PSA levels accurately predict insignificant tumor in stage T1c disease.

Biopsy, Needle↗

A comparison of four severity-adjusted models to predict mortality after coronary artery bypass graft surgery.

OBJECTIVE: To assess the validity of four severity-adjusted models to predict mortality following coronary artery bypass graft surgery by using an independent surgical database. DESIGN: A prospective observational study wherein predicted mortality for each patient was obtained by using four different published severity-adjusted models. SETTING: A university-affiliated teaching community hospital. PATIENTS: Eight hundred sixty-eight consecutive patients who underwent coronary artery bypass graft surgery without accompanying valve or aneurysm repair during the period from 1991 to 1993. INTERVENTIONS: None. MAIN OUTCOME MEASURES: Predicted mortality rates for each model were obtained by averaging individual patient predictions and were compared with actual morality rates. We assessed the accuracy of overall prediction for the total series, as well as compared individual patient predictions created by each model. The discrimination of models was assessed with receiver operating characteristic curves and the Hosmer-Lemeshow goodness-of-fit statistic. RESULTS: The observed crude mortality rate was 3.7%. The predicted mortality rate ranged from 2.8% to 9.2%, despite relatively good discrimination by the models (area under the receiver operating characteristic curve, 0.70 to 0.74). The individual patient mortality predicted by different models varied by as much as a ninefold difference. CONCLUSIONS: The currently used coronary artery bypass graft predictive models, although generally accurate, have significant shortcomings and should be used with caution. The predicted mortality rate following coronary artery bypass graft surgery varied by a factor of 3.3 from lowest to highest, making the choice of model a critical factor when assessing outcome. The use of these models for individual patient risk estimations is risky because of the marked discrepancies in individual predictions created by each model.

Aged↗

Improvement of side-chain modeling in proteins with the self-consistent mean field theory method based on an analysis of the factors influencing prediction.

With the objective of improving side-chain conformation prediction, we have analyzed the influence of various factors on prediction by the Self-Consistent Mean Field Theory method, applied to a set of high resolution x-ray protein structure models. These factors may be classed as variations in the mean field optimization protocol, variations in the potential energy function, and variations in rotamer library completeness. We have developed an optimization protocol that consistently reached lower mean field conformational free energies than two other protocols. This protocol led to an important improvement in prediction. We observed a major improvement in prediction with two more detailed van der Waals parameter sets, which we found to be due mainly to the introduction of scaling of 1-4 interactions. In a comparison of two knowledge-based rotamer libraries of considerably different size, we observed an unexpected decrease in prediction with an increase in library completeness. However, when we introduced a torsion potential term in the potential energy function, we found an important increase in average prediction and in the prediction of almost all residue types with a more complete rotamer set. The two knowledge-based rotamer libraries now became equivalent in terms of average prediction. The results we obtained in an analysis of the effect of the introduction of an additional electrostatic term in the potential energy function were largely inconclusive. However, we found a small increase in average prediction for an electrostatic potential term with a fixed dielectric constant of 15. The combined effect of all the factors we analyzed in this study resulted in average prediction accuracies of 79.9% for X1, 68.1% for X1 + 2, and 1.590 A for global rms deviation (RMSD); the corresponding values for core residues were 88.2%, 78.6%, and 1.171 A. These values represent improvements in average prediction of 6.5% for X1, 9.1% for X1 + 2, and 0.163 A for global RMSD over the original conditions; the corresponding improvements in the core were 5.9%, 9.0%, and 0.180 A, respectively.

Amino Acids↗

A structural alphabet for local protein structures: improved prediction methods.

Three-dimensional protein structures can be described with a library of 3D fragments that define a structural alphabet. We have previously proposed such an alphabet, composed of 16 patterns of five consecutive amino acids, called Protein Blocks (PBs). These PBs have been used to describe protein backbones and to predict local structures from protein sequences. The Q16 prediction rate reaches 40.7% with an optimization procedure. This article examines two aspects of PBs. First, we determine the effect of the enlargement of databanks on their definition. The results show that the geometrical features of the different PBs are preserved (local RMSD value equal to 0.41 A on average) and sequence-structure specificities reinforced when databanks are enlarged. Second, we improve the methods for optimizing PB predictions from sequences, revisiting the optimization procedure and exploring different local prediction strategies. Use of a statistical optimization procedure for the sequence-local structure relation improves prediction accuracy by 8% (Q16 = 48.7%). Better recognition of repetitive structures occurs without losing the prediction efficiency of the other local folds. Adding secondary structure prediction improved the accuracy of Q16 by only 1%. An entropy index (Neq), strongly related to the RMSD value of the difference between predicted PBs and true local structures, is proposed to estimate prediction quality. The Neq is linearly correlated with the Q16 prediction rate distributions, computed for a large set of proteins. An "expected" prediction rate QE16 is deduced with a mean error of 5%.

Amino Acid Sequence↗

Prediction of the conformation and geometry of loops in globular proteins: testing ArchDB, a structural classification of loops.

In protein structure prediction, a central problem is defining the structure of a loop connecting 2 secondary structures. This problem frequently occurs in homology modeling, fold recognition, and in several strategies in ab initio structure prediction. In our previous work, we developed a classification database of structural motifs, ArchDB. The database contains 12,665 clustered loops in 451 structural classes with information about phi-psi angles in the loops and 1492 structural subclasses with the relative locations of the bracing secondary structures. Here we evaluate the extent to which sequence information in the loop database can be used to predict loop structure. Two sequence profiles were used, a HMM profile and a PSSM derived from PSI-BLAST. A jack-knife test was made removing homologous loops using SCOP superfamily definition and predicting afterwards against recalculated profiles that only take into account the sequence information. Two scenarios were considered: (1) prediction of structural class with application in comparative modeling and (2) prediction of structural subclass with application in fold recognition and ab initio. For the first scenario, structural class prediction was made directly over loops with X-ray secondary structure assignment, and if we consider the top 20 classes out of 451 possible classes, the best accuracy of prediction is 78.5%. In the second scenario, structural subclass prediction was made over loops using PSI-PRED (Jones, J Mol Biol 1999;292:195-202) secondary structure prediction to define loop boundaries, and if we take into account the top 20 subclasses out of 1492, the best accuracy is 46.7%. Accuracy of loop prediction was also evaluated by means of RMSD calculations.

Models, Molecular↗

Evaluation of different biological data and computational classification methods for use in protein interaction prediction.

Protein-protein interactions play a key role in many biological systems. High-throughput methods can directly detect the set of interacting proteins in yeast, but the results are often incomplete and exhibit high false-positive and false-negative rates. Recently, many different research groups independently suggested using supervised learning methods to integrate direct and indirect biological data sources for the protein interaction prediction task. However, the data sources, approaches, and implementations varied. Furthermore, the protein interaction prediction task itself can be subdivided into prediction of (1) physical interaction, (2) co-complex relationship, and (3) pathway co-membership. To investigate systematically the utility of different data sources and the way the data is encoded as features for predicting each of these types of protein interactions, we assembled a large set of biological features and varied their encoding for use in each of the three prediction tasks. Six different classifiers were used to assess the accuracy in predicting interactions, Random Forest (RF), RF similarity-based k-Nearest-Neighbor, Naïve Bayes, Decision Tree, Logistic Regression, and Support Vector Machine. For all classifiers, the three prediction tasks had different success rates, and co-complex prediction appears to be an easier task than the other two. Independently of prediction task, however, the RF classifier consistently ranked as one of the top two classifiers for all combinations of feature sets. Therefore, we used this classifier to study the importance of different biological datasets. First, we used the splitting function of the RF tree structure, the Gini index, to estimate feature importance. Second, we determined classification accuracy when only the top-ranking features were used as an input in the classifier. We find that the importance of different features depends on the specific prediction task and the way they are encoded. Strikingly, gene expression is consistently the most important feature for all three prediction tasks, while the protein interactions identified using the yeast-2-hybrid system were not among the top-ranking features under any condition.

Computational Biology↗

Prediction of side-chain conformations on protein surfaces.

An approach is described that improves the prediction of the conformations of surface side chains in crystal structures, given the main-chain conformation of a protein. A key element of the methodology involves the use of the colony energy. This phenomenological term favors conformations found in frequently sampled regions, thereby approximating entropic effects and serving to smooth the potential energy surface. Use of the colony energy significantly improves prediction accuracy for surface side chains with little additional computational cost. Prediction accuracy was quantified as the percentage of side-chain dihedral angles predicted to be within 40 degrees of the angles measured by X-ray diffraction. Use of the colony energy in predictions for single side chains improved the prediction accuracy for chi(1) and chi(1+2) from 65 and 40% to 74 and 59%, respectively. Several other factors that affect prediction of surface side-chain conformations were also analyzed, including the extent of conformational sampling, details of the rotamer library employed, and accounting for the crystallographic environment. The prediction of conformations for polar residues on the surface was generally found to be more difficult than those for hydrophobic residues, except for polar residues participating in hydrogen bonds with other protein groups. For surface residues with hydrogen-bonded side chains, the prediction accuracy of chi(1) and chi(1+2) was 79 and 63%, respectively. For surface polar residues, in general (all side-chain prediction), the accuracy of chi(1) and chi(1+2) was only 73 and 56%, respectively. The most accurate results were obtained using the colony energy and an all-atom description that includes neighboring molecules in the crystal (protein chains and hetero atoms). Here, the accuracy of chi(1) and chi(1+2) predictions for surface side chains was 82 and 73%, respectively. The root mean square deviations obtained for hydrogen-bonding surface side chains were 1.64 and 1.81 A, with and without consideration of crystal packing effects, respectively.

Crystallography, X-Ray↗