Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Review: protein secondary structure prediction continues to rise.

Methods predicting protein secondary structure improved substantially in the 1990s through the use of evolutionary information taken from the divergence of proteins in the same structural family. Recently, the evolutionary information resulting from improved searches and larger databases has again boosted prediction accuracy by more than four percentage points to its current height of around 76% of all residues predicted correctly in one of the three states, helix, strand, and other. The past year also brought successful new concepts to the field. These new methods may be particularly interesting in light of the improvements achieved through simple combining of existing methods. Divergent evolutionary profiles contain enough information not only to substantially improve prediction accuracy, but also to correctly predict long stretches of identical residues observed in alternative secondary structure states depending on nonlocal conditions. An example is a method automatically identifying structural switches and thus finding a remarkable connection between predicted secondary structure and aspects of function. Secondary structure predictions are increasingly becoming the work horse for numerous methods aimed at predicting protein structure and function. Is the recent increase in accuracy significant enough to make predictions even more useful? Because the recent improvement yields a better prediction of segments, and in particular of beta strands, I believe the answer is affirmative. What is the limit of prediction accuracy? We shall see.

Computational Biology↗

A modular concept of HLA for comprehensive peptide binding prediction.

A variety of algorithms have been successful in predicting human leukocyte antigen (HLA)-peptide binding for HLA variants for which plentiful experimental binding data exist. Although predicting binding for only the most common HLA variants may provide sufficient population coverage for vaccine design, successful prediction for as many HLA variants as possible is necessary to understand the immune response in transplantation and immunotherapy. However, the high cost of obtaining peptide binding data limits the acquisition of binding data. Therefore, a prediction algorithm, which applies the binding information from well-studied HLA variants to HLA variants, for which no peptide data exist, is necessary. To this end, a modular concept of class I HLA-peptide binding prediction was developed. Accurate predictions were made for several alleles without using experimental peptide binding data specific to those alleles. We include a comparison of module-based prediction and supertype-based prediction. The modular concept increased the number of predictable alleles from 15 to 75 of HLA-A and 12 to 36 of HLA-B proteins. Under the modular concept, binding data of certain HLA alleles can make prediction possible for numerous additional alleles. We report here a ranking of HLA alleles, which have been identified to be the most informative. Modular peptide binding prediction is freely available to researchers on the web at http://www.peptidecheck.org .

Alleles↗

Lean body mass-based standardized uptake value, derived from a predictive equation, might be misleading in PET studies.

The standardized uptake value (SUV) has gained recognition in recent years as a semiquantitative evaluation parameter in positron emission tomography (PET) studies. However, there is as yet no consensus on the way in which this index should be determined. One of the confusing factors is the normalisation procedure. Among the proposed anthropometric parameters for normalisation is lean body mass (LBM); LBM has been determined by using a predictive equation in most if not all of the studies. In the present study, we assessed the degree of agreement of various LBM predictive equations with a reference method. Secondly, we evaluated the impact of predicted LBM values on a hypothetical value of 2.5 SUV, normalised to LBM (SUV(LBM)), by using various equations. The study population consisted of 153 women, aged 32.3+/-11.8 years (mean+/-SD), with a height of 1.61+/-0.06 m, a weight of 71.1+/-17.5 kg, a body surface area of 1.77+/-0.22 m(2) and a body mass index of 27.6+/-6.9 kg/m(2). LBM (44.2+/-6.6 kg) was measured by a dual-energy X-ray absorptiometry (DEXA) method. A total of nine equations from the literature were evaluated, four of them from recent PET studies. Although there was significant correlation between predicted and measured LBM values, 95% limits of agreement determined by the Bland and Altman method showed a wide range of variation in predicted LBM values as compared with DEXA, no matter which predictive equation was used. Moreover, only one predictive equation was not statistically different in the comparison of means (DEXA and predicted LBM values). It was also shown that the predictive equations used in this study yield a wide range of SUV(LBM) values from 1.78 to 5.16 (29% less or 107% more) for an SUV of 2.5. In conclusion, this study suggests that estimation of LBM by use of a predictive equation may cause substantial error for an individual, and that if LBM is chosen for the SUV normalisation procedure, it should be measured, not predicted.

Absorptiometry, Photon↗

An assessment of protein secondary structure prediction methods based on amino acid sequence.

Five of the several secondary structure prediction methods based on protein amino acid sequence has been computerized, allowing the calculation of joint prediction histograms which have been shown to be superior to any individual prediction. The known structures of about 40 proteins experimentally determined by X-ray crystallography are compared with the predictions resulting from calculated histograms. The accuracy of the predictions for helices is generally much better than for both beta-sheet regions and for turns. The overall agreement between prediction and observation within the amino terminal half of the protein molecules is clearly superior to that for the carboxyl half, suggesting an amino nucleating core. Predictions for smaller proteins and thermally stable proteins are generally good, indicating the sensitivity of the methods to short-range but not long-range interactions. In less than half the cases tested were the predictions useful; there was no way of knowing ahead of time if a favorable prediction would result. Given the lack of dramatic improvement with an increase in data base for the schemes and the generally poor agreement factors, it appears that a perfect predictive algorithm must include a consideration of energy minimization, thermalization, and long-range interactions. Extreme caution is suggested in applying present prediction routines to unknown protein structures.

Amino Acid Sequence↗

Protein structure prediction.

Current methods developed for predicting protein structure are reviewed. The most widely used algorithms of Chou and Fasman and Garnier et al for predicting secondary structure are compared to the most recent ones including sequence similarity methods, neural network, pattern recognition or joint prediction methods. The best of these methods correctly predict 63-65% of the residues in the database with cross-validation for 3 conformations, helix, beta strand and coli with a standard deviation of 6-8% per protein. However, when a homologous protein is already in the database, the accuracy of prediction by the similarity peptide method of Levin and Garnier reaches about 90%. Some conclusions can be drawn on the mechanism of protein folding. As all the prediction methods only use the local sequence for prediction (+/- 8 residues maximum) one can infer that 65% of the conformation of a residue is dictated on average by the local sequence, the rest is brought by the folding. The best predicted proteins or peptide segments are those for which the folding has less effect on the conformation. Presently, prediction of tertiary structure is only of practical use when the structure of a homologous protein is already known. Amino acid alignment to define residues of equivalent spatial position is critical for modelling of the protein. We showed for serine proteases that secondary structure prediction can help to define a better alignment. Non-homologous segments of the polypeptide chain, such as loops, libraries of known loops and/or energy minimization with various force fields, are used without yet giving satisfactory solutions. An example of modelling by homology, aided by secondary structure prediction on 2 regulatory proteins, Fnr and FixK is presented.

Algorithms↗

Can restenosis after coronary angioplasty be predicted from clinical variables?

OBJECTIVES: The purpose of this study was to determine whether variables shown to correlate with restenosis in one group (learning group) could be shown to predict recurrent stenosis in a second group (validation group). BACKGROUND: Restenosis remains a critical limitation after percutaneous transluminal coronary angioplasty. Although several clinical variables have been shown to correlate with restenosis, there are few data concerning attempts to predict recurrent stenosis. METHODS: The source of data was the clinical data base at Emory University. Patients who had had previous coronary surgery and patients who underwent coronary angioplasty in the setting of acute myocardial infarction were excluded. A total of 4,006 patients with angiographic restudy after successful angioplasty were identified. They were classified into a learning group of 2,500 patients and a validation group of 1,506 patients. The correlates of restenosis in the learning group were determined by stepwise logistic regression, and a model was developed to predict the probability of restenosis and was tested in the validation group. By using various cut points for the predicted probability of restenosis, a receiver operating characteristic curve was created. Goodness of fit of the model was evaluated by comparing average predicted probabilities with average observed probabilities within subgroups on the basis of risk level determined by linear regression analysis. RESULTS: In the learning group 1,145 patients had restenosis and 1,355 did not. Correlates of restenosis were severe angina, severe diameter stenosis before angioplasty, left anterior descending coronary artery dilation, diabetes, greater diameter stenosis after angioplasty, hypertension, absence of an intimal tear, eccentric morphology and older patient age. The model derived from the learning group was used to predict restenosis in the validation group. By varying the cut point for the predicted probability of restenosis above which restenosis is diagnosed and below which it is not, a receiver operating characteristic curve was created. The curve was close to the line of identity, reflecting a poor predictive ability. However, the model was shown to fit well with the predicted probability of restenosis correlating well with the observed probability (r = 0.98, p = 0.0001). CONCLUSIONS: Clinical variables provide limited ability to predict definitively whether a particular patient will have restenosis. However, the current model may be used to predict the probability of restenosis, with some uncertainty, at least in well characterized patients who have already had angioplasty.

Angioplasty, Balloon, Coronary↗

Estimation of significant solvent concentration ranges and its application to the enhancement of the accuracy of gradient predictions.

The solvent concentration range actually useful for gradient predictions is significantly narrower than the total range scanned in a gradient run. This range, called "solvent informative range" (SIR), if known with the highest accuracy, allows to predict gradient retention times (t(g) with minimal error. The small size of the SIR supports the application of the linear solvent strength theory (LSST). Furthermore, LSST allows a closed-form solution to the integral required to predict gradient retention times, which eliminates numerical integration, needed with other retention models. A methodology that calculates the SIR by applying error analysis, and uses it to improve the accuracy in the prediction of t(g) from isocratic experiments, is proposed. The importance of those mobile-phase compositions that do not contribute significantly to the prediction of t(g) is selectively attenuated within the prediction algorithm, relying the predictions more heavily on the SIR. As a result, t(g) was found to be predicted with similar accuracy using isocratic training data with regard to predictions based on gradient training data. The approach is useful for all situations where the chromatographer is able to provide predictions of retention at constant solvent concentration, and wish to predict the retention in gradient mode.

Algorithms↗

Computational wear prediction of a total knee replacement from in vivo kinematics.

Wear of ultra-high molecular weight polyethylene bearings in total knee replacements remains a major limitation to the longevity of these clinically successful devices. Few design tools are currently available to predict mild wear in implants based on varying kinematics, loads, and material properties. This paper reports the implementation of a computer modeling approach that uses fluoroscopically measured motions as inputs and predicts patient-specific implant damage using computationally efficient dynamic contact and tribological analyses. Multibody dynamic simulations of two activities (gait and stair) with two loading conditions (70-30 and 50-50 medial-lateral load splits) were generated from fluoroscopic data to predict contact pressure and slip velocity time histories for individual elements on the tibial insert surface. These time histories were used in a computational wear analysis to predict the depth of damage due to wear and creep experienced by each element. Predicted damage areas, volumes, and maximum depths were evaluated against a tibial insert retrieved from the same patient who provided the in vivo motions. Overall, the predicted damage was in close agreement with damage observed on the retrieval. The gait and stair simulations separately predicted the correct location of maximum damage on the lateral side, whereas a combination of gait and stair was required to predict the correct location on the medial side. Predicted maximum damage depths were consistent with the retrieval as well. Total computation time for each damage prediction was less than 30 min. Continuing refinement of this approach will provide a robust tool for accurately predicting clinically relevant wear in total knee replacements.

Aged↗

Prediction error for free monetary reward in the human prefrontal cortex.

Making predictions about future rewards is an important ability for primates, and its neurophysiological mechanisms have been studied extensively. One important approach is to identify neural systems that process errors related to reward prediction (i.e., areas that register the occurrence of unpredicted rewards and the failure of expected rewards). In monkeys that have learned to predict appetitive rewards during reward-directed behaviors, dopamine neurons reliably signal both types of prediction error. The mechanisms in the human brain involved in processing prediction error for monetary rewards are not well understood. Furthermore, nothing is known of how such systems operate when rewards are not contingent on behavior. We used event-related fMRI to localize responses to both classes of prediction error. Subjects were able to predict a monetary reward or a nonreward on the basis of a prior visual cue. On occasional trials, cue-outcome contingencies were reversed (unpredicted rewards and failure of expected rewards). Subjects were not required to make decisions or actions. We compared each type of prediction error trial with its corresponding control trial in which the same prediction did not fail. Each type of prediction error evoked activity in a distinct frontotemporal circuit. Unexpected reward failure evoked activity in the temporal cortex and frontal pole (area 10). Unpredicted rewards evoked activity in the orbitofrontal cortex, the frontal pole, parahippocampal cortex, and cerebellum. Activity time-locked to prediction errors in frontotemporal circuits suggests that they are involved in encoding the associations between visual cues and monetary rewards in the human brain.

Basal Ganglia↗

Enhancing spatial estimates of metal pollutants in raw wastewater irrigated fields using a topsoil organic carbon map predicted from aerial photography.

Various approaches have been used to estimate metal pollutant element (TE) contents at unsampled locations in a 15-ha contaminated site located in the plain of Pierrelaye-Bessancourt (about 24 km Northwest of Paris). 87 samples of soil plough layer were randomly sampled at each mesh of a regular square grid over the whole study area and the total contents of Cd, Cr, Cu, Ni, Pb, and Zn were measured. A first set of 50 measurements, randomly selected from the 87 samples, was used for the prediction and another set of 37 measurements was kept for the validation. Topsoil organic carbon contents (SOC) were measured at 75 sites with 50 measurements sharing the same locations as TE. An aerial photography of the study area showing bare soils was selected for relating brightness intensities and SOC. Mapping procedures used were ordinary kriging (OK), cokriging (COK), collocated cokriging (CC), and kriging with external drift (KED). SOC maps used as exhaustively sampled information in KED and CC of TE were obtained by KED and CC procedures, respectively, accounting for 75 SOC measurements and the brightness intensities of numerical counts provided by the visible bands of the aerial photograph bare soils. Consequently, for each TE, four maps were generated: two maps resulting from KED and CC procedures (KED-SOC75P, CC-SOC75P), another one provided by standard cokriging (COK-TE50SOC75) accounting for TE prediction set plus 75 SOC measurements, and the last one corresponding to that estimated by ordinary kriging from only prediction set measurements (OK50). Three indices: (1) the mean prediction error (ME) and the mean absolute prediction error (|ME|); (2) the root mean square error (RMSE); and (3) the relative improvement (RI) of accuracy, as well as residuals analysis, were computed from the validation set (observed data) and predicted values. On the 37 test data, the results showed that the more accurate predictions were systematically those obtained by kriging accounting for SOC map predicted by KED from 75 SOC measurements and brightness values of the aerial photo (KED-SOC75P) followed closely by CC-SOC75P procedure, except for Cu and Zn where CC-SOC75P appeared to be slightly more accurate than KED-SOC75P. In regard to the RI of accuracy between prediction methods, the results confirmed once for all the benefit of accounting for SOC data set plus the exhaustively sampled information provided by the aerial photography regardless of the considered TE. Nevertheless, for Cd, Pb, and Zn, the RI of accuracy was less than 20% between the two most accurate methods (KED-SOC75P and CC-SOC75P) and standard cokriging in which the information provided by the aerial photography is ignored when mapping. The sensitivity of KED-SOC75P and CC-SOC75P approaches to the sampling density of the target variables (TE) was assessed using 10 random subsets of different sizes (25 and 33 observations) drawn from a prediction set that includes 50 data. Results have shown that the TE estimates by KED-SOC75P and CC-SOC75P approaches using only 25 TE samples were much more accurate than the estimates performed by OK50 and COK-TE50SOC75 approaches that use the whole samples of the prediction set. Moreover, the RI of accuracy was reduced by less than 15% if the original sampling density was reduced by a third.

Agriculture↗

Models to predict emissions of health-damaging pollutants and global warming contributions of residential fuel/stove combinations in China.

Residential energy use in developing countries has traditionally been associated with combustion devices of poor energy efficiency, which have been shown to produce substantial health-damaging pollution, contributing significantly to the global burden of disease, and greenhouse gas (GHG) emissions. Precision of these estimates in China has been hampered by limited data on stove use and fuel consumption in residences. In addition limited information is available on variability of emissions of pollutants from different stove/fuel combinations in typical use, as measurement of emission factors requires measurement of multiple chemical species in complex burn cycle tests. Such measurements are too costly and time consuming for application in conjunction with national surveys. Emissions of most of the major health-damaging pollutants (HDP) and many of the gases that contribute to GHG emissions from cooking stoves are the result of the significant portion of fuel carbon that is diverted to products of incomplete combustion (PIC) as a result of poor combustion efficiencies. The approximately linear increase in emissions of PIC with decreasing combustion efficiencies allows development of linear models to predict emissions of GHG and HDP intrinsically linked to CO2 and PIC production, and ultimately allows the prediction of global warming contributions from residential stove emissions. A comprehensive emissions database of three burn cycles of 23 typical fuel/stove combinations tested in a simulated village house in China has been used to develop models to predict emissions of HDP and global warming commitment (GWC) from cooking stoves in China, that rely on simple survey information on stove and fuel use that may be incorporated into national surveys. Stepwise regression models predicted 66% of the variance in global warming commitment (CO2, CO, CH4, NOx, TNMHC) per 1 MJ delivered energy due to emissions from these stoves if survey information on fuel type was available. Subsequently if stove type is known, stepwise regression models predicted 73% of the variance. Integrated assessment of policies to change stove or fuel type requires that implications for environmental impacts, energy efficiency, global warming and human exposures to HDP emissions can be evaluated. Frequently, this involves measurement of TSP or CO as the major HDPs. Incorporation of this information into models to predict GWC predicted 79% and 78% of the variance respectively. Clearly, however, the complexity of making multiple measurements in conjunction with a national survey would be both expensive and time consuming. Thus, models to predict HDP using simple survey information, and with measurement of either CO/CO2 or TSP/CO2 to predict emission factors for the other HDP have been derived. Stepwise regression models predicted 65% of the variance in emissions of total suspended particulate as grams of carbon (TSPC) per 1 MJ delivered if survey information on fuel and stove type was available and 74% if the CO/CO2 ratio was measured. Similarly stepwise regression models predicted 76% of the variance in COC emissions per MJ delivered with survey information on stove and fuel type and 85% if the TSPC/CO2 ratio was measured. Ultimately, with international agreements on emissions trading frameworks, similar models based on extensive databases of the fate of fuel carbon during combustion from representative household stoves would provide a mechanism for computing greenhouse credits in the residential sector as part of clean development mechanism frameworks and monitoring compliance to control regimes.

Air Pollutants↗

The predictability of corneal flap thickness and tissue laser ablation in laser in situ keratomileusis.

OBJECTIVE: To evaluate the relationship between predicted flap thickness and actual flap thickness and between predicted tissue ablation and actual tissue ablation. DESIGN: Prospective, nonrandomized comparative (self-controlled) trial. PARTICIPANTS: A total of 60 patients (102 eyes) who underwent laser in situ keratomileusis (LASIK). MAIN OUTCOME MEASURES: Subtraction pachymetry was used to determine actual corneal flap thickness and corneal tissue ablation depth. Other measurements included flap diameter and keratometry readings. RESULTS: Actual flap thickness was significantly different (P < 0.0001) from predicted flap thickness. Fifteen eyes had a predicted flap thickness of 160 micrometer and a mean actual flap of 105 micrometer (standard deviation [SD], +/-24. 3 micrometer range, 48-141 micrometer). Sixty-four had a predicted flap of 180 micrometer with an actual flap mean of 125 micrometer (SD, +/-18.5 micrometer range, 82-155 micrometer). Seventeen eyes had a predicted flap of 200 micrometer, with an actual flap mean of 144 micrometer (SD, +/-19.3 micrometer range, 108-187 micrometer). In addition, we found that significantly more tissue (P < 0.0001) was ablated than predicted. Linear regression of the observed ablation on predicted ablation yielded the following relationship: actual ablation = 14.5 + 1.5 (predicted ablation). Neither flap diameter nor flap thickness were found to increase with respect to steeper corneal curvatures. CONCLUSIONS: Actual corneal flap thickness was consistently less than predicted regardless of the depth plate used; actual tissue ablation was consistently greater than predicted tissue ablation for the laser used in this study.

Cornea↗

Prediction of dynamic tendon forces from electromyographic signals: an artificial neural network approach.

Artificial neural networks (ANN) with a backpropagation algorithm were used to predict dynamic tendon forces from electromyographic (EMG) signals. To achieve this goal, tendon forces and EMG-signals were recorded simultaneously in the gastrocnemius muscle of three cats while walking and trotting at different speeds on a motor-driven treadmill. The quality of the tendon force predictions were evaluated for three levels of generalization. First, at the intrasession level, tendon force predictions were made for step cycles from the same experimental session as the step cycles which were used to train the ANN. At this level of generalization very good results were obtained. Second, at the intrasubject level, tendon force predictions were made for one cat walking at a given speed while the ANN was trained with data from the same animal walking at different speeds. For the intrasubject predictions, the quality of the results depended on the walking speed for which the predictions were made: for the speeds at the low and high extremes, the predictions were worse than for the intermediate speeds. The cross-correlation coefficients between predicted and actual force time histories ranged from 0.78 to 0.91. Third, at the intersubject level, tendon forces were predicted for one animal walking at a given speed while the ANN was trained with data from the remaining two animals walking at the corresponding speed. The cross-correlation coefficients between predicted and actual force time histories ranged from 0.72 to 0.98. It was concluded that the ANN-approach is a powerful technique to predict dynamic tendon forces from EMG-signals.

Animals↗

A future prediction type artificial heart system.

The demand of the biological system needs to be predicted to consider the quality of life (QOL) of a patient with an artificial heart system. The purpose of this study was the prediction of the imminent cardiac output and the predictive control for an artificial heart. For that purpose, autonomic nerve information was applied in this study. Nervous sympathicus action potentials were measured, and a prediction function of cardiac output was made using the sympathetic tone and preload and after-load measurement with multiple regression analysis. The predicted value showed significant correlation with the measured value after 2.9 s. Currently, however, long-term instrumentation of the nervous sympathicus potential is difficult. Thus, hemodynamic fluctuations, which recently have attracted attention, were used in this study. A prediction function using the Mayer wave, which represented nervous sympathicus, was determined. As a result, mid-term prediction became possible. Furthermore, a measurement of the vagal nerve was used as a possible long-term prediction parameter. For long-term prediction, Hurst exponent analysis was used in this study. Vagal nerve discharges in the changing position showed alteration of long-term determination. In conclusion, the future prediction control of an artificial heart takes shape using these prediction functions.

Action Potentials↗

Comparisons of nomograms and urologists' predictions in prostate cancer.

When applying nomograms to a clinical setting it is essential to know how their predictions compare with clinicians'. Comparisons exist outside of the prostate cancer literature. We reviewed these comparisons and conducted 2 experiments comparing predictions of clinicians with prostate cancer nomograms. By using Medline, we searched studies from January 1966 to July 1999 that compared human predictions with nomogram predictions. Next, we conducted 2 experiments: (1) 17 urologists were presented with 10 case vignettes and asked to predict the 5-year recurrence-free probabilities for each patient; (2) case presentations of 63 prostate cancer patients (including full clinical histories with complete diagnostic data and surgical findings) were made to a group of 25 clinicians who were asked to predict organ-confined disease. We found 22 published studies comparing human experts with nomograms, greater than half (13 of 22) showed the nomogram performing above the level of the human expert. Our first experiment showed urologist modification of 165 nomogram predictions led to a decrease in prediction accuracy (c-index decreased from.67 to.55, P <.05). In our second experiment, clinician predictions of organ-confined disease were comparable to the nomogram (area under the receiver operating characteristic curve [AUC] 0.78 and 0.79, respectively). A mixed-model suggests the nomogram did not augment clinician prediction accuracy (doctor excess error 1.4%, P =.75, 95% confidence interval [CI]: -10.9% to 8.2%). Our data suggest that nomograms do not seem to diminish predictive accuracy and they may be of significant benefit in certain clinical decision making settings.

Carcinoma↗

Adaptive prediction of internal target motion using external marker motion: a technical study.

An adaptive prediction approach was developed to infer internal target position by external marker positions. First, a prediction model (or adaptive neural network) is developed to infer target position from its former positions. For both internal target and external marker motion, two networks with the same type are created. Next, a linear model is established to correlate the prediction errors of both neural networks. Based on this, the prediction error of an internal target position can be reconstructed by the linear combination of the prediction errors of the external markers. Finally, the next position of the internal target is estimated by the network and subsequently corrected by the reconstructed prediction error. In a similar way, future positions are inferred as their previous positions are predicted and corrected. This method was examined by clinical data. The results demonstrated that an improvement (10% on average) of correlation between predicted signal and real internal motion was achieved, in comparison with the correlation between external markers and internal target motion. Based on the clinical data (with correlation coefficient 0.75 on average) observed between external marker and internal target motions, a prediction error (23% on average) of internal target position was achieved. The preliminary results indicated that this method is helpful to improve the predictability of internal target motion with the additional information of external marker signals. A consistent correlation between external and internal signals is important for prediction accuracy.

Algorithms↗

A paradigm for class prediction using gene expression profiles.

We propose a general framework for prediction of predefined tumor classes using gene expression profiles from microarray experiments. The framework consists of 1) evaluating the appropriateness of class prediction for the given data set, 2) selecting the prediction method, 3) performing cross-validated class prediction, and 4) assessing the significance of prediction results by permutation testing. We describe an application of the prediction paradigm to gene expression profiles from human breast cancers, with specimens classified as positive or negative for BRCA1 mutations and also for BRCA2 mutations. In both cases, the accuracy of class prediction was statistically significant when compared to the accuracy of prediction expected by chance. The framework proposed here for the application of class prediction is designed to reduce the occurrence of spurious findings, a legitimate concern for high-dimensional microarray data. The prediction paradigm will serve as a good framework for comparing different prediction methods and may accelerate the development of molecular classifiers that are clinically useful.

Algorithms↗

Linear regression models for solvent accessibility prediction in proteins.

The relative solvent accessibility (RSA) of an amino acid residue in a protein structure is a real number that represents the solvent exposed surface area of this residue in relative terms. The problem of predicting the RSA from the primary amino acid sequence can therefore be cast as a regression problem. Nevertheless, RSA prediction has so far typically been cast as a classification problem. Consequently, various machine learning techniques have been used within the classification framework to predict whether a given amino acid exceeds some (arbitrary) RSA threshold and would thus be predicted to be "exposed," as opposed to "buried." We have recently developed novel methods for RSA prediction using nonlinear regression techniques which provide accurate estimates of the real-valued RSA and outperform classification-based approaches with respect to commonly used two-class projections. However, while their performance seems to provide a significant improvement over previously published approaches, these Neural Network (NN) based methods are computationally expensive to train and involve several thousand parameters. In this work, we develop alternative regression models for RSA prediction which are computationally much less expensive, involve orders-of-magnitude fewer parameters, and are still competitive in terms of prediction quality. In particular, we investigate several regression models for RSA prediction using linear L1-support vector regression (SVR) approaches as well as standard linear least squares (LS) regression. Using rigorously derived validation sets of protein structures and extensive cross-validation analysis, we compare the performance of the SVR with that of LS regression and NN-based methods. In particular, we show that the flexibility of the SVR (as encoded by metaparameters such as the error insensitivity and the error penalization terms) can be very beneficial to optimize the prediction accuracy for buried residues. We conclude that the simple and computationally much more efficient linear SVR performs comparably to nonlinear models and thus can be used in order to facilitate further attempts to design more accurate RSA prediction methods, with applications to fold recognition and de novo protein structure prediction methods.

Amino Acids↗