Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Predictive accuracy of rural physicians' stated retention plans.

CONTEXT: The retention of rural physicians is a difficult phenomenon to study because job changes--the outcome of interest--take years to unfold. One common way to study retention is to ask rural practitioners through surveys how much longer they expect to remain in their current positions and use these statements of "anticipated retention" as an expedient proxy measure of actual retention. PURPOSE: To test the predictive accuracy of rural physicians' stated retention plans and test the hypotheses that predictions are more accurate for certain physicians, such as those with more experience, more control of their work situations, and at less risk for job burnout. METHODS: A 1991 mail survey (national stratified random sample) prospectively queried rural physicians' retention plans, and a follow-up survey 5 to 6 years later determined if and when respondents (N = 405, 67.5% combined response rate) had moved. FINDINGS: Retention predictions for the entire cohort corresponded remarkably well to the group's actual retention, with the proportion remaining each year deviating by only a few percentage points from what the group collectively expected. Predictions for individuals were also moderately accurate: 4 of 5 physicians who predicted remaining at least 5 years did so; 2 of 3 who predicted remaining less than 5 years indeed left before 5 years. Predictions of job changes in less than 2 years tended to be more accurate than predictions of 2 to 5 years. Physicians' predictions were more accurate when they worked in practices they owned (greater control) and were on-call 2 or fewer times each week (lower burnout risk). Accuracy was not greater with any of 5 measures of experience. CONCLUSIONS: Rural generalist physicians are moderately accurate when reporting how much longer they will remain in their jobs, validating the use of anticipated retention in rural health workforce studies.

Adult↗

Femoral neck trabecular patterns predict osteoporotic fractures.

In this paper we show that texture analysis of femoral neck trabecular patterns can be used to predict osteoporotic fractures. The study is based on a sample of 123 women aged 44-66 years with and without fractures. We analyzed trabecular patterns using the Co-occurrence Matrix texture analysis algorithm and compared the predictive utility of the textural data with densitometry. Logistic regression was used to estimate the predictive utility, exp(B), of clinical and textural data per standard deviation. Reproducibility was also demonstrated using paired films at 1-year intervals (CoV=4.5%). Bone mass estimated by DEXA measurements of the spine and hip were the most predictive of fractures giving a two-fold increase in fractures per s.d. bone mass loss (95% CI: 1.2-3.1, p<0.005). Age was also highly predictive with fracture risk increasing by 1.07-fold per year (95% CI: 1.01-1.14, p<0.02). Trabecular texture was found to give a lower, but significant, prediction of fracture of 1.5-fold per s.d. trabecular pattern loss (95% CI: 0.96-2.31, p<0.05). Combining age, weight, and trabecular texture increased the fracture prediction to 1.78-fold per s.d. (95% CI: 1.19-2.67). Combining trabecular texture with densitometry increased the predictive ability to 2.06-fold per s.d. (95% CI: 1.32-3.22) and combined with age and weight as well increased exp(B) to 2.1-fold per s.d. (95% CI: 1.32-3.35). This shows that osteoporotic trabecular texture changes can be "measured." Moreover, the combination of age, weight, and trabecular texture is more predictive than either alone. We propose therefore that this trabecular texture analysis is both reproducible and clinically meaningful. The application of such methods could be used to improve the estimation of fracture risk in conjunction with other clinical data, or where densitometry data cannot be obtained (e.g., in retrospective studies).

Absorptiometry, Photon↗

Neural models for predicting viral vaccine targets.

We applied artificial neural networks (ANN) for the prediction of targets of immune responses that are useful for study of vaccine formulations against viral infections. Using a novel data representation, we developed a system termed MULTIPRED that can predict peptide binding to multiple related human leukocyte antigens (HLA). This implementation showed high accuracy in the prediction of the promiscuous peptides that bind to five HLA-A2 allelic variants. MULTIPRED is useful for the identification of peptides that bind multiple HLA-A2 variants as a group. By implementing ANN as a classification engine, we enabled both the prediction of peptides binding to multiple individual HLA-A2 molecules and the prediction of promiscuous binders using a single model. The ANN MULTIPRED predicts peptide binding to HLA-A*0205 with excellent accuracy (area under the receiver operating characteristic curve--AROC>0.90), and to HLA-A*0201, HLA-A*0204 and HLA-A*0206 with high accuracy (AROC>0.85). Antigenic regions with high density of binders ("antigenic hot-spots") represent best targets for vaccine design. MULTIPRED not only predicts individual 9-mer binders but also predicts antigenic hot spots. Two HLA-A2 hot-spots in Severe Acute Respiratory Syndrome Coronavirus (SARS-CoV) membrane protein were predicted by using MULTIPRED.

Algorithms↗

Human neural learning depends on reward prediction errors in the blocking paradigm.

Learning occurs when an outcome deviates from expectation (prediction error). According to formal learning theory, the defining paradigm demonstrating the role of prediction errors in learning is the blocking test. Here, a novel stimulus is blocked from learning when it is associated with a fully predicted outcome, presumably because the occurrence of the outcome fails to produce a prediction error. We investigated the role of prediction errors in human reward-directed learning using a blocking paradigm and measured brain activation with functional magnetic resonance imaging. Participants showed blocking of behavioral learning with juice rewards as predicted by learning theory. The medial orbitofrontal cortex and the ventral putamen showed significantly lower responses to blocked, compared with nonblocked, reward-predicting stimuli. In reward-predicting control situations, deactivation in orbitofrontal cortex and ventral putamen occurred at the time of unpredicted reward omissions. Responses in discrete parts of orbitofrontal cortex correlated with the degree of behavioral learning during, and after, the learning phase. These data suggest that learning in primary reward structures in the human brain correlates with prediction errors in a manner that complies with principles of formal learning theory.

Adult↗

Neurons in the monkey superior colliculus predict the visual result of impending saccadic eye movements.

1. Previous experiments have shown that visual neurons in the lateral intraparietal area (LIP) respond predictively to stimuli outside their classical receptive fields when an impending saccade will bring those stimuli into their receptive fields. Because LIP projects strongly to the intermediate layers of the superior colliculus, we sought to demonstrate similar predictive responses in the monkey colliculus. 2. We studied the behavior of 90 visually responsive neurons in the superficial and intermediate layers of the superior colliculus of two rhesus monkeys (Macaca mulatta) when visual stimuli or the locations of remembered stimuli were brought into their receptive fields by a saccade. 3. Thirty percent (18/60) of intermediate layer visuomovement cells responded predictively before a saccade outside the movement field of the neuron when that saccade would bring the location of a stimulus into the receptive field. Each of these neurons did not respond to the stimulus unless an eye movement brought it into its receptive field, nor did it discharge in association with the eye movement unless it brought a stimulus into its receptive field. 4. These neurons were located in the deeper parts of the intermediate layers and had relatively larger receptive fields and movement fields than the cells at the top of the intermediate layers. 5. The predictive responses of most of these neurons (16/18, 89%) did not require that the stimulus be relevant to the monkey's rewarded behavior. However, for some neurons the predictive response was enhanced when the stimulus was the target of a subsequent saccade into the neuron's movement field. 6. Most neurons with predictive responses responded with a similar magnitude and latency to a continuous stimulus that remained on after the saccade, and to the same stimulus when it was only flashed for 50 ms coincident with the onset of the saccade target and thus never appeared within the cell's classical receptive field. 7. The visual response of neurons in the intermediate layers of the colliculus is suppressed during the saccade itself. Neurons that showed predictive responses began to discharge before the saccade, were suppressed during the saccade, and usually resumed discharging after the saccade. 8. Three neurons in the intermediate layers responded tonically from stimulus appearance to saccade without a presaccadic burst. These neurons responded predictively to a stimulus that was going to be the target for a second saccade, but not to an irrelevant flashed stimulus. 9. No superficial layer neuron (0/27) responded predictively when a stimulus would not be brought into their receptive fields by a saccade.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Influence of instruction, prediction, and afferent sensory information on the postural organization of step initiation.

1. Our previous study showed that two distinct postural modifications occurred when subjects were instructed to step, rather than maintain stance, in response to a backward surface translation: 1) the automatic postural responses to the surfaces perturbation were reduced in magnitude and 2) the anticipatory postural adjustments promoting foot-off were shortened in duration. This study investigates the extent to which task instruction, prediction of perturbation velocity, and afferent sensory information related to perturbation velocity are responsible for these postural modification. 2. Eleven human subjects were instructed in advance, to either maintain stance or step forward in response to a backward surface translation. Four different velocities of translation were used to perturb equilibrium. To assess the influence of predicted versus actual velocity information, the surface translations were presented in both a blocked order of increasing perturbation velocity (predictable) and a random order (unpredictable). Lower-extremity electromyographs (EMGs), ground reaction forces, and movement kinematics were quantified for both the automatic postural responses to perturbation and the anticipatory postural adjustments for step initiation. 3. The instruction to step was not solely responsible for the suppression of the automatic postural response. Prediction of perturbation velocity was required for significant suppression of the early automatic postural response when subjects stepped in response to the perturbation. When compared with the stance condition, the magnitude of the initial 50 ms of the automatic response in bilateral soleus and the left limb gastrocnemius (initial stance limb) was significantly reduced only when the perturbation velocities were presented in a blocked order. The magnitude of the automatic response was not reduced in the gastrocnemius of the right limb, which was always the initial swing limb and recruited for heel-off in the step conditions. This asymmetrical reduction of the gastrocnemius suggests that modification of the response was specific to the instruction, rather than a general decrease in the extensor muscle excitability. 4. The suppression of the early automatic postural response involved a change in the bias of the response. Despite the reduced magnitude during the predictable velocity step condition, the slope (i.e., gain) of the response with increasing velocities was not different from that of the stance condition. Thus the excitability of the automatic response was reduced by a relatively constant amount for each velocity when the perturbation velocity was predictable. 5. In contrast to the importance of velocity prediction for modification of the automatic postural response, actual velocity information was used for modification of the anticipatory postural adjustments when step was initiated in response to the surface perturbation. Regardless of whether the perturbation velocities were presented in a blocked or random order, the anticipatory postural adjustments were rapidly initiated and the duration of the postural adjustments for step initiation was shortened as the velocity of perturbation increased. 6. We conclude that the CNS uses prediction of perturbation velocity to modify the excitability of early automatic postural responses when the postural goal changes. In contrast, actual afferent velocity information can be used to modify the duration of the anticipatory postural adjustments for a voluntary step in response to perturbation. Thus the CNS utilizes feed-forward prediction to modify peripherally triggered postural responses, and utilizes immediate afferent information to modify the centrally initiated postural adjustments associated with voluntary movement.

Adaptation, Physiological↗

Gene expression profiling on lung cancer outcome prediction: present clinical value and future premise.

DNA microarray has been widely used in cancer research to better predict clinical outcomes and potentially improve patient management. The new approach provides accurate tumor classification and outcome predictions, such as tumor stage, metastatic status, and patient survival, and offers some hope for individualized medicine. However, growing evidence suggests that gene-based prediction is not stable and little is known about the prediction power of gene expression profiles compared with well-known clinical and pathologic predictors. This review summarized up-to-date publications in microarray-based lung cancer clinical outcome prediction and conducted secondary analyses for those with sufficient sample sizes and associated clinical information. Among the most commonly used analytic approaches, unsupervised clustering mainly recaptures tumor histology and provides variable degrees of prediction for tumor stage, lymph node status, or survival. Overall, most studies lack an independent validation. Supervised learning and testing generally offer a better prediction. Noted is that when conventional predictors of age, gender, stage, cell type, and tumor grade are considered collectively, the predictive advantage of the gene expression profiles diminishes. We conclude that outcome prediction from gene expression signatures selected by current analytic approaches can be mostly explained by well-known conventional predictors, particularly histologic subtype and grade of differentiation. A strategy for establishing independent or more accurate signatures is commented.

Cell Differentiation↗

Molecular prediction of response to 5-fluorouracil and interferon-alpha combination chemotherapy in advanced hepatocellular carcinoma.

PURPOSE: The prognosis of hepatocellular carcinoma (HCC) is very poor, particularly in patients with tumors that have invaded the major branches of the portal vein. Combination chemotherapy with intra-arterial 5-fluorouracil and subcutaneous interferon-alpha has shown promising results for such advanced HCC, but it is important to develop the ability to accurately predict chemotherapeutic responses. EXPERIMENTAL DESIGN: We analyzed the expression of 3,080 genes using a polymerase chain reaction-based array in 20 HCC patients who were treated with combination chemotherapy after reduction surgery. After unsupervised analyses, a supervised classification method for predicting chemotherapeutic responses was constructed. To minimize the number of predictive genes, we used a random permutation test to select only significant (P < 0.01) genes. A leave-one-out cross-validation confirmed the gene selection. We also prepared an additional 11 cases for validation of predictive performance. RESULTS: Hierarchical clustering analysis and principal component analysis with all 3,080 genes revealed distinct gene expression patterns in responders (those with complete response or partial response) and nonresponders (those with stable disease or progressive disease) to the combination chemotherapy. Using a weighted-voting classification method with either all genes or only significant genes as assessed by permutation testing, the objective responses to treatment were correctly predicted in 17 of 20 cases (accuracy, 85%; positive predictive value, 100%; negative predictive value, 80%). Moreover, patients in the validation dataset could be classified into two distinct prognostic groups using 63 predictive genes. CONCLUSIONS: Molecular analysis of 63 genes can predict the response of patients with advanced HCC and major portal vein tumor thrombi to combination chemotherapy with 5-fluorouracil and interferon-alpha.

Adult↗

Modeling and risk prediction in the current era of interventional cardiology: a report from the National Heart, Lung, and Blood Institute Dynamic Registry.

BACKGROUND: Validation of in-hospital mortality models after percutaneous coronary interventions using multicenter data remains limited. METHODS AND RESULTS: This study evaluated whether multivariable mortality models developed during the pre-stent era by New York State, American College of Cardiology (ACC)-National Cardiovascular Data Registry, Northern New England Cooperative Group, Cleveland Clinic Foundation, and the University of Michigan are relevant in patients undergoing percutaneous coronary intervention in the 1997 to 1999 National Heart, Lung, and Blood Institute Dynamic Registry. Of 4448 Dynamic Registry patients, 73% received > or =1 stent and 28% received a IIB/IIIA receptor inhibitor. In-hospital mortality occurred in 64 patients (1.4%). The New York state model predicted mortality in 69 patients (1.5%; 95% confidence bounds [CI], 0.89% to 1.70%); Northern New England predicted mortality in 60 patients (1.3%; 95% CI, 1.0% to 1.7%); and Cleveland Clinic predicted mortality in 76 patients (1.7%; 95% CI, 1.3% to 2.1%). Among high-risk subgroups, with these 3 models, observed and predicted in-hospital mortality rates in general were not different. The other 2 models yielded different results. The University of Michigan predicted fewer deaths (n=47; 1.1%; 95% CI, 0.7% to 1.3%), and the ACC Registry model predicted 603 deaths (13.5%; 95% CI, 12.6% to 14.4%). Using the ACC Registry model, predicted mortality was higher than observed in each subgroup. CONCLUSIONS: Application of 5 mortality risk models developed from different data sets to patients undergoing percutaneous coronary intervention in the Dynamic Registry predicted, in 3 models, mortality rates that were not significantly different than those observed. In both high and low risk subgroups, the University of Michigan slightly underpredicted mortality, and the ACC Registry predicted significantly higher mortality than that observed.

Angioplasty, Balloon, Coronary↗

Donor and recipient predicted lung volume and lung size after heart-lung transplantation.

Lung volumes after heart-lung transplantation (HLT) were recorded and compared with measurements at the time of assessment for surgery and the predicted values for recipients. The influence of donor lung size and recipients' underlying lung disease was evaluated. All patients underwent HLT between April 1984 and April 1991, and only those 82 who survived for at least 6 mo were studied. Mean total lung capacity (TLC) at preoperative assessment was 112% (SD = 28%) of the value predicted for recipients. One month after HLT, mean TLC was 83% (SD = 15%) of the predicted value but increased to 100% (SD = 15%) after 9 mo. No further change in average TLC occurred for 5 yr subsequently. The mean TLC of patients with emphysema before surgery was 164% (SD = 26%) of the predicted value and fell to the predicted value within 1 mo of HLT. The TLC in patients with primary pulmonary hypertension before surgery was close to the predicted value, but postoperative predicted TLC was achieved later than in emphysema patients. A donor-versus-recipient difference in TLC of more than 1L at the time of assessment did not influence the adaptation to the predicted value. FEV1 and vital capacity (VC) rose from means of 70% (SD = 25%) and 63% (SD = 20%) at 1 mo to 96% (SD = 27%) and 91% (SD = 18%), respectively, at 9 mo after HLT. After HLT, TLC returns to the predicted value for the recipient, and not to the preoperative TLC.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

Predictions made by psychiatrists and psychiatric nurses of violence by patients.

Studies have shown that it is difficult for psychiatrists to accurately predict which patients will be violent while hospitalized. The authors compared the predictions of 14 psychiatrists and nine psychiatric nurses who independently evaluated 308 patients consecutively admitted to a hospital in Israel and rated their likelihood of becoming violent. The psychiatrists and nurses also completed a general questionnaire about the criteria they used to predict violence. No significant differences were found in the accuracy of predictions between the two professional groups or in the criteria they used to predict violence. The total predictive value, or proportion of all cases predicted correctly, was 82 percent for the psychiatrists and 84 percent for the nurses. The predictions of the two groups coincided for 83 percent of the patients. The results suggest that psychiatrists and psychiatric nurses make similarly accurate predictions of violence and use similar criteria for making them.

Attitude↗

The mixed blessings of self-knowledge in behavioral prediction: enhanced discrimination but exacerbated bias.

Four experiments demonstrate that self-knowledge provides a mixed blessing in behavioral prediction, depending on how accuracy is measured. Compared with predictions of others, self-knowledge tends to decrease overall accuracy by increasing bias (the mean difference between predicted behavior and reality) but tends to increase overall accuracy by also enhancing discrimination (the correlation between predicted behavior and reality). Overall, participants' self-predictions overestimated the likelihood that they would engage in desirable behaviors (bias), whereas peer predictions were relatively unbiased. However, self-predictions also were more strongly correlated with individual differences in actual behavior (discrimination) than were peer predictions. Discussion addresses the costs and benefits of self-knowledge in behavioral prediction and the broader implications of measuring judgmental accuracy of judgment in terms of bias versus discrimination.

Adolescent↗

Predictive versus measured energy expenditure using limits-of-agreement analysis in hospitalized, obese patients.

BACKGROUND: Accuracy of predictive formulae is crucial for therapeutic planning because indirect calorimetry measurement is not always possible or cost effective. Energy requirements are more difficult to predict in the acutely ill obese patient compared with lean patients because of an increased resting energy expenditure per lean body mass and a variable stress response to illness. METHODS: A retrospective review of 726 patients identified 57 patients (32 spontaneous breathing, S; 25 ventilator dependent, V) with body mass indexes of 30-50 kg/m2. Limits-of-agreement analysis determined bias (the mean difference between measured and predicted values) and precision (the standard deviation of the bias) to evaluate the accuracy of predictive formulae compared with measured resting energy expenditure (MREE) by a Deltatrac Metabolic Monitor. Predictive accuracy was determined within+/-10% MREE. The predictive formulae examined were: variations of the Harris-Benedict equations using ideal, adjusted weights of 25% and 50% and actual weights with stress factors ranging from 1.0 to 1.5; the Ireton-Jones equation for obesity; the Ireton-Jones equations for hospitalized patients (S and V); and the ratio of 21 kcalories per kilogram actual weight. RESULTS: The mean MREE was 21 kcal/kg actual weight. The adjusted Harris-Benedict average weight equation was optimal for predicting MREE for the combined S and V sets (bias = 182+/-123; 67%+/-10% MREE), as well as the S subset (bias = 159+/-112; 69%+/-10% MREE). CONCLUSIONS: The Harris-Benedict equations using the average of actual and ideal weight and a stress factor of 1.3 most accurately predicted MREE in acutely ill, obese patients with BMIs of 30-50 kg/m2. Predictive formulae were least accurate for obese, ventilator-dependent patients.

Bias↗

A comparison of predictive outcomes of APACHE II and SAPS II in a surgical intensive care unit.

The Acute Physiologic Score and Chronic Health Evaluation (APACHE) II and the Simplified Acute Physiologic Scale (SAPS) II are two of the more commonly employed predictors of outcome and performance in the intensive care unit setting. However, controversy persists about whether the scores generated by these systems have similar predictive value. This study compared the predicted mortalities derived from APACHE II and SAPS II and contrasted them to the actual mortality in a surgical intensive care unit (SICU). Data for 1665 patients admitted to the SICU between July 1994 and August 1997 were entered into an SICU computerized database. From recorded demographic, hemodynamic, and laboratory data, APACHE II and SAPS II scores were obtained with corresponding predicted mortalities. Patients were stratified by age into categories of less than and greater than 65 years old. Predicted mortalities by APACHE II and SAPS II were compared for each group. An additional analysis included a comparison of survivors and nonsurvivors. There was no significant difference in predicted mortality between APACHE II and SAPS II in any of the groups. Actual mortality was 30 of 486 (6.2%) in patients less than 65 years of age and 73 of 1179 (6.2%) in patients 65 years of age or greater. The APACHE II and SAPS II predicted mortalities (mean +/- SD) for patients less than 65 years of age were 10.5% +/- 10.6% and 10.9% +/- 13.3%, respectively (P > .05). The APACHE II and SAPS II predicted mortalities in patients 65 years of age or greater were 19.1% +/- 17.8% and 18.7% +/- 21.0%, respectively (P > .05). Similarly, when patients were stratified by survival status, no significant difference was present between groups. However, in individual patients, a difference between APACHE II and SAPS II scores was often present. We conclude that although disparities between APACHE II and SAPS II predicted mortalities in individual patients may be significant, APACHE II and SAPS II have similar predictive value in a large SICU patient population. However, both APACHE II and SAPS II systems overestimate mortality in SICU patients. Based on our results, we conclude that either system can be used to measure quality of care in the SICU; however, neither system can be reliably applied to a single patient.

APACHE↗

Improved prediction of critical residues for protein function based on network and phylogenetic analyses.

BACKGROUND: Phylogenetic approaches are commonly used to predict which amino acid residues are critical to the function of a given protein. However, such approaches display inherent limitations, such as the requirement for identification of multiple homologues of the protein under consideration. Therefore, complementary or alternative approaches for the prediction of critical residues would be desirable. Network analyses have been used in the modelling of many complex biological systems, but only very recently have they been used to predict critical residues from a protein's three-dimensional structure. Here we compare a couple of phylogenetic approaches to several different network-based methods for the prediction of critical residues, and show that a combination of one phylogenetic method and one network-based method is superior to other methods previously employed. RESULTS: We associate a network with each member of a set of proteins for which the three-dimensional structure is known and the critical residues have been previously determined experimentally. We show that several network-based centrality measurements (connectivity, 2-connectivity, closeness centrality, betweenness and cluster coefficient) accurately detect residues critical for the protein's function. Phylogenetic approaches render predictions as reliable as the network-based measurements, although, interestingly, the two general approaches tend to predict different sets of critical residues. Hence we propose a hybrid method that is composed of one network-based calculation--the closeness centrality--and one phylogenetic approach--the Conseq server. This hybrid approach predicts critical residues more accurately than the other methods tested here. CONCLUSION: We show that network analysis can be used to improve the prediction of amino acids critical for protein function, when utilized in combination with phylogenetic approaches. It is proposed that such improvement is due to the complementary nature of these approaches: network-based methods tend to predict as critical those residues that are highly connected and internal (i.e., non-surface), although some surface residues are indeed identified as critical by network analyses; whereas residues chosen by phylogenetic approaches display a lower overall probability of being surface inaccessible.

Computational Biology↗

PSSM-based prediction of DNA binding sites in proteins.

BACKGROUND: Detection of DNA-binding sites in proteins is of enormous interest for technologies targeting gene regulation and manipulation. We have previously shown that a residue and its sequence neighbor information can be used to predict DNA-binding candidates in a protein sequence. This sequence-based prediction method is applicable even if no sequence homology with a previously known DNA-binding protein is observed. Here we implement a neural network based algorithm to utilize evolutionary information of amino acid sequences in terms of their position specific scoring matrices (PSSMs) for a better prediction of DNA-binding sites. RESULTS: An average of sensitivity and specificity using PSSMs is up to 8.7% better than the prediction with sequence information only. Much smaller data sets could be used to generate PSSM with minimal loss of prediction accuracy. CONCLUSION: One problem in using PSSM-derived prediction is obtaining lengthy and time-consuming alignments against large sequence databases. In order to speed up the process of generating PSSMs, we tried to use different reference data sets (sequence space) against which a target protein is scanned for PSI-BLAST iterations. We find that a very small set of proteins can actually be used as such a reference data without losing much of the prediction value. This makes the process of generating PSSMs very rapid and even amenable to be used at a genome level. A web server has been developed to provide these predictions of DNA-binding sites for any new protein from its amino acid sequence. AVAILABILITY: Online predictions based on this method are available at http://www.netasa.org/dbs-pssm/

Algorithms↗

Improving the specificity of high-throughput ortholog prediction.

BACKGROUND: Orthologs (genes that have diverged after a speciation event) tend to have similar function, and so their prediction has become an important component of comparative genomics and genome annotation. The gold standard phylogenetic analysis approach of comparing available organismal phylogeny to gene phylogeny is not easily automated for genome-wide analysis; therefore, ortholog prediction for large genome-scale datasets is typically performed using a reciprocal-best-BLAST-hits (RBH) approach. One problem with RBH is that it will incorrectly predict a paralog as an ortholog when incomplete genome sequences or gene loss is involved. In addition, there is an increasing interest in identifying orthologs most likely to have retained similar function. RESULTS: To address these issues, we present here a high-throughput computational method named Ortholuge that further evaluates previously predicted orthologs (including those predicted using an RBH-based approach) - identifying which orthologs most closely reflect species divergence and may more likely have similar function. Ortholuge analyzes phylogenetic distance ratios involving two comparison species and an outgroup species, noting cases where relative gene divergence is atypical. It also identifies some cases of gene duplication after species divergence. Through simulations of incomplete genome data/gene loss, we show that the vast majority of genes falsely predicted as orthologs by an RBH-based method can be identified. Ortholuge was then used to estimate the number of false-positives (predominantly paralogs) in selected RBH-predicted ortholog datasets, identifying approximately 10% paralogs in a eukaryotic data set (mouse-rat comparison) and 5% in a bacterial data set (Pseudomonas putida - Pseudomonas syringae species comparison). Higher quality (more precise) datasets of orthologs, which we term "ssd-orthologs" (supporting-species-divergence-orthologs), were also constructed. These datasets, as well as Ortholuge software that may be used to characterize other species' datasets, are available at http://www.pathogenomics.ca/ortholuge/ (software under GNU General Public License). CONCLUSION: The Ortholuge method reported here appears to significantly improve the specificity (precision) of high-throughput ortholog prediction for both bacterial and eukaryotic species. This method, and its associated software, will aid those performing various comparative genomics-based analyses, such as the prediction of conserved regulatory elements upstream of orthologous genes.

Algorithms↗

Network-based de-noising improves prediction from microarray data.

BACKGROUND: Prediction of human cell response to anti-cancer drugs (compounds) from microarray data is a challenging problem, due to the noise properties of microarrays as well as the high variance of living cell responses to drugs. Hence there is a strong need for more practical and robust methods than standard methods for real-value prediction. RESULTS: We devised an extended version of the off-subspace noise-reduction (de-noising) method to incorporate heterogeneous network data such as sequence similarity or protein-protein interactions into a single framework. Using that method, we first de-noise the gene expression data for training and test data and also the drug-response data for training data. Then we predict the unknown responses of each drug from the de-noised input data. For ascertaining whether de-noising improves prediction or not, we carry out 12-fold cross-validation for assessment of the prediction performance. We use the Pearson's correlation coefficient between the true and predicted response values as the prediction performance. De-noising improves the prediction performance for 65% of drugs. Furthermore, we found that this noise reduction method is robust and effective even when a large amount of artificial noise is added to the input data. CONCLUSION: We found that our extended off-subspace noise-reduction method combining heterogeneous biological data is successful and quite useful to improve prediction of human cell cancer drug responses from microarray data.

Algorithms↗