Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “random forest”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A comparison of decision tree ensemble creation techniques.

We experimentally evaluate bagging and seven other randomization-based approaches to creating an ensemble of decision tree classifiers. Statistical tests were performed on experimental results from 57 publicly available data sets. When cross-validation comparisons were tested for statistical significance, the best method was statistically more accurate than bagging on only eight of the 57 data sets. Alternatively, examining the average ranks of the algorithms across the group of data sets, we find that boosting, random forests, and randomized trees are statistically significantly better than bagging. Because our results suggest that using an appropriate ensemble size is important, we introduce an algorithm that decides when a sufficient number of classifiers has been created for an ensemble. Our algorithm uses the out-of-bag error estimate, and is shown to result in an accurate ensemble for those methods that incorporate bagging into the construction of the ensemble.

Algorithms↗

Population genetic structure of two rare tree species (Colubrina oppositifolia and Alphitonia ponderosa, Rhamnaceae) from Hawaiian dry and mesic forests using random amplified polymorphic DNA markers.

Hawaiian dry and mesic forests contain an increasingly rare assemblage of species due to habitat destruction, invasive alien weeds and exotic pests. Two rare Rhamnaceae species in these ecosystems, Colubrina oppositifolia and Alphitonia ponderosa, were examined using random amplified polymorphic DNA (RAPD) markers to determine the genetic structure of the populations and the amount of variation relative to other native Hawaiian species. Relative variation is lower than with other Hawaiian species, although this is probably not a consequence of genetic bottleneck. Larger populations of both species contain the highest levels of genetic diversity and smaller populations generally the least as determined by number of polymorphic loci, estimated heterozygosity, and Shannon's index of genetic diversity. Populations on separate islands were readily discernible for both species as were two populations of C. oppositifolia on Hawai'i island (North and South Kona populations). Substructure among Kaua'i subpopulations of A. ponderosa that were ecologically separated was also evident. Although population diversity is thought to have remained at predisturbance levels, population size continues to decline as recruitment is either absent or does not keep pace with senescence of mature plants. Recovery efforts must focus on control of alien species if these and other endemic dry and mesic forest species are to persist.

Colubrina↗

Application and comparison of classification algorithms for recognition of Alzheimer's disease in electrical brain activity (EEG).

The early detection of subjects with probable Alzheimer's disease (AD) is crucial for effective appliance of treatment strategies. Here we explored the ability of a multitude of linear and non-linear classification algorithms to discriminate between the electroencephalograms (EEGs) of patients with varying degree of AD and their age-matched control subjects. Absolute and relative spectral power, distribution of spectral power, and measures of spatial synchronization were calculated from recordings of resting eyes-closed continuous EEGs of 45 healthy controls, 116 patients with mild AD and 81 patients with moderate AD, recruited in two different centers (Stockholm, New York). The applied classification algorithms were: principal component linear discriminant analysis (PC LDA), partial least squares LDA (PLS LDA), principal component logistic regression (PC LR), partial least squares logistic regression (PLS LR), bagging, random forest, support vector machines (SVM) and feed-forward neural network. Based on 10-fold cross-validation runs it could be demonstrated that even tough modern computer-intensive classification algorithms such as random forests, SVM and neural networks show a slight superiority, more classical classification algorithms performed nearly equally well. Using random forests classification a considerable sensitivity of up to 85% and a specificity of 78%, respectively for the test of even only mild AD patients has been reached, whereas for the comparison of moderate AD vs. controls, using SVM and neural networks, values of 89% and 88% for sensitivity and specificity were achieved. Such a remarkable performance proves the value of these classification algorithms for clinical diagnostics.

Aged↗

Discrimination of intact mycobacteria at the strain level: a combined MALDI-TOF MS and biostatistical analysis.

New methodologies for surveillance and identification of Mycobacterium tuberculosis are required to stem the spread of disease worldwide. In addition, the ability to discriminate mycobacteria at the strain level may be important to contact or source case investigations. To this end, we are developing MALDI-TOF MS methods for the identification of M. tuberculosis in culture. In this report, we describe the application of MALDI-TOF MS, as well as statistical analysis including linear discriminant and random forest analysis, to 16 medically relevant strains from four species of mycobacteria, M. tuberculosis, M. avium, M. intracellulare, and M. kansasii. Although species discrimination can be accomplished on the basis of unique m/z values observed in the MS fingerprint spectrum, discrimination at the strain level is predicted on the relative abundance of shared m/z values among strains within a species. For the 16 mycobacterial strains investigated in the present study, it is possible to unambiguously identify strains within a species on the basis of MALDI-TOF MS data. The error rate for classification of individual strains using linear discriminant analysis was 0.053 using 37 m/z variables, whereas the error rate for classification of individual strains using random forest analysis was 0.023 using only 18 m/z variables. In addition, using random forest analysis of MALDI-TOF MS data, it was possible to correctly classify bacterial strains as either M. tuberculosis or non-tuberculous with 100% accuracy.

Bacterial Proteins↗

Few amino acid positions in rpoB are associated with most of the rifampin resistance in Mycobacterium tuberculosis.

BACKGROUND: Mutations in rpoB, the gene encoding the beta subunit of DNA-dependent RNA polymerase, are associated with rifampin resistance in Mycobacterium tuberculosis. Several studies have been conducted where minimum inhibitory concentration (MIC, which is defined as the minimum concentration of the antibiotic in a given culture medium below which bacterial growth is not inhibited) of rifampin has been measured and partial DNA sequences have been determined for rpoB in different isolates of M. tuberculosis. However, no model has been constructed to predict rifampin resistance based on sequence information alone. Such a model might provide the basis for quantifying rifampin resistance status based exclusively on DNA sequence data and thus eliminate the requirements for time consuming culturing and antibiotic testing of clinical isolates. RESULTS: Sequence data for amino acid positions 511-533 of rpoB and associated MIC of rifampin for different isolates of M. tuberculosis were taken from studies examining rifampin resistance in clinical samples from New York City and throughout Japan. We used tree-based statistical methods and random forests to generate models of the relationships between rpoB amino acid sequence and rifampin resistance. The proportion of variance explained by a relatively simple tree-based cross-validated regression model involving two amino acid positions (526 and 531) is 0.679. The first partition in the data, based on position 531, results in groups that differ one hundredfold in mean MIC (1.596 micrograms/ml and 159.676 micrograms/ml). The subsequent partition based on position 526, the most variable in this region, results in a > 354-fold difference in MIC. When considered as a classification problem (susceptible or resistant), a cross-validated tree-based model correctly classified most (0.884) of the observations and was very similar to the regression model. Random forest analysis of the MIC data as a continuous variable, a regression problem, produced a model that explained 0.861 of the variance. The random forest analysis of the MIC data as discrete classes produced a model that correctly classified 0.942 of the observations with sensitivity of 0.958 and specificity of 0.885. CONCLUSIONS: Highly accurate regression and classification models of rifampin resistance can be made based on this short sequence region. Models may be better with improved (and consistent) measurements of MIC and more sequence data.

Amino Acids↗

The challenge for genetic epidemiologists: how to analyze large numbers of SNPs in relation to complex diseases.

Genetic epidemiologists have taken the challenge to identify genetic polymorphisms involved in the development of diseases. Many have collected data on large numbers of genetic markers but are not familiar with available methods to assess their association with complex diseases. Statistical methods have been developed for analyzing the relation between large numbers of genetic and environmental predictors to disease or disease-related variables in genetic association studies. In this commentary we discuss logistic regression analysis, neural networks, including the parameter decreasing method (PDM) and genetic programming optimized neural networks (GPNN) and several non-parametric methods, which include the set association approach, combinatorial partitioning method (CPM), restricted partitioning method (RPM), multifactor dimensionality reduction (MDR) method and the random forests approach. The relative strengths and weaknesses of these methods are highlighted. Logistic regression and neural networks can handle only a limited number of predictor variables, depending on the number of observations in the dataset. Therefore, they are less useful than the non-parametric methods to approach association studies with large numbers of predictor variables. GPNN on the other hand may be a useful approach to select and model important predictors, but its performance to select the important effects in the presence of large numbers of predictors needs to be examined. Both the set association approach and random forests approach are able to handle a large number of predictors and are useful in reducing these predictors to a subset of predictors with an important contribution to disease. The combinatorial methods give more insight in combination patterns for sets of genetic and/or environmental predictor variables that may be related to the outcome variable. As the non-parametric methods have different strengths and weaknesses we conclude that to approach genetic association studies using the case-control design, the application of a combination of several methods, including the set association approach, MDR and the random forests approach, will likely be a useful strategy to find the important genes and interaction patterns involved in complex diseases.

Editorial↗

Artificial intelligence in treatment prediction for skeletal Class III malocclusion: A systematic review.

In skeletal Class III patients, treatment options range from orthodontics to orthognathic surgery. Choosing the optimal approach requires a comprehensive clinical evaluation, which may be supported by AI tools. The aim of this study was to assess the performance of AI models in predicting the need for orthognathic surgery and in identifying predictors influencing treatment decisions. A PRISMA-guided electronic database search (PubMed, Web of Science; 2009-2024; English/French) was performed to identify studies using machine learning (ML) or deep learning (DL) on cephalometric and clinical data. After screening and assessment for eligibility, 15 studies were critically appraised. Model performance was summarized using accuracy, sensitivity, specificity, and the area under the curve (AUC). ML algorithms (particularly Random Forest and XGBoost) and DL models (ResNet-based convolutional neural networks (CNNs)) achieved high accuracy for predicting surgical need. Frequently selected predictors included Wits appraisal, ANB angle, the maxillomandibular ratio (Mx/Md), overjet, and the divergence of the lower gonial angle. AI methods show promise for assisting treatment decisions in Class III malocclusion, with Random Forest and XGBoost performing well on tabular cephalometric data and CNNs on imaging. Larger, multicentre datasets and external validation are needed to improve reliability, address bias, and support clinical implementation.

Humans↗

Predicting ACL injury risk in athletes: A systematic review of machine learning-based models.

BACKGROUND: Early ACL injury risk identification in athletes is essential. This systematic review examines machine learning (ML) models for predicting ACL injuries, evaluating their methodological quality, performance, and reliability. METHOD: A comprehensive electronic search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases, supplemented by Google Scholar for grey literature, covering articles published between January 1, 2015, and August 30, 2025. Eligible studies were appraised using the Prediction Model Study Risk of Bias Assessment Tool (PROBAST) for methodological quality and risk of bias, and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines for quality of evidence. RESULTS: Ten studies were included. PROBAST showed eight studies had moderate risk of bias and two low risk. TRIPOD found only two studies met quality criteria. ML models included logistic regression (n = 5), support vector machines (n = 4), k-nearest neighbor (n = 3), decision trees (n = 3), random forests (n = 5), neural networks (n = 2), linear discriminant analysis (n = 1), and pre-trained CNNs (n = 1). AUC ranged from 0.63 to 0.98. Accuracy (reported in six studies) ranged from 26% to 95%; however, these values should be interpreted with caution due to the absence of confidence intervals, lack of class imbalance handling, and limited external validation across studies. Tree-based ensemble methods such as random forest achieved competitive accuracy (74-86%), while SVM, a non-ensemble classifier, reported accuracy ranging from 71% to 95%; however, the highest values were obtained in studies with notably small sample sizes (n = 12 to n = 39), raising concerns about overfitting and generalizability. CONCLUSION: Current ML algorithms show promise for identifying athletes at high ACL injury risk and detecting relevant risk factors. Although study quality was generally satisfactory, future research should prioritize external validation and model interpretability to support clinical translation.

Humans↗

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans↗

Bioinformatics Analysis and Experimental Validation of Key Genes Associated With Hypoxia and Ischemia in Myocardial Infarction.

BACKGROUND: This study aimed to screen and identify core hypoxia-ischemia-related genes associated with myocardial infarction (MI). METHOD: Two transcriptomic datasets, GSE97320 and GSE48060, were retrieved from the Gene Expression Omnibus (GEO) database. After data integration and batch effect elimination, differential expression analysis was performed to screen differentially expressed genes (DEGs), and the corresponding visualization analysis was conducted. Hypoxia-ischemia-related genes were acquired from the GeneCards database; hypoxia-ischemia related genes (HIRGs) were subsequently identified by intersecting the retrieved genes with screened DEGs. Gene Ontology (GO) functional enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were implemented to explore the biological functions and underlying signaling pathways of HIRGs. A combination of protein-protein interaction (PPI) network analysis and random forest (RF) algorithm was applied to screen hub genes from HIRGs. The external GEO dataset GSE66360 was utilized to validate the expression patterns of candidate hub genes. Furthermore, an acute myocardial infarction (AMI) mouse model was established, and quantitative real-time polymerase chain reaction (qPCR) was performed to detect the mRNA expression levels of hub genes in myocardial tissues for in&#xa0;vivo validation. RESULTS: A total of 633 DEGs and 308 hypoxia-ischemia-related genes were screened in the present study, among which 21 overlapping HIRGs were obtained. PLAUR and IL1B were finally identified as two hub genes from HIRGs based on PPI network and random forest algorithm. The qPCR results revealed that the expression levels of PLAUR and IL1B were significantly upregulated in the AMI group compared with the sham operation group (p&#x2009;<&#x2009;0.05). CONCLUSION: The present findings demonstrated that PLAUR and IL1B serve as pivotal genes involved in the pathological hypoxia-ischemia process of AMI. These two genes may act as novel biomarkers and promising therapeutic targets for the recognition and clinical intervention of hypoxia-ischemia injury following AMI.

Myocardial Infarction↗

Urinary biomarker profiling in transitional cell carcinoma.

Urinary biomarkers or profiles that allow noninvasive detection of recurrent transitional cell carcinoma (TCC) of the bladder are urgently needed. We obtained duplicate proteomic (SELDI) profiles from 227 subjects (118 TCC, 77 healthy controls and 32 controls with benign urological conditions) and used linear mixed effects models to identify peaks that are differentially expressed between TCC and controls and within TCC subgroups. A Random Forest classifier was trained on 130 profiles to develop an algorithm to predict the presence of TCC in a randomly selected initial test set (n = 54) and an independent validation set (n = 43) several months later. Twenty two peaks were differentially expressed between all TCC and controls (p < 10(-7)). However potential confounding effects of age, sex and analytical run were identified. In an age-matched sub-set, 23 peaks were differentially expressed between TCC and combined benign and healthy controls at the 0.005 significance level. Using the Random Forest classifier, TCC was predicted with 71.7% sensitivity and 62.5% specificity in the initial set and with 78.3% sensitivity and 65.0% specificity in the validation set after 6 months, compared with controls. Several peaks of importance were also identified in the linear mixed effects model. We conclude that SELDI profiling of urine samples can identify patients with TCC with comparable sensitivities and specificities to current tumor marker tests. This is the first time that reproducibility has been demonstrated on an independent test set analyzed several months later. Identification of the relevant peaks may facilitate multiplex marker assay development for detection of recurrent disease.

Adult↗

Relating HIV-1 sequence variation to replication capacity via trees and forests.

The problem of relating genotype (as represented by amino acid sequence) to phenotypes is distinguished from standard regression problems by the nature of sequence data. Here we investigate an instance of such a problem where the phenotype of interest is HIV-1 replication capacity and contiguous segments of protease and reverse transcriptase sequence constitutes genotype. A variety of data analytic methods have been proposed in this context. Shortcomings of select techniques are contrasted with the advantages afforded by tree-structured methods. However, tree-structured methods, in turn, have been criticized on grounds of only enjoying modest predictive performance. A number of ensemble approaches (bagging, boosting, random forests) have recently emerged, devised to overcome this deficiency. We evaluate random forests as applied in this setting, and detail why prediction gains obtained in other situations are not realized. Other approaches including logic regression, support vector machines and neural networks are also applied. We interpret results in terms of HIV-1 reverse transcriptase structure and function.

Journal Article↗

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation↗

Translating microarray data for diagnostic testing in childhood leukaemia.

BACKGROUND: Recent findings from microarray studies have raised the prospect of a standardized diagnostic gene expression platform to enhance accurate diagnosis and risk stratification in paediatric acute lymphoblastic leukaemia (ALL). However, the robustness as well as the format for such a diagnostic test remains to be determined. As a step towards clinical application of these findings, we have systematically analyzed a published ALL microarray data set using Robust Multi-array Analysis (RMA) and Random Forest (RF). METHODS: We examined published microarray data from 104 ALL patients specimens, that represent six different subgroups defined by cytogenetic features and immunophenotypes. Using the decision-tree based supervised learning algorithm Random Forest (RF), we determined a small set of genes for optimal subgroup distinction and subsequently validated their predictive power in an independent patient cohort. RESULTS: We achieved very high overall ALL subgroup prediction accuracies of about 98%, and were able to verify the robustness of these genes in an independent panel of 68 specimens obtained from a different institution and processed in a different laboratory. Our study established that the selection of discriminating genes is strongly dependent on the analysis method. This may have profound implications for clinical use, particularly when the classifier is reduced to a small set of genes. We have demonstrated that as few as 26 genes yield accurate class prediction and importantly, almost 70% of these genes have not been previously identified as essential for class distinction of the six ALL subgroups. CONCLUSION: Our finding supports the feasibility of qRT-PCR technology for standardized diagnostic testing in paediatric ALL and should, in conjunction with conventional cytogenetics lead to a more accurate classification of the disease. In addition, we have demonstrated that microarray findings from one study can be confirmed in an independent study, using an entirely independent patient cohort and with microarray experiments being performed by a different research team.

Cell Line, Tumor↗

Efficacy of the NMIC-150 system in identifying extended-spectrum beta-lactamases in clinical isolates.

Extended-spectrum beta-lactamases (ESBLs) are significant contributors to the growing global crisis of antimicrobial resistance. This study evaluated the performance of the NMIC-150 System for susceptibility testing of third-generation cephalosporins (3GCs) and assessed whether ceftazidime-avibactam and aztreonam-avibactam could identify ESBL-producing carbapenem-resistant Enterobacterales (CREs). A total of 278 non-duplicate clinical isolates (Klebsiella pneumoniae, E. coli, and Proteus mirabilis) were analyzed. Antimicrobial susceptibility was determined using reference broth microdilution (BMD) and the NMIC-150 System. ESBL production was defined as an &#x2265;eight-fold reduction in the minimum inhibitory concentration (MIC) of 3GCs in the presence of clavulanic acid, according to CLSI criteria. Whole-genome sequencing was performed to characterize ESBL and carbapenemase genes among 3GC-resistant isolates. A Random Forest model was used to predict ESBL-producing isolates based on MIC values. The NMIC-150 System demonstrated over 90% categorical and essential agreement with BMD for ceftazidime and ceftriaxone, along with robust predictive performance via Random Forest analysis. These findings suggest that the NMIC-150 System is a reliable platform for 3GC susceptibility testing and that an &#x2265;eight-fold MIC reduction with ceftazidime-avibactam or aztreonam-avibactam may serve as a phenotypic indicator of ESBL production in CRE isolates. In conclusion, the NMIC-150 System shows potential for routine antimicrobial resistance surveillance and may facilitate the rapid identification of ESBL-producing CREs in clinical settings.

Microbial Sensitivity Tests↗

Development of linear, ensemble, and nonlinear models for the prediction and interpretation of the biological activity of a set of PDGFR inhibitors.

A QSAR modeling study has been done with a set of 79 piperazyinylquinazoline analogues which exhibit PDGFR inhibition. Linear regression and nonlinear computational neural network models were developed. The regression model was developed with a focus on interpretative ability using a PLS technique. However, it also exhibits a good predictive ability after outlier removal. The nonlinear CNN model had superior predictive ability compared to the linear model with a training set error of 0.22 log(IC50) units (R2 = 0.93) and a prediction set error of 0.32 log(IC50) units (R2 = 0.61). A random forest model was also developed to provide an alternate measure of descriptor importance. This approach ranks descriptors, and its results confirm the importance of specific descriptors as characterized by the PLS technique. In addition the neural network model contains the two most important descriptors indicated by the random forest model.

Linear Models↗

Survival ensembles.

We propose a unified and flexible framework for ensemble learning in the presence of censoring. For right-censored data, we introduce a random forest algorithm and a generic gradient boosting algorithm for the construction of prognostic and diagnostic models. The methodology is utilized for predicting the survival time of patients suffering from acute myeloid leukemia based on clinical and genetic covariates. Furthermore, we compare the diagnostic capabilities of the proposed censored data random forest and boosting methods, applied to the recurrence-free survival time of node-positive breast cancer patients, with previously published findings.

Algorithms↗

Prediction of clinical drug efficacy by classification of drug-induced genomic expression profiles in vitro.

Assays of drug action typically evaluate biochemical activity. However, accurately matching therapeutic efficacy with biochemical activity is a challenge. High-content cellular assays seek to bridge this gap by capturing broad information about the cellular physiology of drug action. Here, we present a method of predicting the general therapeutic classes into which various psychoactive drugs fall, based on high-content statistical categorization of gene expression profiles induced by these drugs. When we used the classification tree and random forest supervised classification algorithms to analyze microarray data, we derived general "efficacy profiles" of biomarker gene expression that correlate with anti-depressant, antipsychotic and opioid drug action on primary human neurons in vitro. These profiles were used as predictive models to classify naïve in vitro drug treatments with 83.3% (random forest) and 88.9% (classification tree) accuracy. Thus, the detailed information contained in genomic expression data is sufficient to match the physiological effect of a novel drug at the cellular level with its clinical relevance. This capacity to identify therapeutic efficacy on the basis of gene expression signatures in vitro has potential utility in drug discovery and drug target validation.

Algorithms↗