Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Estimating the association of antimicrobial resistance genes with minimum inhibitory concentration in Escherichia coli: an observational study.

BACKGROUND: Surveillance and prediction of antibiotic resistance in Escherichia coli relies on curated databases of genes and mutations. We aimed to quantify the effect of acquiring specific genetic elements on minimum inhibitory concentrations (MICs) for particular antibiotic-species combinations, addressing the current scarcity of such data in existing databases. METHODS: For this observational study, we evaluated a collection of E coli isolates with linked whole-genome sequencing and MIC data, originating from human urinary or bloodstream infections obtained from the Oxford University Hospitals National Health Service Foundation Trust in Oxfordshire, UK. We used multivariable interval regression models to estimate the change in MIC (with 95% CIs) for specific antibiotics associated with the acquisition of antibiotic resistance genes and associated mutations in the National Center for Biotechnology Information AMRFinder database, with and without an adjustment for population structure. We then tested the ability of these models to predict MIC and binary resistance or susceptibility using leave-one-out cross-validation. FINDINGS: We evaluated 2875 E coli isolates obtained during 2013-2018 and 2020. Although most ARGs and resistance mutations (89 [80%] of 111) were associated with an increased MIC, a much smaller number (27 [24%] of 111) was found to be putatively independently resistance-conferring (ie, associated with an MIC above the European Committee on Antimicrobial Susceptibility Testing breakpoint) when acquired in isolation. We found evidence of differential effects of acquired ARGs and resistance mutations between different generations of cephalosporin antibiotics and showed that sub-breakpoint variation in MIC can be linked to genetic mechanisms of resistance. 20 697 (83·3%; range 52·9-97·7 across all antibiotics) of 24 858 MICs were correctly exactly predicted and 23 677 (95·2%; 87·3-97·7) of 24 858 MICs were predicted to within one doubling dilution. INTERPRETATION: Quantitative estimates of the independent effect of the acquisition of ARGs on MIC add to the interpretability and utility of existing databases. Compared with approaches using machine learning models, the use of these estimates yields similar or better performance in the prediction of antibiotic resistance phenotype with more readily interpretable results. The methods outlined here could be readily applied to other antibiotic-pathogen combinations. FUNDING: The National Institute for Health and Care Research (NIHR) and the Medical Research Council (MRC).

Escherichia coli↗

Dietary Polyphenol Acteoside-Related Molecular Signatures in Clear Cell Renal Cell Carcinoma: Multi-Omics Profiling and Functional Validation of IMPDH1.

Clear cell renal cell carcinoma (ccRCC) is characterized by substantial metabolic and molecular heterogeneity, but the disease-relevant programs associated with acteoside, a dietary polyphenol, remain poorly understood. We integrated predicted acteoside targets with bulk, single-cell, and spatial transcriptomic data from ccRCC and combined molecular subtyping with cross-cohort machine-learning analysis. Acteoside-related signatures were preferentially enriched in malignant compartments and increased with tumor grade and stage. Consensus clustering identified two molecular subtypes with distinct biological and clinical features. C1 was associated with immune activation, metabolic activity, and more favorable survival, whereas C2 showed greater genomic instability, reduced renal epithelial differentiation, and poorer outcomes. We further benchmarked multiple machine-learning strategies and established a 10-gene prognostic model that retained predictive performance across independent cohorts, with IMPDH1 emerging as the strongest risk-associated feature. Functional experiments confirmed the biological relevance of IMPDH1: its knockdown suppressed ccRCC cell proliferation, DNA synthesis, colony formation, and migration, whereas overexpression produced the opposite effects. Together, these findings indicate that acteoside-related molecular signatures capture clinically relevant heterogeneity in ccRCC and provide a framework for linking dietary-polyphenol-related molecular space with tumor biology. The identification and functional validation of IMPDH1 further highlight its potential importance in ccRCC progression.

IMPDH1↗

Oncogenic EME1 promotes tumor progression and immune modulation in human cancers with therapeutic targeting potential.

BACKGROUND: EME1, a critical DNA repair endonuclease, has emerged as a potential oncogene implicated in genome instability and cancer progression. However, its pan-cancer roles, prognostic significance, immune interactions, and therapeutic targeting remain underexplored. METHODS: We conducted a comprehensive pan-cancer analysis integrating multi-omics data from public databases, including TIMER2.0, GEPIA2, TISIDB, and cBioPortal, to evaluate EME1 expression, genetic alterations, and their association with clinical outcomes, immune infiltration, and molecular pathways. Virtual screening of 3180 FDA-approved drugs and molecular dynamics (MD) simulations were employed to identify and validate potential EME1 inhibitors. RESULTS: EME1 was significantly overexpressed in various human cancers and positively associated with advanced tumor grade and stage. High EME1 expression and mutations were linked to poor overall and disease-free survival. Immunogenomic profiling revealed strong positive correlations between EME1 and myeloid-derived suppressor cells (MDSCs), alongside a negative association with endothelial cell function, suggesting immunosuppressive roles. Machine learning models based on EME1-associated genes demonstrated high predictive accuracy for liver hepatocellular carcinoma (AUC > 0.90). Virtual screening identified eight promising drug candidates, including Everolimus and Dioscin, with strong binding affinities. MD simulations confirmed the stability of these interactions, particularly for Dioscin. CONCLUSION: This study reveals the multifaceted oncogenic roles of EME1 in tumor progression, immune evasion, and prognosis. It proposes EME1 as a promising biomarker and therapeutic target across multiple cancer types. The identified drug candidates warrant further in vitro and in vivo validation for potential repurposing in EME1-targeted cancer therapy.

EME1↗

Sparse deconvolution of cell type medleys in spatial transcriptomics.

Mapping cell distributions across spatial locations with whole-genome coverage is essential for understanding cellular responses and signaling However, current deconvolution models aim to estimate the proportions of distinct cell types in each spatial transcriptomics spot by integrating reference single-cell data. These models often assume strong overlap between the reference and spatial datasets, neglecting biology-grounded constraints such as sparsity and cell-type variations, as well as technical sparsity. As a result, these methods rely on over-permissive algorithms that ignore given constraints leading to inaccurate predictions, particularly in heterogeneous or unmatched datasets. We introduce Weight-Induced Sparse Regression (WISpR), a machine learning algorithm that integrates spot-specific hyperparameters and sparsity-driven modeling. Unlike conventional approaches that neglect biology-grounded constraints, WISpR accurately predicts cell-type distributions while preserving biological coherence, i.e., spatially and functionally consistent cell-type localization, even in unmatched datasets. Benchmarking against five alternative methods across ten datasets, WISpR consistently outperformed competitors and predicted cellular landscapes in both normal and cancerous tissues. By leveraging sparse cell-type arrangements, WISpR provides biologically informed, high-resolution cellular maps. Its ability to decode tissue organization in both healthy and diseased states highlights WISpR's practical utility for spatial transcriptomics, particularly in challenging settings involving noise, sparsity, or reference mismatches.

Humans↗

Predicting cancer drug response by proteomic profiling.

PURPOSE: Accurate prediction of an individual patient's drug response is an important prerequisite of personalized medicine. Recent pharmacogenomics research in chemosensitivity prediction has studied the gene-drug correlation based on transcriptional profiling. However, proteomic profiling will more directly solve the current functional and pharmacologic problems. We sought to determine whether proteomic signatures of untreated cells were sufficient for the prediction of drug response. EXPERIMENTAL DESIGN: In this study, a machine learning model system was developed to classify cell line chemosensitivity exclusively based on proteomic profiling. Using reverse-phase protein lysate microarrays, protein expression levels were measured by 52 antibodies in a panel of 60 human cancer cell (NCI-60) lines. The model system combined several well-known algorithms, including random forests, Relief, and the nearest neighbor methods, to construct the protein expression--based chemosensitivity classifiers. The classifiers were designed to be independent of the tissue origin of the cells. RESULTS: A total of 118 classifiers of the complete range of drug responses (sensitive, intermediate, and resistant) were generated for the evaluated anticancer drugs, one for each agent. The accuracy of chemosensitivity prediction of all the evaluated 118 agents was significantly higher (P < 0.02) than that of random prediction. Furthermore, our study found that the proteomic determinants for chemosensitivity of 5-fluorouracil were also potential diagnostic markers of colon cancer. CONCLUSIONS: The results showed that it was feasible to accurately predict chemosensitivity by proteomic approaches. This study provides a basis for the prediction of drug response based on protein markers in the untreated tumors.

Antineoplastic Agents↗

Computational prediction of the chromosome-damaging potential of chemicals.

We report on the generation of computer-based models for the prediction of the chromosome-damaging potential of chemicals as assessed in the in vitro chromosome aberration (CA) test. On the basis of publicly available CA-test results of more than 650 chemical substances, half of which are drug-like compounds, we generated two different computational models. The first model was realized using the (Q)SAR tool MCASE. Results obtained with this model indicate a limited performance (53%) for the assessment of a chromosome-damaging potential (sensitivity), whereas CA-test negative compounds were correctly predicted with a specificity of 75%. The low sensitivity of this model might be explained by the fact that the underlying 2D-structural descriptors only describe part of the molecular mechanism leading to the induction of chromosome aberrations, that is, direct drug-DNA interactions. The second model was constructed with a more sophisticated machine learning approach and generated a classification model based on 14 molecular descriptors, which were obtained after feature selection. The performance of this model was superior to the MCASE model, primarily because of an improved sensitivity, suggesting that the more complex molecular descriptors in combination with statistical learning approaches are better suited to model the complex nature of mechanisms leading to a positive effect in the CA-test. An analysis of misclassified pharmaceuticals by this model showed that a large part of the false-negative predicted compounds were uniquely positive in the CA-test but lacked a genotoxic potential in other mutagenicity tests of the regulatory testing battery, suggesting that biologically nonsignificant mechanisms could be responsible for the observed positive CA-test result. Since such mechanisms are not amenable to modeling approaches it is suggested that a positive prediction made by the model reflects a biologically significant genotoxic potential. An integration of the machine-learning model as a screening tool in early discovery phases of drug development is proposed.

Chromosomes↗

Machine learning-guided risk stratification in elderly AML based on genomic, immunophenotypic and therapeutic profiles.

BACKGROUND: Elderly patients with acute myeloid leukemia (AML) exhibit considerable biological and clinical heterogeneity, hindering precise prognosis. Existing prognostic systems inadequately capture the complexity of elderly AML due to their reliance on data from younger cohorts and omission of key factors like immunophenotypic markers and therapeutic profiles. This study aimed to develop and internally validate a machine learning-based prognostic model specifically tailored to elderly AML patients. METHODS: A total of 156 patients were analyzed using a two-stage modeling strategy. Clinical and genomic variables were modeled first, followed by independent analysis of immunophenotypic features. Feature selection was performed using multilayer perceptron (MLP) and random forest (RF), while multivariate Cox regression was used for final model construction. Internal validation was conducted using 1000 bootstrap iterations to assess model stability and performance. RESULTS: The model demonstrated strong predictive performance, with a concordance index (C-index) of 0.702. Time-dependent area under the curve (AUC) and calibration plots confirmed accurate prediction of 1-, 3-, and 5-year overall survival. Decision curve analysis indicated favorable net benefit across a range of threshold probabilities. Key independent prognostic factors identified included TP53 mutations, high CD13 expression, and IDH2 mutations. CONCLUSION: This model provides a robust and interpretable tool for individualized risk stratification in elderly AML. By integrating genomic, immunophenotypic, and therapeutic variables, it may help optimize treatment decisions and improve outcomes for this vulnerable population. Future efforts should focus on external validation and integration of dynamic biomarkers.

Humans↗

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive↗

Major depletion of insulin sensitivity-associated taxa in the gut microbiome of persons living with HIV controlled by antiretroviral drugs.

BACKGROUND: Persons living with HIV (PWH) harbor an altered gut microbiome (higher abundance of Prevotella and lower abundance of Bacillota and Ruminococcus lineages) compared to non-infected individuals. Some of these alterations are linked to sexual preference and others to the HIV infection. The relationship between these lineages and metabolic alterations, often present in aging PWH, has been poorly investigated. METHODS: In this study, we compared fecal metagenomes of 25 antiretroviral-treatment (ART)-controlled PWH to three independent control groups of 25 non-infected matched individuals by means of univariate analyses and machine learning methods. Moreover, we used two external datasets to validate predictive models of PWH classification. Next, we searched for associations between clinical and biological metabolic parameters with taxonomic and functional microbiome profiles. Finally, we compare the gut microbiome in 7 PWH after a 17-week ART switch to raltegravir/maraviroc. RESULTS: Three major enterotypes (Prevotella, Bacteroides and Ruminococcaceae) were present in all groups. The first Prevotella enterotype was enriched in PWH, with several of characteristic lineages associated with poor metabolic profiles (low HDL and adiponectin, high insulin resistance (HOMA-IR)). Conversely butyrate-producing lineages were markedly depleted in PWH independently of sexual preference and were associated with a better metabolic profile (higher HDL and adiponectin and lower HOMA-IR). Accordingly with the worst metabolic status of PWH, butyrate production and amino-acid degradation modules were associated with high HDL and adiponectin and low HOMA-IR. Random Forest models trained to classify PWH vs. control on taxonomic abundances displayed high generalization performance on two external holdout datasets (ROC AUC of 80-82%). Finally, no significant alterations in microbiome composition were observed after switching to raltegravir/maraviroc. CONCLUSION: High resolution metagenomic analyses revealed major differences in the gut microbiome of ART-controlled PWH when compared with three independent matched cohorts of controls. The observed marked insulin resistance could result both from enrichment in Prevotella lineages, and from the depletion in species producing butyrate and involved into amino-acid degradation, which depletion is linked with the HIV infection.

Humans↗

Prediction of alpha-turns in proteins using PSI-BLAST profiles and secondary structure information.

In this paper a systematic attempt has been made to develop a better method for predicting alpha-turns in proteins. Most of the commonly used approaches in the field of protein structure prediction have been tried in this study, which includes statistical approach "Sequence Coupled Model" and machine learning approaches; i) artificial neural network (ANN); ii) Weka (Waikato Environment for Knowledge Analysis) Classifiers and iii) Parallel Exemplar Based Learning (PEBLS). We have also used multiple sequence alignment obtained from PSIBLAST and secondary structure information predicted by PSIPRED. The training and testing of all methods has been performed on a data set of 193 non-homologous protein X-ray structures using five-fold cross-validation. It has been observed that ANN with multiple sequence alignment and predicted secondary structure information outperforms other methods. Based on our observations we have developed an ANN-based method for predicting alpha-turns in proteins. The main components of the method are two feed-forward back-propagation networks with a single hidden layer. The first sequence-structure network is trained with the multiple sequence alignment in the form of PSI-BLAST-generated position specific scoring matrices. The initial predictions obtained from the first network and PSIPRED predicted secondary structure are used as input to the second structure-structure network to refine the predictions obtained from the first net. The final network yields an overall prediction accuracy of 78.0% and MCC of 0.16. A web server AlphaPred (http://www.imtech.res.in/raghava/alphapred/) has been developed based on this approach.

Amino Acid Sequence↗

Diagnosing anorexia based on partial least squares, back propagation neural network, and support vector machines.

Support vector machine (SVM), as a novel type of learning machine, for the first time, was used to develop a predictive model for early diagnosis of anorexia. It was based on the concentration of six elements (Zn, Fe, Mg, Cu, Ca, and Mn) and the age extracted from 90 cases. Compared with the results obtained from two other classifiers, partial least squares (PLS) and back-propagation neural network (BPNN), the SVM method exhibited the best whole performance. The accuracies for the test set by PLS, BPNN, and SVM methods were 52%, 65%, and 87%, respectively. Moreover, the models we proposed could also provide some insight into what factors were related to anorexia.

Anorexia↗

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 684 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

Journal Article↗

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease.

OBJECTIVE: To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. METHODS: For this critical review, Medline, Embase and IEEE were searched from inception to 1 January 2025. Included were studies describing machine learning algorithms designed to specifically compare output of cardiovascular risk assessment with the FRS. Commentaries, letters, unpublished work or non-peer-reviewed papers were excluded.Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, two reviewers screened titles and abstracts independently, then populated a purpose-built data extraction form. A subsequent qualitative thematic analysis focused on algorithms' strengths, added value, potential harms, unintended consequences and equity implications.The main outcome assessed was whether, among healthy adults, the algorithm improved CVD risk prediction relative to the FRS. RESULTS: Of 707 studies retrieved, 29 met inclusion criteria. 23 reported improved predictive ability relative to the FRS. Most datasets and/or medical records used included sociodemographic predictors of CVD not included among FRS inputs. Some added costly diagnostic tests like CT angiography to FRS screening indicators. When they were defined, inputs and outcomes such as hypertension or myocardial infarction did not always adhere to FRS values. Statistical significance was generally taken as a proxy for clinical significance. Some algorithms overestimated the number at risk compared with the FRS without discussing whether that larger proportion might be at risk of overdiagnosis rather than CVD, while a few decreased the proportion found to be at risk. CONCLUSIONS: Use of artificial intelligence to improve accuracy of risk assessment for CVD demonstrates the technological capacity to merge known sociodemographic predictors with biologic variables and examine non-linear interactions among these. Still needed to achieve patient benefit is clinical insight, adherence to screening principles and cost-benefit assessment of inputs selected.

Humans↗

STRUMP-I: Structure-based machine learning approach to pMHC-I binding prediction using force field energy features.

The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on Major Histocompatibility Complex class I (MHC-I) molecules, collectively termed peptide-MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting neoantigens in cancers, the ability to accurately predict which peptides will bind to the diverse population of MHC alleles is critically important. Current computational methods for pMHC-I prediction fall broadly into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and energetics of pMHC binding. While sequence-based methods have been popularly used, their performance is dependent on the size and quality of training data. On the other hands, while structure-based approaches can generalize better across diverse MHC alleles, they traditionally depend on identifying a single global minimum energy conformation, an assumption that often fails due to the inherent binding promiscuity of MHC-I molecules. To address these limitations, we developed a STRUMP-I (STRUcture-based pMHC Prediction (for class I)), a novel pMHC binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine-learning features. STRUMP-I achieves performance comparable to state-of-the-art sequence-based models while significantly outperforming them on MHC alleles with limited representation in training data. Furthermore, STRUMP-I demonstrates strong synergy when integrated with sequence-based methods, notably enhancing prediction precision. The robustness and generalizability of STRUMP-I were confirmed by evaluating its predictive performance on independent, previously unseen datasets, including an experimentally validated cancer neoantigen dataset. This combined approach advances our capability to reliably identify clinically relevant neoantigen targets. The source code and trained models are available at https://github.com/yoonjoolab/STRUMP-I.

energy optimization↗

Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.

Diffuse large B-cell lymphoma (DLBCL), the most common lymphoid malignancy in adults, is curable in less than 50% of patients. Prognostic models based on pre-treatment characteristics, such as the International Prognostic Index (IPI), are currently used to predict outcome in DLBCL. However, clinical outcome models identify neither the molecular basis of clinical heterogeneity, nor specific therapeutic targets. We analyzed the expression of 6,817 genes in diagnostic tumor specimens from DLBCL patients who received cyclophosphamide, adriamycin, vincristine and prednisone (CHOP)-based chemotherapy, and applied a supervised learning prediction method to identify cured versus fatal or refractory disease. The algorithm classified two categories of patients with very different five-year overall survival rates (70% versus 12%). The model also effectively delineated patients within specific IPI risk categories who were likely to be cured or to die of their disease. Genes implicated in DLBCL outcome included some that regulate responses to B-cell-receptor signaling, critical serine/threonine phosphorylation pathways and apoptosis. Our data indicate that supervised learning classification techniques can predict outcome in DLBCL and identify rational targets for intervention.

Antineoplastic Combined Chemotherapy Protocols↗

Concept formation vs. logistic regression: predicting death in trauma patients.

This study compares two classification models used to predict survival of injured patients entering the emergency department. Concept formation is a machine learning technique that summarizes known examples cases in the form of a tree. After the tree is constructed, it can then be used to predict the classification of new cases. Logistic regression, on the other hand, is a statistical model that allows for a quantitative relationship for a dichotomous event with several independent variables. The outcome (dependent) variable must have only two choices, e.g. does or does not occur, alive or dead, etc. The result of this model is an equation which is then used to predict the probability of class membership of a new case. The two models were evaluated on a trauma registry database composed of information on all trauma patients admitted in 1992 to a Level I trauma center. A total of 2155 records. representing all trauma patients admitted for more than 24 h or who died in the Emergency Department, were grouped into two databases as follows: (1) discharge status of 'died' (containing 151 records), and (2) any discharge status other than 'died' (containing 2004 records). Both databases contained the same variables.

Artificial Intelligence↗

Patient-specific models for predicting the outcomes of patients with community acquired pneumonia.

We investigated two patient-specific and four population-wide machine learning methods for predicting dire outcomes in community acquired pneumonia (CAP) patients. Predicting dire outcomes in CAP patients can significantly influence the decision about whether to admit the patient to the hospital or to treat the patient at home. Population-wide methods induce models that are trained to perform well on average on all future cases. In contrast, patient-specific methods specifically induce a model for a particular patient case. We trained the models on a set of 1601 patient cases and evaluated them on a separate set of 686 cases. One patient-specific method performed better than the population-wide methods when evaluated within a clinically relevant range of the ROC curve. Our study provides support for patient-specific methods being a promising approach for making clinical predictions.

Algorithms↗