Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Toward a Better Paradigm for Head and Neck Cancer Treatment Applying AI (HNC-TACTIC): Protocol for an International Cohort Study of Electronic Health Records.

BACKGROUND: Head and neck squamous cell carcinomas (HNSCCs) cause considerable morbidity and mortality. Multimodal treatment strategies can cause significant toxicity, and therapy options are limited for recurrent disease. Immunotherapy has emerged as a promising approach. However, patient response variability underscores the need for better predictive markers. OBJECTIVE: This study aims to use artificial intelligence to develop two predictive models in patients with HNSCC to assess (1) progression or recurrence following primary curative treatment and (2) long-term survival after immunotherapy schemes in recurrent and metastatic disease. This study will also describe the characteristics of patients with early, locally advanced, and recurrent or metastatic cancers. METHODS: This is a retrospective, observational study of data captured in electronic health records (EHRs) from participating hospitals between January 1, 2014, and December 31, 2021. This study's population comprises adults diagnosed with HNSCC at any stage. Study variables, including demographics, comorbidities, clinical variables, treatments, and outcomes, will be extracted using EHRead, a technology that applies natural language processing and machine learning to extract and analyze structured and unstructured clinical information in deidentified EHRs. Predictive models based on dynamic risk stratification for treatment response and progression or recurrence will be developed using multivariable logistic regressions, decision tree classifiers, and random forest approaches. Descriptive and outcome analyses will be shown for different anatomic subsites and stratified by stage and treatment. RESULTS: This study began enrolling sites in July 2021 and is currently ongoing. By December 2025, data from 10 centers has been collected, comprising a total of 151,934,990 EHRs from 2,159,719 patients. CONCLUSIONS: Development of predictive models using artificial intelligence will advance clinical understanding of HNSCC to improve patient outcomes.

Humans↗

Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.

Diffuse large B-cell lymphoma (DLBCL), the most common lymphoid malignancy in adults, is curable in less than 50% of patients. Prognostic models based on pre-treatment characteristics, such as the International Prognostic Index (IPI), are currently used to predict outcome in DLBCL. However, clinical outcome models identify neither the molecular basis of clinical heterogeneity, nor specific therapeutic targets. We analyzed the expression of 6,817 genes in diagnostic tumor specimens from DLBCL patients who received cyclophosphamide, adriamycin, vincristine and prednisone (CHOP)-based chemotherapy, and applied a supervised learning prediction method to identify cured versus fatal or refractory disease. The algorithm classified two categories of patients with very different five-year overall survival rates (70% versus 12%). The model also effectively delineated patients within specific IPI risk categories who were likely to be cured or to die of their disease. Genes implicated in DLBCL outcome included some that regulate responses to B-cell-receptor signaling, critical serine/threonine phosphorylation pathways and apoptosis. Our data indicate that supervised learning classification techniques can predict outcome in DLBCL and identify rational targets for intervention.

Antineoplastic Combined Chemotherapy Protocols↗

Concept formation vs. logistic regression: predicting death in trauma patients.

This study compares two classification models used to predict survival of injured patients entering the emergency department. Concept formation is a machine learning technique that summarizes known examples cases in the form of a tree. After the tree is constructed, it can then be used to predict the classification of new cases. Logistic regression, on the other hand, is a statistical model that allows for a quantitative relationship for a dichotomous event with several independent variables. The outcome (dependent) variable must have only two choices, e.g. does or does not occur, alive or dead, etc. The result of this model is an equation which is then used to predict the probability of class membership of a new case. The two models were evaluated on a trauma registry database composed of information on all trauma patients admitted in 1992 to a Level I trauma center. A total of 2155 records. representing all trauma patients admitted for more than 24 h or who died in the Emergency Department, were grouped into two databases as follows: (1) discharge status of 'died' (containing 151 records), and (2) any discharge status other than 'died' (containing 2004 records). Both databases contained the same variables.

Artificial Intelligence↗

Patient-specific models for predicting the outcomes of patients with community acquired pneumonia.

We investigated two patient-specific and four population-wide machine learning methods for predicting dire outcomes in community acquired pneumonia (CAP) patients. Predicting dire outcomes in CAP patients can significantly influence the decision about whether to admit the patient to the hospital or to treat the patient at home. Population-wide methods induce models that are trained to perform well on average on all future cases. In contrast, patient-specific methods specifically induce a model for a particular patient case. We trained the models on a set of 1601 patient cases and evaluated them on a separate set of 686 cases. One patient-specific method performed better than the population-wide methods when evaluated within a clinically relevant range of the ROC curve. Our study provides support for patient-specific methods being a promising approach for making clinical predictions.

Algorithms↗

Linking MRI radiomics to transcriptomics-based radiosensitivity in lower-grade glioma: A radiogenomic framework.

BACKGROUND: RSI is a transcriptomics-based biomarker associated with radiotherapy outcomes, but its clinical application is constrained by the requirement for tumor tissue and RNA sequencing. This study investigates whether MRI-derived radiomic features can reflect RSI-defined intrinsic radiosensitivity in lower-grade glioma.This addresses a critical gap arising from the limited availability of matched imaging and genomic data in routine clinical practice. METHODS: MRI-derived radiomic features were extracted from FLAIR images of lower-grade glioma patients obtained from TCIA and matched with transcriptomic data from TCGA. A total of 107 patients with both MRI and RNA sequencing data were included in the radiogenomic analysis. Radiomic features were ranked using a Borda-based ensemble feature selection strategy. Five supervised machine-learning classifiers were trained to predict RSI-based radiosensitivity classification, and model interpretability was assessed using SHAP within radiogenomic framework. RESULTS: Classification performance increased with feature number and stabilized at compact subset of 13 radiomic features. Logistic regression showed stable performance with an AUC of 0.82 (95 % CI: 0.71-0.93). SHAP analysis indicated that heterogeneity-related texture features were dominant contributors to model predictions, with many associated with the RR phenotype, while others were linked to the RS phenotype. CONCLUSION: An MRI-based radiomic signature enables non-invasive prediction of RSI-defined radiosensitivity in lower-grade glioma. Rather than offering an immediately deployable clinical tool, this study establishes a proof-of-concept radiogenomic framework demonstrating that intrinsic radiosensitivity, traditionally assessed through invasive molecular assays, can be approximated using quantitative imaging features. These findings highlight the potential of imaging-based radiosensitivity assessment and provide a foundation for future radiogenomic investigations.

Lower-grade glioma↗

Sample handling for mass spectrometric proteomic investigations of human sera.

Proteomic investigations of sera are potentially of value for diagnosis, prognosis, choice of therapy, and disease activity assessment by virtue of discovering new biomarkers and biomarker patterns. Much debate focuses on the biological relevance and the need for identification of such biomarkers while less effort has been invested in devising standard procedures for sample preparation and storage in relation to model building based on complex sets of mass spectrometric (MS) data. Thus, development of standardized methods for collection and storage of patient samples together with standards for transportation and handling of samples are needed. This requires knowledge about how sample processing affects MS-based proteome analyses and thereby how nonbiological biased classification errors are avoided. In this study, we characterize the effects of sample handling, including clotting conditions, storage temperature, storage time, and freeze/thaw cycles, on MS-based proteomics of human serum by using principal components analysis, support vector machine learning, and clustering methods based on genetic algorithms as class modeling and prediction methods. Using spiking to artificially create differentiable sample groups, this integrated approach yields data that--even when working with sample groups that differ more than may be expected in biological studies--clearly demonstrate the need for comparable sampling conditions for samples used for modeling and for the samples that are going into the test set group. Also, the study emphasizes the difference between class prediction and class comparison studies as well as the advantages and disadvantages of different modeling methods.

Humans↗

A leakage-aware genomic prediction pipeline for meropenem resistance in Klebsiella pneumoniae using transformer-based resistome representation learning.

MOTIVATION: Antimicrobial resistance (AMR) in Klebsiella pneumoniae, particularly to carbapenems such as meropenem, is a major global health problem. Machine learning is increasingly used to predict resistance from genomic markers; however, many models fail to capture high-level gene-gene interactions and may exhibit inflated performance due to lineage-biased prediction. Existing genomic prediction models largely rely on flat feature representations that fail to capture epistatic gene interactions, and commonly suffer from inflated performance estimates due to phylogenetic data leakage. To address these limitations simultaneously, a leakage-aware hybrid TabTransformer-CatBoost pipeline was developed, combining self-attention-based resistome representation learning with gradient boosting classification under clade-aware data partitioning. A self-attention encoder converts sparse gene presence-absence profiles into contextualized latent embeddings, which are subsequently classified using gradient boosting to capture lineage-aware AMR patterns. RESULTS: The proposed architecture outperformed classical baselines including Logistic Regression, Random Forest, XGBoost, and optimized CatBoost models. Internal accuracy reached 92.59% for the Chained Hybrid configuration (area under the receiver operating characteristic curve, AUROC = 0.8670, F1 = 0.8537). Performance gains primarily originated from the embedding stage, as confirmed by ablation analysis. External validation across independent multinational cohorts (n = 305) demonstrated generalizability (AUROC = 0.8105; F1 = 0.7552). Permutation testing produced near-zero Matthews Correlation Coefficient (MCC) = 0.0091, indicating predictions reflect genuine biological signal rather than noise. These results establish attention-based genomic embedding with gradient boosting as a scalable, interpretable, and leakage-aware framework for clinical AMR prediction. AVAILABILITY AND IMPLEMENTATION: The source code for the TabTransformer-CatBoost framework, including preprocessing pipelines and pre-trained embeddings, is available at https://github.com/SibelKervanci/kp-meropenem-tabtransformer.

Journal Article↗

Experiments with AdaBoost.RT, an improved boosting scheme for regression.

The application of boosting technique to regression problems has received relatively little attention in contrast to research aimed at classification problems. This letter describes a new boosting algorithm, AdaBoost.RT, for regression problems. Its idea is in filtering out the examples with the relative estimation error that is higher than the preset threshold value, and then following the AdaBoost procedure. Thus, it requires selecting the suboptimal value of the error threshold to demarcate examples as poorly or well predicted. Some experimental results using the M5 model tree as a weak learning machine for several benchmark data sets are reported. The results are compared to other boosting methods, bagging, artificial neural networks, and a single M5 model tree. The preliminary empirical comparisons show higher performance of AdaBoost.RT for most of the considered data sets.

Algorithms↗

Demonstration of two novel methods for predicting functional siRNA efficiency.

BACKGROUND: siRNAs are small RNAs that serve as sequence determinants during the gene silencing process called RNA interference (RNAi). It is well know that siRNA efficiency is crucial in the RNAi pathway, and the siRNA efficiency for targeting different sites of a specific gene varies greatly. Therefore, there is high demand for reliable siRNAs prediction tools and for the design methods able to pick up high silencing potential siRNAs. RESULTS: In this paper, two systems have been established for the prediction of functional siRNAs: (1) a statistical model based on sequence information and (2) a machine learning model based on three features of siRNA sequences, namely binary description, thermodynamic profile and nucleotide composition. Both of the two methods show high performance on the two datasets we have constructed for training the model. CONCLUSION: Both of the two methods studied in this paper emphasize the importance of sequence information for the prediction of functional siRNAs. The way of denoting a bio-sequence by binary system in mathematical language might be helpful in other analysis work associated with fixed-length bio-sequence.

Algorithms↗

Advancing precision tacrolimus therapy: a systems genetics dissection in BXD platform.

BACKGROUND: Tacrolimus is a core immunosuppressant in organ transplantation, but its narrow therapeutic window and significant pharmacokinetic variability hinder precision dosing. Although CYP3A5-guided strategies have established clinical relevance for tacrolimus initial dose adjustment, they do not fully account for the marked interindividual variability in tacrolimus exposure, highlighting the need for complementary models to decode more complex genetic regulation. This study aimed to identify candidate genetic modulators of tacrolimus metabolism and develop an integrated predictive framework for individualized therapy. METHODS: Using 46 BXD recombinant inbred mouse strains, we characterized transcriptomics and machine learning, and validated key genes. We then constructed a clinical model using data from 168 renal transplant recipients. RESULTS: We identified 19 genomic loci associated with tacrolimus pharmacokinetic traits and supported DBP/CYP2A6 as candidate modulators associated with tacrolimus disposition. The clinical prediction model, incorporating these genes and clinical variables, achieved robust AUROC. CONCLUSIONS: These findings support a polygenic contribution to tacrolimus metabolism and provide an experimental and computational framework for identifying candidate modulators relevant to individualized dosing. The BXD mouse platform offers a systems-genetics approach for mechanistic discovery that may inform future translational studies on tacrolimus precision dosing.

Animals↗

Flow rate of some pharmaceutical diluents through die-orifices relevant to mini-tableting.

The effects of cylindrical orifice length and diameter on the flow rate of three commonly used pharmaceutical direct compression diluents (lactose, dibasic calcium phosphate dihydrate and pregelatinised starch) were investigated, besides the powder particle characteristics (particle size, aspect ratio, roundness and convexity) and the packing properties (true, bulk and tapped density). Flow rate was determined for three different sieve fractions through a series of miniature tableting dies of different orifice diameter (0.4, 0.3 and 0.2 cm) and thickness (1.5, 1.0 and 0.5 cm). It was found that flow rate decreased with the increase of the orifice length for the small diameter (0.2 cm) but for the large diameter (0.4 cm) was increased with the orifice length (die thickness). Flow rate changes with the orifice length are attributed to the flow regime (transitional arch formation) and possible alterations in the position of the free flowing zone caused by pressure gradients arising from the flow of self-entrained air, both above the entrance in the die orifice and across it. Modelling by the conventional Jones-Pilpel non-linear equation and by two machine learning algorithms (lazy learning, LL, and feed-forward back-propagation, FBP) was applied and predictive performance of the fitted models was compared. It was found that both FBP and LL algorithms have significantly higher predictive performance than the Jones-Pilpel non-linear equation, because they account both dimensions of the cylindrical die opening (diameter and length). The automatic relevance determination for FBP revealed that orifice length is the third most influential variable after the orifice diameter and particle size, followed by the bulk density, the difference between bulk and tapped densities and the particle convexity.

Algorithms↗

In silico prediction of pregnane X receptor activators by machine learning approaches.

Pregnane X receptor (PXR) regulates drug metabolism and is involved in drug-drug interactions. Prediction of PXR activators is important for evaluating drug metabolism and toxicity. Computational pharmacophore and quantitative structure-activity relationship models have been developed for predicting PXR activators. Because of the structural diversity of PXR activators, more efforts are needed for exploring methods applicable to a broader spectrum of compounds. We explored three machine learning methods (MLMs) for predicting PXR activators, which were trained and tested by using significantly higher number of compounds, 128 PXR activators (98 human) and 77 PXR non-activators, than those of previous studies. The recursive feature-selection method was used to select molecular descriptors relevant to PXR activator prediction, which are consistent with conclusions from other computational and structural studies. In a 10-fold cross-validation test, our MLM systems correctly predicted 81.2 to 84.0% of PXR activators, 80.8 to 85.0% of hPXR activators, 61.2 to 70.3% of PXR nonactivators, and 67.7 to 73.6% of hPXR nonactivators. Our systems also correctly predicted 73.3 to 86.7% of 15 newly published hPXR activators. MLMs seem to be useful for predicting PXR activators and for providing clues to physicochemical features of PXR activation.

Artificial Intelligence↗

A compression-based approach for coding sequences identification. I. Application to prokaryotic genomes.

Most of the gene prediction algorithms for prokaryotes are based on Hidden Markov Models or similar machine-learning approaches, which imply the optimization of a high number of parameters. The present paper presents a novel method for the classification of coding and non-coding regions in prokaryotic genomes, based on a suitably defined compression index of a DNA sequence. The main features of this new method are the non-parametric logic and the costruction of a dictionary of words extracted from the sequences. These dictionaries can be very useful to perform further analyses on the genomic sequences themselves. The proposed approach has been applied on some prokaryotic complete genomes, obtaining optimal scores of correctly recognized coding and non-coding regions. Several false-positive and false-negative cases have been investigated in detail, which have revealed that this approach can fail in the presence of highly structured coding regions (e.g., genes coding for modular proteins) or quasi-random non-coding regions (e.g., regions hosting non-functional fragments of copies of functional genes; regions hosting promoters or other protein-binding sequences). We perform an overall comparison with other gene-finder software, since at this step we are not interested in building another gene-finder system, but only in exploring the possibility of the suggested approach.

Algorithms↗

Genetic mapping and predictive modeling of paralog synthetic lethality.

Paralogs are abundant in the human genome and thought to be a primary source of synthetic lethality, yet the vast paralogome remains largely uncharacterized. A digenic screen of 36,648 paralogous pairs in the human genome revealed that synthetic lethalities were infrequent and varied in penetrance in different tumor backgrounds. We hypothesized that the variable penetrance of synthetic lethalities resulted from complex polygenic interactions with different cellular contexts. A machine learning classifier of a subset of paralog pairs tested across 49 cancer models revealed that endogenous perturbations in related pathways predicted paralog synthetic lethality. Further, predictive modeling of paralog synthetic lethality showed that the strength of synthetic lethal interactions was largely due to the overlap and essentiality of the protein-protein interaction networks shared by the paralog pairs. Collectively, this study tested 36,648 digenic paralog interactions and delineated the key feature classes that underlie the heterogeneity of paralog synthetic lethalities.

Humans↗

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype‑dependent opioid consumption over 72 h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non‑carriers, despite reporting similar subjective pain scores. This consistent genotype‑dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3↗

Transcriptome-based high-frequency recurrence index predicts frequent recurrence in non-muscle-invasive bladder cancer after Bacillus Calmette-Guérin therapy.

BACKGROUND: High-frequency recurrence (HfR,&#x2009;&#x2265;&#x2009;2 recurrences) in non-muscle-invasive bladder cancer (NMIBC) poses a significant clinical burden. Current risk models, such as the European Organization for Research and Treatment of Cancer (EORTC), the European Association of Urology (EAU), and the UROMOL classification, offer limited predictive accuracy for identifying patients at risk for frequent recurrence despite appropriate treatment. METHODS: A 75-gene high-frequency recurrence index (HfRI) was constructed by selecting recurrence-associated genes using differential expression and Cox regression analyses. The HfRI was computed as a weighted sum of normalized gene expression values. The model was trained on a discovery cohort and validated in multiple cohorts (n&#x2009;=&#x2009;1379) using machine-learning approaches. Clinical relevance was assessed using recurrence-free survival (RFS) and Cox models, and predictive performance was compared with that of the EORTC, EAU, and UROMOL classifications using the area under the curve (AUC) and the concordance index (c-index). RESULTS: The HfRI robustly stratified patients into high-risk and low-risk groups across six independent NMIBC cohorts. Patients classified as HfRI-high had a significantly greater likelihood of experiencing&#x2009;&#x2265;&#x2009;2 recurrences (&#x3c7;2, p&#x2009;=&#x2009;0.001) and showed markedly reduced RFS (log-rank test, p&#x2009;<&#x2009;0.001). The adverse prognostic effect of the HfRI persisted even among patients treated with BCG therapy (log-rank test, p&#x2009;=&#x2009;0.02). Multivariate analysis revealed that the HfRI was an independent predictor of HfR (HR&#x2009;=&#x2009;2.82, 95% CI&#x2009;=&#x2009;1.89-4.20, p&#x2009;<&#x2009;0.001). Compared with established clinical risk classifiers, the HfRI demonstrated superior predictive performance (AUC&#x2009;=&#x2009;0.736, c-index&#x2009;=&#x2009;0.673) in terms of the EORTC (AUC&#x2009;=&#x2009;0.594), EAU (AUC&#x2009;=&#x2009;0.557) risk groups, and UROMOL2021 (AUC&#x2009;=&#x2009;0.596) classification. Pathway analysis revealed that HfRI-high tumors were characterized by upregulation of cell cycle progression and DNA replication pathways, accompanied by suppression of immune signaling pathways. These biological features provide a mechanistic explanation for the reduced responsiveness to intravesical BCG therapy, underscoring the role of HfRI not only as a predictor of recurrence risk but also as a biomarker capable of identifying patients unlikely to benefit from standard BCG treatment. CONCLUSIONS: HfRI represents a robust, transcriptome-based tool for predicting frequent recurrence in NMIBC patients. The HfRI supports earlier identification of patients at risk of high-frequency recurrence, thereby supporting personalized treatment strategies.

Humans↗

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine↗