Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Integrative multi-omics and single-cell analysis identifies EGFR pathway activation and metabolic reprogramming as potential synthetic lethal vulnerabilities in resistance to the FGFR inhibitor AZD4547.

BACKGROUND: Although fibroblast growth factor receptor (FGFR) inhibitors (FGFRi) have demonstrated clinical promise, the inevitable emergence of acquired resistance remains a critical bottleneck, severely compromising their long-term clinical efficacy. The pan-cancer molecular landscape and heterogeneous mechanisms driving this resistance, ranging from genetic alterations to dynamic network rewiring, remain poorly understood. METHODS: We integrated large-scale pharmacogenomic profiling of the FGFR inhibitor AZD4547 from the GDSC2 and PRISM databases with single-cell RNA sequencing to dissect the multi-omics landscape of FGFRi resistance across 312 cell lines from 8 cancer types. This multi-omics framework was further extended by machine learning modeling and systematic synthetic lethality screening to uncover actionable therapeutic targets. In vitro viability assays and western blot analysis were subsequently conducted to experimentally evaluate the predicted FGFR-EGFR synthetic lethality. RESULTS: Our dual-database analysis unveiled a multi-dimensional atlas of FGFRi resistance. We identified cancer-specific genomic drivers, such as ELF4 amplification in glioblastoma, alongside key transcriptomic markers including UCP2 and FSCN1, highlighting a shift towards metabolic reprogramming and epithelial-mesenchymal transition (EMT). Single-cell analysis unveiled that resistance is linked to the heterogeneous enrichment of baseline subpopulations characterized by distinct metaprograms, including cell-cycle dysregulation. Furthermore, a random forest model built on a LASSO-derived transcriptomic signature was constructed, demonstrating promising predictive capability for AZD4547 sensitivity (mean test-set AUC = 0.73, 95% CI [0.63, 0.80]); the signature generalized well to erdafitinib but showed limited transferability to some other FGFR inhibitors (e.g. pemigatinib, BGJ398). Most notably, our synthetic lethal screening revealed a convergent reliance on compensatory RTK signaling (specifically EGFR pathway enrichment) and downstream MAPK/PI3K cascades in resistant phenotypes, providing converging computational evidence for EGFR pathway activation as an adaptive bypass mechanism. This predicted synthetic lethality was experimentally supported in two FGFR-dependent cell line models (RT112 and CCLP1), in which combined FGFR-EGFR inhibition produced marked synergistic antiproliferative effects. CONCLUSIONS: This study establishes a comprehensive multi-omics atlas of resistance to the FGFR inhibitor AZD4547, delineating convergent mechanisms of metabolic reprogramming and EGFR-mediated bypass signaling. Our findings characterize the resistance as a dynamic network rewiring and nominate rational combination strategies to overcome this therapeutic bottleneck. While FGFR-EGFR co-inhibition is experimentally supported, metabolic co-targeting remains a computationally derived, hypothesis-generating strategy.

Benzamides

Integrated analysis of plasma metabolomics and proteomics reveals the biological characteristics of damp-heat and stasis-toxin syndrome in colorectal cancer.

OBJECTIVE: To investigate the biological attributes of core syndromes in colorectal cancer, namely, the damp-heat and stasis-toxin syndrome (SRYD). METHODS: Between October 2021 and October 2022, a cohort comprising 40 patients with colorectal cancer (CRC) diagnosed with damp-heat and stasis-toxin syndrome (SRYD group), 40 patients with CRC without this syndrome (non-SRYD group), and 40 healthy controls (Normal group) was recruited at Jiangsu Province Hospital of Chinese Medicine. Untargeted metabolomics analysis was conducted on plasma samples from all 120 participants, while differential protein analysis using four-dimensional data-independent acquisition proteomics was performed on 20 randomly selected samples per group. A combined analysis of proteomics and metabolomics data followed, and the identified potential diagnostic biomarkers were subsequently used to train and validate multiple machine learning models. RESULTS: Proteomic analysis revealed 130 differential proteins in the colorectal cancer with damp-heat and stasis-toxin syndrome (CRC-SRYD) group, enriched in pathways including complement and coagulation cascades, as well as nuclear factor kappa-B (NF-κB) signaling. Metabolomic analysis identified 584 differential metabolites within the same group, showing enrichment in pathways such as primary bile acid biosynthesis, central carbon metabolism in cancer, and glucagon signaling. Integrated pathway analysis indicated heightened activity of the NF-κB signaling pathway in the CRC-SRYD group. A biomarker panel, comprising 6 proteins and 9 metabolites selected through the ReliefF algorithm, was used to construct a diagnostic model with random forest, achieving an accuracy of 93.33%, sensitivity of 80.00%, and specificity of 100%. CONCLUSION: This study systematically elucidates plasma metabolomic and proteomic alterations in patients with CRC, establishing a robust diagnostic model for CRC syndrome (CRC-SRYD). Further investigation is warranted to clarify the underlying molecular mechanisms and biological foundations.

Humans

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer

Long COVID in Elderly COPD Patients: Clinical Features, Pulmonary Function Decline, and Proteomic Insights.

BACKGROUND: Elderly patients with chronic obstructive pulmonary disease (COPD) face a heightened risk of developing long coronavirus disease (COVID); however the exact clinical characteristics and underlying mechanisms remain unclear. METHODS: We enrolled 85 elderly COPD patients, of whom 43 reported newly onset persistent fatigue (the most dominant complaint of long COVID) within 1 year after severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, and they were allocated to the Long-COVID group. The remaining 42 patients were assigned to the Control group. Patients completed questionnaires, pulmonary function tests, chest CT, routine laboratory tests, and blood proteomic analysis. RESULTS: Long-COVID patients had a longer course of COPD (> 5 years, 76.8% vs 52.4%) and duration of SARS-CoV-2 infection (10.0 days vs 7.0 days) (All P < 0.05), higher symptom burden, worse pulmonary ventilation function and a more rapid decrease in DLCO (All P < 0.05). Proteomic analysis indicated disruptions in inflammation and energy metabolism, potentially underlying long COVID in these patients. The machine learning model identified wheezing, the duration of SARS-CoV-2 infection, EIF2S3 (eukaryotic translation initiation factor 2 subunit gamma), current FEV1/FVC (%), and the course of COPD as key features distinguishing Long-COVID patients, and exhibited excellent performance. CONCLUSION: Elderly COPD patients with a longer COPD course and duration of COVID-19 are more prone to develop long COVID, with decreased pulmonary ventilation and diffusion ability. Disordered inflammation regulation and energy metabolism may be the potential mechanisms, highlighting the importance of monitoring inflammation and metabolic dysregulation in elderly COPD patients after recovery from COVID-19.

Humans

Evidence for synaptic plasticity in the cerebellar cortex.

The learning machine model of the cerebellum by Marr and Albus contains a special type of synaptic plasticity. Experimental evidence for this synaptic plasticity has been meager, but very recently positive evidence has become available. Ito, Sakurai and Tongroach (7) demonstrated the occurrence of a long-lasting depression in mossy fiber responsiveness of Purkinje cells subsequent to conjuctive stimulation of mossy fibers and climbing fibers. A similar long-lasting depression was shown to occurred in sensitivity of Purkinje cell dendrites to a putative neurotransmitter of parallel fibers, i.e., glutamate. Furthermore, Ito and Kano (6) produced a long-lasting depression in the molecular layer of the cerebellar cortex by simultaneous direct stimulation of parallel fibers and climbing fibers. These long-lasting depressions appear to represent a synaptic plasticity of the form proposed by Albus.

Animals

Automated Machine Learning Tools to Build Regression Models for Schizosaccharomyces pombe Omics Data.

Machine learning is a powerful tool for analyzing biological data and making useful predictions. The surge of biological data from high-throughput omics technologies has raised the need for modeling approaches capable of tackling such amounts of data, which is pivotal to understanding the nature of complex molecular systems. Here, we show how to construct a simple model using automated machine learning (AutoML) to predict protein abundance in Schizosaccharomyces pombe, using data obtained from codon usage bias and quantitative proteomics.

Machine Learning

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics&#x2013;based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics&#x2013;based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

PredIL13: Stacking a variety of machine and deep learning methods with ESM-2 language model for identifying IL13-inducing peptides.

Interleukin (IL)-13 has emerged as one of the recently identified cytokine. Since IL-13 causes the severity of COVID-19 and alters crucial biological processes, it is urgent to explore novel molecules or peptides capable of including IL-13. Computational prediction has received attention as a complementary method to in-vivo and in-vitro experimental identification of IL-13 inducing peptides, because experimental identification is time-consuming, laborious, and expensive. A few computational tools have been presented, including the IL13Pred and iIL13Pred. To increase prediction capability, we have developed PredIL13, a cutting-edge ensemble learning method with the latest ESM-2 protein language model. This method stacked the probability scores outputted by 168 single-feature machine/deep learning models, and then trained a logistic regression-based meta-classifier with the stacked probability score vectors. The key technology was to implement ESM-2 and to select the optimal single-feature models according to their absolute weight coefficient for logistic regression (AWCLR), an indicator of the importance of each single-feature model. Especially, the sequential deletion of single-feature models based on the iterative AWCLR ranking (SDIWC) method constructed the meta-classifier consisting of the top 16 single-feature models, named PredIL13, while considering the model's accuracy. The PredIL13 greatly outperformed the-state-of-the-art predictors, thus is an invaluable tool for accelerating the detection of IL13-inducing peptide within the human genome.

Humans

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans

Integrated single-cell transcriptomics, Mendelian randomization, and machine learning identify CEBPZ as an immune-related biomarker in oral lichen planus.

BACKGROUND: Oral lichen planus (OLP) is a chronic, immune-mediated oral mucosal disease with complex pathophysiology and potential for malignant transformation. Understanding its molecular basis is critical for the development of precise diagnostic and therapeutic strategies. OBJECTIVES: We aimed to identify key immune-related biomarkers and characterize cellular dynamics in OLP, with a particular focus on the role of CEBPZ in disease pathogenesis. MATERIAL AND METHODS: We analyzed single-cell RNA sequencing (scRNA-seq) data from OLP lamina propria samples (GSE211630) to identify disease-specific T-cell subpopulations using high-dimensional weighted gene co-expression network analysis (hdWGCNA) for oxidative stress-related gene modules.-data-based Mendelian randomization (SMR) integrated FinnGen genome-wide association study (GWAS; 342,499 Europeans) data with Genotype-Tissue Expression (GTEx) expression quantitative trait loci (eQTL) data to identify causal genes. Machine learning (ML) models (least absolute shrinkage and selection operator (LASSO) and convolutional neural network (CNN)) were developed using bulk RNA-seq datasets (GSE52130 and GSE38616) for diagnostic purposes. RESULTS: We identified OLP-specific T-cell populations (clusters 0, 3, 5, 7, 13, and 15) with enhanced migration inhibition factor (MIF) pathway signaling toward B cells and monocytes. Two oxidative stress-associated modules contained hub genes, including CEBPZ. Summary-data-based Mendelian randomization analysis identified 231 OLP-associated genes, with CEBPZ uniquely intersecting LASSO-selected markers (odds ratio (OR) = 1.057, 95% confidence interval (95% CI) = 1.013-1.102, p = 0.010). Machine learning models achieved area under the curve (AUC) values ranging from 0.653 to 0.745, with the CNN model reaching a validation accuracy of 0.735. CEBPZ showed elevated expression in OLP T cells and correlated with enhanced MIF-(CD74+CXCR4) signaling. CONCLUSIONS: This integrative approach identifies CEBPZ as a pivotal biomarker linking genetic susceptibility, oxidative stress, and immune dysregulation in OLP. Our diagnostic models offer promising tools for OLP management.

CEBPZ

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans

Inferring metabolic objectives and trade-offs in single cells during embryogenesis.

While proliferating cells optimize their metabolism to produce biomass, the metabolic objectives of cells that perform non-proliferative tasks are unclear. The opposing requirements for optimizing each objective result in a trade-off that forces single cells to prioritize their metabolic needs and optimally allocate limited resources. Here, we present single-cell optimization objective and trade-off inference (SCOOTI), which infers metabolic objectives and trade-offs in biological systems by integrating bulk and single-cell omics data, using metabolic modeling and machine learning. We validated SCOOTI by identifying essential genes from CRISPR-Cas9 screens in embryonic stem cells, and by inferring the metabolic objectives of quiescent cells, during different cell-cycle phases. Applying this to embryonic cell states, we observed a decrease in metabolic entropy upon development. We further uncovered a trade-off between glutathione and biosynthetic precursors in one-cell zygote, two-cell embryo, and blastocyst cells, potentially representing a trade-off between pluripotency and proliferation. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis

MyESL: A Software for Evolutionary Sparse Learning in Molecular Phylogenetics and Genomics.

Evolutionary sparse learning uses supervised machine learning to build evolutionary models where genomic sites loci are parameters. It uses the Least Absolute Shrinkage and Selection Operator with bi-level sparsity to connect a specific phylogenetic hypothesis with sequence variation across genomic loci. The MyESL software addresses the need for open-source tools to perform evolutionary sparse learning analyses, offering features to preprocess input phylogenomic alignments, post-process output models to generate molecular evolutionary metrics, and make Least Absolute Shrinkage and Selection Operator regression adaptable and efficient for phylogenetic trees and alignments. The core of MyESL, which constructs models with logistic regressions using bi-level sparsity, is written in C++. Its input data preprocessing and result post-processing tools are developed in Python. Compared to other tools, MyESL is more computationally efficient and provides evolution-friendly inputs and outputs. These features have already enabled the use of MyESL in two phylogenomic applications, one to identify outlier sequences and fragile clades in inferred phylogenies and another to build genetic models of convergent traits. In addition to the use in a Python environment, MyESL is available as a standalone executable compatible across multiple platforms, which can be directly integrated into scripts and third-party software. The source code, executable, and documentation for MyESL are openly accessible at https://github.com/kumarlabgit/MyESL.

Phylogeny

Integrated transcriptome analysis and machine learning to construct a homeostatic model of acetylation for bladder cancer and validate the key gene CES1.

BACKGROUND: Bladder cancer (BLCA) is one of the most common malignant tumors of the urinary system. Protein acetylation (PA) plays a critical role in regulating multiple biological processes (BPs), cellular homeostasis, and cancer-related signaling pathways. This study aimed to construct a homeostatic model of acetylation for BLCA using integrated transcriptome analysis and machine learning and to validate the key gene CES1. METHODS: RNA sequencing (RNA-seq) and clinical data were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases. Acetylation-related differentially expressed genes (DEGs) in BLCA were screened using differential expression analysis (DEA). An acetylation homeostatic model was constructed via univariate, machine learning-based least absolute shrinkage and selection operator (LASSO) and multivariate Cox regression analyses, followed by validation in multiple cohorts. Single-cell RNA-seq analysis was used to explore gene expression patterns in diverse cell types. Enrichment analysis (EA), immune infiltration, and drug sensitivity analysis (DSA) were performed to characterize molecular features of different risk groups. Finally, the biological function of CES1 as the key gene was verified by in vitro knockdown experiments. RESULTS: We established a robust acetylation homeostatic model consisting of five genes, which effectively predicted overall survival (OS) and served as an independent prognostic factor in BLCA. High-risk patients showed significantly poorer prognosis, distinct immune infiltration profiles, and differential drug sensitivity. CES1 was identified and validated as the key gene in this model, which was highly expressed in BLCA and associated with poor prognosis. Knockdown of CES1 markedly suppressed cell proliferation, invasion, and migration, and reduced intracellular coenzyme A (CoA) levels, thereby regulating PA homeostasis. CONCLUSIONS: We developed and validated a novel acetylation homeostatic model for survival stratification and personalized treatment guidance in BLCA, based on integrated transcriptome analysis and machine learning. CES1 is closely associated with intracellular CoA levels and the malignant progression of BLCA. Its potential association with PA homeostasis requires further mechanistic validation, and it may act as a candidate therapeutic biomarker for BLCA.

Bladder cancer (BLCA)

Drug design by machine learning: the use of inductive logic programming to model the structure-activity relationships of trimethoprim analogues binding to dihydrofolate reductase.

The machine learning program GOLEM from the field of inductive logic programming was applied to the drug design problem of modeling structure-activity relationships. The training data for the program were 44 trimethoprim analogues and their observed inhibition of Escherichia coli dihydrofolate reductase. A further 11 compounds were used as unseen test data. GOLEM obtained rules that were statistically more accurate on the training data and also better on the test data than a Hansch linear regression model. Importantly machine learning yields understandable rules that characterized the chemistry of favored inhibitors in terms of polarity, flexibility, and hydrogen-bonding character. These rules agree with the stereochemistry of the interaction observed crystallographically.

Artificial Intelligence

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation