Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.

MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.

CRISPR-Cas Systems↗

Application of machine learning to structural molecular biology.

A technique of machine learning, inductive logic programming implemented in the program GOLEM, has been applied to three problems in structural molecular biology. These problems are: the prediction of protein secondary structure; the identification of rules governing the arrangement of beta-sheets strands in the tertiary folding of proteins; and the modelling of a quantitative structure activity relationship (QSAR) of a series of drugs. For secondary structure prediction and the QSAR, GOLEM yielded predictions comparable with contemporary approaches including neural networks. Rules for beta-strand arrangement are derived and it is planned to contrast their accuracy with those obtained by human inspection. In all three studies GOLEM discovered rules that provided insight into the stereochemistry of the system. We conclude machine learning used together with human intervention will provide a powerful tool to discover patterns in biological sequences and structures.

Amino Acid Sequence↗

Prediction of the isoelectric point of an amino acid based on GA-PLS and SVMs.

The support vector machine (SVM), as a novel type of a learning machine, for the first time, was used to develop a QSPR model that relates the structures of 35 amino acids to their isoelectric point. Molecular descriptors calculated from the structure alone were used to represent molecular structures. The seven descriptors selected using GA-PLS, which is a sophisticated hybrid approach that combines GA as a powerful optimization method with PLS as a robust statistical method for variable selection, were used as inputs of RBFNNs and SVM to predict the isoelectric point of an amino acid. The optimal QSPR model developed was based on support vector machines, which showed the following results: the root-mean-square error of 0.2383 and the prediction correlation coefficient R=0.9702 were obtained for the whole data set. Satisfactory results indicated that the GA-PLS approach is a very effective method for variable selection, and the support vector machine is a very promising tool for the nonlinear approximation.

Amino Acids↗

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic↗

Harnessing deep learning for proteome-scale detection of amyloid signaling motifs.

MOTIVATION: Amyloid signaling sequences adopt the cross-β fold that is capable of self-replication in the templating process. Propagation of the amyloid fold from the receptor to the effector protein is used for signal transduction in the immune response pathways in animals, fungi, and bacteria. So far, a dozen of families of amyloid signaling motifs (ASMs) have been classified. Unfortunately, due to the wide variety of ASMs it is difficult to identify them in large protein databases available, which limits the possibility of conducting experimental studies. To date, various deep learning (DL) models have been applied across a range of protein-related tasks, including domain family classification and the prediction of protein structure and protein-protein interactions. RESULTS: In this study, we develop tailor-made bidirectional LSTM and BERT-based architectures to model ASM, and compare their performance against a state-of-the-art machine learning grammatical model. Our research is focused on developing a discriminative model of generalized ASMs, capable of detecting ASMs in large datasets. The DL-based models are trained on a diverse set of motif families and a global negative set, and used to identify ASMs from remotely related families. We analyze how both models represent the data and demonstrate that the DL-based approaches effectively detect ASMs, including novel motifs, even at the genome scale. AVAILABILITY AND IMPLEMENTATION: The models are provided as a Python package, asmscan-bilstm, and a Docker image at https://github.com/chrispysz/asmscan-proteinbert-run. The source code can be accessed at https://github.com/jakub-galazka/asmscan-bilstm and https://github.com/chrispysz/asmscan-proteinbert. Data and results are at https://github.com/wdyrka-pwr/ASMscan.

Deep Learning↗

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation↗

Comparison of classic statistical methods and machine learning approaches to classify readiness.

MOTIVATION: Predicting physical and cognitive readiness in warfighters is critical for mission success. These predictions can be improved by identifying key biomarkers using multiple omics modalities. The MASTR-E study conducted by McKetney and colleagues is one of the most comprehensive multi-omics studies of saliva samples collected from warfighters, which also applied classic linear statistical (CLS) techniques to discover key biomarkers of readiness. Aligning with McKetney et al.'s assumptions, we operationalize readiness as a binary proxy, where pre-mission samples are labeled as "ready" to reflect a rested, unstressed physiological baseline, while post-mission samples are labeled "not ready" to reflect cumulative physical and cognitive load from the mission. As such, readiness here is not a direct biological or physiological construct, but an inferred state likely dominated by stress-related physiological changes. This assumption and definition is discussed further in the Introduction and Limitations sections. Here, we apply machine learning (ML) analyses to better assess generalizability, consider hidden interactions, and identify nonlinear patterns in the data. We investigated whether ML approaches could predict readiness and identify relevant biomarkers. ML models were trained on proteomics-only or metabolomics-only datasets to classify participants as ready or not ready and important model features were considered as putative biomarkers. Training and testing datasets were curated for two objectives: (i) recognize biomolecular signatures indicative of readiness within the same donor and (ii) assess generalizability across warfighters by withholding donors for testing. RESULTS: Proteomics-based models achieved AUCs of 0.907 ± 0.034 and 0.860 ± 0.063 for Objectives 1 and 2, respectively. Metabolomics-based models achieved Objective 1 AUC of 0.994 ± 0.007 and Objective 2 AUC of 0.993 ± 0.010. Comparative analysis with existing literature validates the model's feature importances, but the identified putative biomarkers significantly differ from those discovered through CLS analyses, as only one ML-identified biomarker overlapping with those identified through CLS methods. We show that these ML models and identified features are more robust to noise and generalizable across participants than those identified using CLS methods. AVAILABILITY: The analysis pipelines are provided as Jupyter notebooks, including all code and documentation, and are available publicly on GitHub at {https://github.com/netrias/ReadinessClassification}.

Machine Learning↗

The Role of Artificial Intelligence for Intimate Partner Violence Prevention: A Systematic Review.

INTRODUCTION: Intimate partner violence (IPV), encompassing physical, sexual, emotional and economic abuse, remains a pervasive global health concern. Traditional prevention efforts face obstacles such as underreporting, delayed detection and limited personalised support. Emerging artificial intelligence (AI) approaches offer new opportunities to enhance IPV prevention. AIM: This systematic review maps and synthesises evidence on AI-driven tools in IPV prevention based on studies published between 2004 and 2024. METHODS: Following PRISMA 2020 guidelines and PROSPERO registration, we searched PubMed, Embase, CINAHL, PsycINFO, IEEE Xplore and Web of Science. Eligible studies explicitly evaluated AI technologies targeting IPV prediction, screening, intervention or support delivery. Study quality was appraised using the Mixed Methods Appraisal Tool (MMAT). RESULTS: Of 1304 records initially identified, 41 studies met eligibility criteria. AI applications ranged from machine learning (ML) for risk prediction and natural language processing (NLP) for IPV detection in clinical and social media data, to image analysis for forensic evaluation and chatbot-based support. Predictive modelling demonstrated strong discriminative performance, while NLP-based screening detected IPV with notable sensitivity. Chatbots showed feasibility and user acceptability, but evidence of their direct impact on reducing IPV incidence was limited, with one randomised controlled trial showing a modest reduction. Key challenges identified included algorithmic bias, data privacy risks and barriers to integration across health and social care systems. DISCUSSION: AI-informed interventions show promise for improving IPV detection, risk assessment, and scalable support, but questions remain about long-term effectiveness, ethical fairness, transparency and equitable implementation. Future interdisciplinary research should address these concerns to responsibly deploy AI in IPV prevention. RELEVANCE TO CLINICAL PRACTICE: The findings highlight the importance of trauma-informed, culturally responsive care and provider training in AI applications. Nurse-led innovation and policy advocacy will be crucial for safe, equitable integration of AI in IPV prevention.

Artificial Intelligence↗

Beyond data and technology: the need for new thinking to enable the era of precision prevention.

BACKGROUND: Global flagship initiatives increasingly advocate for proactive health maintenance to alleviate the growing burden on reactive, disease-focused healthcare systems. Precision prevention is conceived as the targeted modulation of causal pathways across the disease continuum, from latent risk and pre-disease states to clinical manifestation, surpassing conventional public health prevention strategies that prioritise managing population-level risk factors. Traditional discovery and implementation models, however, remain poorly aligned with the pace and breadth of scientific and technological advances. This review outlines key barriers to scaling precision prevention and argues for the integration of conceptual, methodological, and policy perspectives into a single implementation‑oriented framework. MAIN: Individualised risk stratification lies at the core of precision prevention. Genomics serves as a stable substrate for lifetime susceptibility assessment, while meaningful prediction in multifactorial chronic disease requires additional risk monitoring using dynamic intermediate molecular markers and high-resolution exposomic data. Machine learning and other artificial intelligence (AI) methods are increasingly helpful tools for integrating large, heterogeneous and temporally structured real-world data to generate personalised predictions of health trajectories. Trustworthy AI-enabled risk prediction or decision-support systems are expected to provide transparency about model logic, assumptions and performance. In discovery, existing diagnostic classifications and conventional case-control designs can obscure mechanistic heterogeneity. Shifting toward precision phenotyping and biologically grounded disease redefinition could reveal a new layer of molecular understanding. Evidence generation strategies that reflect the temporal change of disease, including high‑risk enrichment, surrogate endpoints, and adaptive, trajectory-based monitoring, are particularly important for common conditions with prolonged latency periods (e.g., cancer, cardiovascular disease). Features often dismissed as "noise", such as stochastic molecular variation and minimal exposures, may in fact encode meaningful individual-level signals and thus merit investigation. CONCLUSION: To shift healthcare from reactive treatment toward proactive health maintenance requires coordinated action from stakeholders to reshape the pillars of discovery, reform outcome assessments and modernise implementation strategies.

Humans↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets.

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Escherichia coli↗

Self-optimizing MPC of melt temperature in injection moulding.

The parameters in plastic injection moulding are highly nonlinear and interacting. Good control of plastic melt temperature for injection moulding is very important in reducing operator setup time, assuring consistent product quality, and preventing thermal degradation of the melt. Step response testing was performed on the barrel heating zones on an industrial injection moulding machine (IMM). The open loop responses indicated a high degree of process coupling between the heating zones. From these experimental step responses, a multiple-input-multiple-output model predictive control strategy was developed and practically implemented. The requirement of negligible overshoot is important to the plastics industry for preventing material overheating and wastage, and reducing machine operator setup time. A generic learning and self-optimizing MPC methodology was developed and implemented on the IMM to control melt temperature for any polymer to be moulded on any machine having different electrical heater capacities. The control performance was tested for varying setpoint trajectories typical of normal machine operations. The results showed that the predictive controller provided good control of melt temperature for all zones with negligible oscillations, and, therefore, eliminated material degradation and extended machine setup time.

Computer Simulation↗

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning↗

Development and Validation an Integrated Deep Learning Model to Assist Eosinophilic Chronic Rhinosinusitis Diagnosis: A Multicenter Study.

BACKGROUND: The assessment of eosinophilic chronic rhinosinusitis (eCRS) lacks accurate non-invasive preoperative prediction methods, relying primarily on invasive histopathological sections. This study aims to use computed tomography (CT) images and clinical parameters to develop an integrated deep learning model for the preoperative identification of eCRS and further explore the biological basis of its predictions. METHODS: A total of 1098 patients with sinus CT images were included from two hospitals and were divided into training, internal, and external test sets. The region of interest of sinus lesions was manually outlined by an experienced radiologist. We utilized three deep learning models (3D-ResNet, 3D-Xception, and HR-Net) to extract features from CT images and calculate deep learning scores. The clinical signature and deep learning score were inputted into a support vector machine for classification. The receiver operating characteristic curve, sensitivity, specificity, and accuracy were used to evaluate the integrated deep learning model. Additionally, proteomic analysis was performed on 34 patients to explore the biological basis of the model's predictions. RESULTS: The area under the curve of the integrated deep learning model to predict eCRS was 0.851 (95% confidence interval [CI]: 0.77-0.93) and 0.821 (95% CI: 0.78-0.86) in the internal and external test sets. Proteomic analysis revealed that in patients predicted to be eCRS, 594 genes were dysregulated, and some of them were associated with pathways and biological processes such as chemokine signaling pathway. CONCLUSIONS: The proposed integrated deep learning model could effectively predict eCRS patients. This study provided a non-invasive way of identifying eCRS to facilitate personalized therapy, which will pave the way toward precision medicine for CRS.

Humans↗

Singularities in mixture models and upper bounds of stochastic complexity.

A learning machine which is a mixture of several distributions, for example, a gaussian mixture or a mixture of experts, has a wide range of applications. However, such a machine is a non-identifiable statistical model with a lot of singularities in the parameter space, hence its generalization property is left unknown. Recently an algebraic geometrical method has been developed which enables us to treat such learning machines mathematically. Based on this method, this paper rigorously proves that a mixture learning machine has the smaller Bayesian stochastic complexity than regular statistical models. Since the generalization error of a learning machine is equal to the increase of the stochastic complexity, the result of this paper shows that the mixture model can attain the more precise prediction than regular statistical models if Bayesian estimation is applied in statistical inference.

Bayes Theorem↗

Granular support vector machines with association rules mining for protein homology prediction.

OBJECTIVE: Protein homology prediction between protein sequences is one of critical problems in computational biology. Such a complex classification problem is common in medical or biological information processing applications. How to build a model with superior generalization capability from training samples is an essential issue for mining knowledge to accurately predict/classify unseen new samples and to effectively support human experts to make correct decisions. METHODOLOGY: A new learning model called granular support vector machines (GSVM) is proposed based on our previous work. GSVM systematically and formally combines the principles from statistical learning theory and granular computing theory and thus provides an interesting new mechanism to address complex classification problems. It works by building a sequence of information granules and then building support vector machines (SVM) in some of these information granules on demand. A good granulation method to find suitable granules is crucial for modeling a GSVM with good performance. In this paper, we also propose an association rules-based granulation method. For the granules induced by association rules with high enough confidence and significant support, we leave them as they are because of their high "purity" and significant effect on simplifying the classification task. For every other granule, a SVM is modeled to discriminate the corresponding data. In this way, a complex classification problem is divided into multiple smaller problems so that the learning task is simplified. RESULTS AND CONCLUSIONS: The proposed algorithm, here named GSVM-AR, is compared with SVM by KDDCUP04 protein homology prediction data. The experimental results show that finding the splitting hyperplane is not a trivial task (we should be careful to select the association rules to avoid overfitting) and GSVM-AR does show significant improvement compared to building one single SVM in the whole feature space. Another advantage is that the utility of GSVM-AR is very good because it is easy to be implemented. More importantly and more interestingly, GSVM provides a new mechanism to address complex classification problems.

Algorithms↗

ANGLE: a sequencing errors resistant program for predicting protein coding regions in unfinished cDNA.

In the process of making full-length cDNA, predicting protein coding regions helps both in the preliminary analysis of genes and in any succeeding process. However, unfinished cDNA contains artifacts including many sequencing errors, which hinder the correct evaluation of coding sequences. Especially, predictions of short sequences are difficult because they provide little information for evaluating coding potential. In this paper, we describe ANGLE, a new program for predicting coding sequences in low quality cDNA. To achieve error-tolerant prediction, ANGLE uses a machine-learning approach, which makes better expression of coding sequence maximizing the use of limited information from input sequences. Our method utilizes not only codon usage, but also protein structure information which is difficult to be used for stochastic model-based algorithms, and optimizes limited information from a short segment when deciding coding potential, with the result that predictive accuracy does not depend on the length of an input sequence. The performance of ANGLE is compared with ESTSCAN on four dataset each of them having a different error rate (one frame-shift error or one substitution error per 200-500 nucleotides) and on one dataset which has no error. ANGLE outperforms ESTSCAN by 9.26% in average Matthews's correlation coefficient on short sequence dataset (< 1000 bases). On long sequence dataset, ANGLE achieves comparable performance.

Algorithms↗

Senescent fibroblasts drive CD8+ T cell dysfunction in colorectal cancer via CD36-mediated lipid transfer and peroxidation.

BACKGROUND: Functional exhaustion of tumor-infiltrating CD8+ T cells represents a hallmark of colorectal cancer (CRC) immunosuppression, though its mechanistic drivers remain elusive. Given the established correlation between CRC progression and stromal senescence characterized by pathological lipid accumulation and impaired immunity, we investigated whether and how senescent fibroblasts actively regulate CD8+ T cell dysfunction. METHODS: Single-cell RNA sequencing (scRNA-seq) analysis was conducted to unveil the diverse fibroblast populations and the significant lipid metabolism changes between senescent fibroblasts and non-senescent fibroblasts in human CRC specimens and adjacent normal mucosa. Machine-learning identified senescent fibroblasts with a distinct gene signature. Cell-cell communication analysis was used to evaluate the interactions between senescent fibroblasts and CD8+ T cells in colorectal cancer. Co-culture experiments were conducted among senescent fibroblasts, CD8+ T cells and patient-derived organoids of CRC (CRC-PDOs), with the results evaluated with high-content imaging and propidium iodide/Hoechst 33,342 staining. Flow cytometry, ELISA and lipid pulse-chase with BODIPY FL C16 were performed to detect the alterations of CD8+ T cell cytotoxic function and metabolic status. AOM/DSS-induced CRC mouse model was used to conduct in vivo validation to evaluate whether senolytics could suppress CRC progression. Patients from the Cancer Genome Atlas colorectal cancer cohort were stratified into CD36-high and CD36-low groups by median expression, and drug sensitivity for GDSC2 compounds was predicted computationally using the oncoPredict R package. RESULTS: ScRNA-seq demonstrated the specific cell population presence and divergence of senescent fibroblasts between neoplastic and histologically normal adjacent cell clusters in CRC. Random Forest was employed for cell senescence classification. Feature importance analysis identified five genes as key contributors to the model&#x2019;s decision process. Cell-cell communication analysis revealed enhanced interactions between senescent fibroblasts and CD8+ T cells in CRC. Co-culture of senescent fibroblasts significantly impaired the cytotoxic functions of CD8+ T cells on CRC-PDOs, which was reflected by the declined proportions of granzyme B (GZMB) + and interferon gamma (IFN&#x3b3;) + CD8+ T cells and enhanced viability of CRC-PDOs. Mechanistically, the co-culture with senescent fibroblasts promoted the lipid shuttling into CD8+ T cells to induce lipid peroxidation and downstream impairment of cytotoxicity. Furthermore, the inhibition of CD36, the specific scavenger receptor for lipid uptake of CD8+ T cells, effectively suppressed lipid transfer and peroxidation thereby preserving the effector functions of CD8+ T cells and ultimately promoting tumor apoptosis. Complementarily, in vivo senolytic treatment significantly suppressed CRC progression in AOM-DSS CRC mouse models. Top 12 therapeutic agents were identified significantly enhanced predicted efficacy in CD36-high tumors. CONCLUSIONS: Our study identified a substantial population of senescent fibroblasts in human CRC through single cell transcriptomics, machine-learning and clinical biopsies. These senescent fibroblasts impair CD8+ T cell-mediated killing of CRC-PDOs via CD36-dependent lipid transfer, suggesting senolytic targeting of stromal cells as a promising immunotherapeutic strategy for CRC.

Colorectal Neoplasms↗