Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets.

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Escherichia coli↗

A comprehensive meta-analysis of tissue resident memory T cells and their roles in shaping immune microenvironment and patient prognosis in non-small cell lung cancer.

Tissue-resident memory T cells (TRM) are a specialized subset of long-lived memory T cells that reside in peripheral tissues. However, the impact of TRM-related immunosurveillance on the tumor-immune microenvironment (TIME) and tumor progression across various non-small-cell lung cancer (NSCLC) patient populations is yet to be elucidated. Our comprehensive analysis of multiple independent single-cell and bulk RNA-seq datasets of patient NSCLC samples generated reliable, unique TRM signatures, through which we inferred the abundance of TRM in NSCLC. We discovered that TRM abundance is consistently positively correlated with CD4+ T helper 1 cells, M1 macrophages, and resting dendritic cells in the TIME. In addition, TRM signatures are strongly associated with immune checkpoint and stimulatory genes and the prognosis of NSCLC patients. A TRM-based machine learning model to predict patient survival was validated and an 18-gene risk score was further developed to effectively stratify patients into low-risk and high-risk categories, wherein patients with high-risk scores had significantly lower overall survival than patients with low-risk. The prognostic value of the risk score was independently validated by the Cancer Genome Atlas Program (TCGA) dataset and multiple independent NSCLC patient datasets. Notably, low-risk NSCLC patients with higher TRM infiltration exhibited enhanced T-cell immunity, nature killer cell activation, and other TIME immune responses related pathways, indicating a more active immune profile benefitting from immunotherapy. However, the TRM signature revealed low TRM abundance and a lack of prognostic association among lung squamous cell carcinoma patients in contrast to adenocarcinoma, indicating that the two NSCLC subtypes are driven by distinct TIMEs. Altogether, this study provides valuable insights into the complex interactions between TRM and TIME and their impact on NSCLC patient prognosis. The development of a simplified 18-gene risk score provides a practical prognostic marker for risk stratification.

Humans↗

RNAcare: integrating clinical data with transcriptomic evidence using rheumatoid arthritis as a case study.

BACKGROUND: Gene expression analysis is a crucial tool for uncovering the biological mechanisms that underlie differences between patient subgroups, offering insights that can inform clinical decisions. However, despite its potential, gene expression analysis remains challenging for clinicians due to the specialised skills required to access, integrate, and analyse large datasets. Existing tools primarily focus on RNA-Seq data analysis, providing user-friendly interfaces but often falling short in several critical areas: they typically do not integrate clinical data, lack support for patient-specific analyses, and offer limited flexibility in exploring relationships between gene expression and clinical outcomes in disease cohorts. Users, including clinicians with a general knowledge of transcriptomics, however, who may have limited programming experience, are increasingly seeking tools that go beyond traditional analysis. To overcome these issues, computational tools must incorporate advanced techniques, such as machine learning, to better understand how gene expression correlates with patient symptoms of interest. RESULTS: Our RNAcare platform, addresses these limitations by offering an interactive and reproducible solution specifically designed for analysing transcriptomic data from patient samples in a clinical context. This enables researchers to directly integrate gene expression data with clinical features, perform exploratory data analysis, and identify patterns among patients with similar diseases. By enabling users to integrate transcriptomic and clinical data, and customise the target label, the platform facilitates the analysis of the relationships between gene expression and clinical symptoms like pain and fatigue. This allows users to generate hypotheses and illustrative visualisations/reports to support their research. As proof of concept, we use RNAcare to link inflammation-related genes to pain and fatigue in rheumatoid arthritis (RA) and detect signatures in the drug response group, confirming previous findings. CONCLUSION: We present a novel computational platform allowing the interpretation of clinical and transcriptomics data in real-time. The platform can be used for data generated by the user, such as the patient data presented here or using published datasets. The platform is available at https://rna-care.mvls.gla.ac.uk/ , and its source code is https://github.com/sii-scRNA-Seq/RNAcare/ .

Humans↗

Distinct immune-metabolic phenotypes underlie poor coronary collateral circulation.

BACKGROUND: Coronary collateral circulation (CCC) significantly impacts myocardial perfusion and clinical outcomes in coronary artery disease patients, yet the underlying molecular heterogeneity remains inadequately characterized. OBJECTIVE: To identify distinct molecular phenotypes in patients with poor CCC, validate these phenotypes using clinical parameters, and evaluate their prognostic implications. METHODS: This study enrolled 149 patients (80 with good CCC and 69 with poor CCC) for high-throughput proteomic profiling. Unsupervised consensus clustering identified molecular subtypes within poor CCC patients, followed by differential expression analysis and KEGG pathway enrichment. Boruta feature selection was implemented, and multiple machine learning algorithms were tested on clinical data, with XGBoost optimization (accuracy 80.0%, F1-score 80.31%) and SHAP value interpretation. External validation was performed using the MIMIC database. Kaplan-Meier analysis and Cox regression models assessed major adverse cardiovascular events (MACE). RESULTS: Two distinct phenotypes emerged among poor CCC patients: Cluster 1 (n&#x2009;=&#x2009;39, Complement-Driven Vascular Remodeling [CDVR]) and Cluster 2 (n&#x2009;=&#x2009;30, Immuno-Thrombotic Myocardial Dysfunction [ITMD]). An XGBoost model incorporating fasting glucose, eosinophil percentage, and HbA1c achieved excellent discrimination (AUC&#x2009;>&#x2009;0.91). External validation confirmed the phenotype-specific clinical patterns. Notably, Cluster 2 demonstrated significantly higher MACE incidence compared to Cluster 1 (Log-rank p&#x2009;<&#x2009;0.05), with KEGG analysis revealing significant upregulation of platelet activation, diabetic cardiomyopathy, and metabolic pathways in the ITMD phenotype. CONCLUSION: Poor CCC encompasses distinct immune-metabolic phenotypes that can be accurately classified using integrated proteomic-clinical modeling. This classification enables more precise risk stratification and may guide personalized therapeutic strategies for coronary artery disease patients with inadequate collateralization.

Humans↗

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans↗

Lactylation-related immune-metabolic dysregulation defines prognostic and therapeutic stratification in lung adenocarcinoma.

BACKGROUND: Lactylation links lactate metabolism with inflammatory signaling and immune regulation in tumors. However, its cellular distribution and translational value in lung adenocarcinoma (LUAD) remain unclear. METHODS: Single-cell RNA-sequencing datasets GSE189357 and GSE171145 were integrated to characterize lactylation-related activity, intercellular communication, and malignant epithelial cell states in LUAD. Single-cell-derived lactylation-related differentially expressed genes were mapped to TCGA-LUAD and multiple GEO cohorts. Univariate Cox regression and machine learning algorithms were used to construct a lactylation-related prognostic signature (LRPS). The associations of LRPS with prognosis, immunotherapy response, drug sensitivity, genomic alterations, immune infiltration, and inflammation- and metabolism-related pathways were evaluated. KRT7 was further validated using virtual knockout analysis, spatial transcriptomics, and in vitro and in vivo experiments. RESULTS: lactylation-related transcriptional activity showed heterogeneous distribution across LUAD cell populations and was associated with altered cell-cell communication. In malignant epithelial cells, LRTS-high and LRTS-low states exhibited distinct metabolic, inflammatory, and tumor-related pathway activities. LRPS showed stable prognostic performance in TCGA-LUAD and multiple GEO cohorts and remained an independent prognostic factor. Low LRPS was associated with greater potential benefit from immunotherapy, whereas different LRPS groups displayed distinct drug sensitivity, genomic alteration, and immune microenvironment patterns. KRT7 was highly expressed in LUAD and associated with poor prognosis. KRT7 knockdown suppressed LUAD cell proliferation, migration, invasion, colony formation, and tumor growth in vivo. CONCLUSIONS: This study identifies lactylation-related immune-metabolic dysregulation as a clinically relevant feature of LUAD and develops a single-cell-guided LRPS for prognosis and therapeutic stratification. KRT7 emerged as an LRPS-related functional candidate with experimentally supported roles in malignant LUAD phenotypes.

Immunotherapy↗

Rhythm profiling using COFE reveals multi-omic circadian rhythms in human cancers in vivo.

The study of ubiquitous circadian rhythms in human physiology requires regular measurements across time. Repeated sampling of the different internal tissues that house circadian clocks is both practically and ethically infeasible. Here, we present a novel unsupervised machine learning approach (COFE) that can use single high-throughput omics samples (without time labels) from individuals to reconstruct circadian rhythms across cohorts. COFE can simultaneously assign time labels to samples and identify rhythmic data features used for temporal reconstruction, while also detecting invalid orderings. With COFE, we discovered widespread de novo circadian gene expression rhythms in 11 different human adenocarcinomas using data from The Cancer Genome Atlas (TCGA) database. The arrangement of peak times of core clock gene expression was conserved across cancers and resembled a healthy functional clock except for the mistiming of a few key genes. Moreover, rhythms in the transcriptome were strongly associated with the cancer-relevant proteome. The rhythmic genes and proteins common to all cancers were involved in metabolism and the cell cycle. Although these rhythms were synchronized with the cell cycle in many cancers, they were uncoupled with clocks in healthy matched tissue. The targets of most of FDA-approved and potential anti-cancer drugs were rhythmic in tumor tissue with different amplitudes and peak times. These findings emphasize the utility of considering "time" in cancer therapy, and suggest a focus on clocks in healthy tissue rather than free-running clocks in cancer tissue. Our approach thus creates new opportunities to repurpose data without time labels to study circadian rhythms.

Humans↗

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation↗

Integrative multi-omics analyses suggest a candidate microbial metabolite-associated host gene network in ulcerative colitis.

Ulcerative colitis (UC) is associated with gut microbial dysbiosis, but the host molecular alterations potentially linked to microbially derived metabolites remain incompletely understood. We integrated Mendelian randomization (MR), microbial metabolite annotation, computational target prediction, colonic transcriptomics, network analysis, and machine learning. MiBioGen microbiome GWAS data were used as exposures and FinnGen Release 12 ULCERENTER as the outcome. Metabolites linked to MR-prioritized taxa were retrieved from GutMGene, and human targets were predicted using SwissTargetPrediction and SEA. UC-related genes were defined by integrating differential expression analysis and WGCNA and then intersected with predicted metabolite targets. MR prioritized one family and eight genera showing nominal genetically supported associations with UC, but none remained significant after Benjamini-Hochberg FDR correction. Three prioritized genera were linked to 15 microbe-metabolite records, corresponding to 13 unique metabolites; nine were retained for target prediction, yielding 277 unique predicted human targets. Transcriptomic analysis identified 1,530 DEGs and a 312-gene MEgrey60 module, with 273 overlapping genes, producing 1,569 unique UC-related genes. Their intersection with the 277 predicted targets yielded 47 candidate genes. Enrichment analyses highlighted mainly metabolic and lipid-related processes. Random Forest showed the highest mean AUC across the two independent external benchmarking cohorts, and SHAP prioritized EPHX1, HSD17B2, IGFBP5, and MMP10. IBDome analysis showed inflammation-associated expression differences in these genes. This study provides a genomics-informed, hypothesis-generating framework that prioritizes candidate microbe-metabolite-host relationships in UC for future experimental validation.

Humans↗

Metagenomic analyses reveal E. coli-derived siderophores as potential signatures for breast cancer.

BACKGROUND: Breast cancer remains a leading cause of cancer-related mortality in women. Recent evidence implicates the gut microbiome and metabolites in breast cancer pathogenesis. This study explores associations between gut microbial species, their predicted metabolites, and breast cancer to uncover potential mechanistic insights. METHODS: Comprehensive metagenomic analyses were conducted on the gut microbiome of pre- and postmenopausal breast cancer patients, where microbial species were profiled through AMPHORA2 and metabolites were predicted through antiSMASH. Multivariate association analysis was used to identify significant associations between specific microbial species, predicted metabolites, and breast cancer status. A custom ensemble machine learning classifier was developed to classify pre- and postmenopausal breast cancer cases and controls based on microbial and predicted metabolite features. Additionally, a synthetic microbiome dataset was generated through MIDASim to validate the reproducibility of the ML results. Using our results, we explored the underlying dynamics of identified taxa and metabolite in breast cancer through literature and statistical support. RESULTS: Our analysis identified 471 microbial species and predicted 40 key metabolites in the metagenomic data. Multivariate analysis identified significant positive associations (p-value&#x2009;<&#x2009;0.05) of E. coli, siderophore, and thiopeptide with breast cancer. The custom ensemble model achieved accuracy and AUC as high as 78% and 90%, respectively, in classifying pre- and postmenopausal cases and controls. The high-ranking features i.e., E. coli, siderophore, and thiopeptide were consistent with the results of the multivariate association analysis, thereby substantiating their biological significance. Using these findings, we propose a mechanistic model in which E. coli secretes siderophores under iron-limited conditions in breast cancer patients, for iron sequestration from the host, which can potentially promote angiogenesis and tumor progression. CONCLUSION: Our findings suggest that microbial iron acquisition mechanisms may play a critical role in breast cancer pathophysiology. Functional validation of these mechanisms is needed to assess therapeutic potential. This study highlights gut microbiota and their metabolites as promising targets for breast cancer research and intervention.

Breast Neoplasms↗

Prediction of gene expression using histone modification patterns extracted by Particle Swarm Optimization.

MOTIVATION: Histone modifications play an important role in transcription regulation. Although the general importance of some histone modifications for transcription regulation has been previously established, the relevance of others and their interaction is subject to ongoing research. By training Machine Learning models to predict a gene's expression and explaining their decision making process, we can get hints on how histone modifications affect transcription. In previous studies, trained models were either hardly explainable or the models were trained solely on the abundance of histone modifications. Based on other studies, which used histone modification patterns, rather than their abundance, to identify potential regulatory elements, we hypothesize the histone modification pattern in a gene's promoter to be more predictive for gene expression. We used an optimization algorithm to extract predictive histone modification profiles. RESULTS: Our algorithm called PatternChrome achieved an average area under curve (AUC) score of 0.9029 over 56 samples for binary classification, outperforming all previous algorithms for the same task. We explained the models decisions to deduce the effect of specific features, certain histone modifications or promoter positions on transcription regulation. Although the predictive histone modification patterns were extracted for each sample separately, they can be used to predict gene expression in other samples, implying that the created patterns are largely generalizable. Interestingly, the impact of histone modifications on gene regulation appears predominantly indifferent to cellular specificity. Through explanation of the classifier's decisions, we substantiate established literature knowledge while concurrently revealing novel insights into the intricate landscape of transcriptional regulation via histone modification. AVAILABILITY AND IMPLEMENTATION: The code for the PatternChrome algorithm, the scripts for the analyses and the required data can be found at (https://gitlab.gwdg.de/MedBioinf/generegulation/patternchrome).

Humans↗

A cfDNA fragmentomics classifier for noninvasive differentiation of benign and malignant renal masses.

Noninvasive differentiation of malignant and benign renal masses remains a major clinical challenge, particularly for radiologically indeterminate lesions. Here, we developed and validated a plasma cell-free DNA (cfDNA) fragmentomics-based machine learning classifier for renal mass characterization. The model was trained on 331 participants (171 cancer, 160 benign) and independently validated on 144 participants (73 cancer, 71 benign). Three cfDNA fragmentation features, including copy number variation (CNV), fragmentation-based methylation (FRAGMA), and nucleosome footprint (NF), derived from low-pass whole-genome sequencing, were integrated into an ensemble framework. The model achieved strong discriminative performance, with area under the curve (AUC) values of 0.956 in the training cohort and 0.946 in the validation cohort, outperforming individual feature-based models. At a predefined operating threshold corresponding to 90% sensitivity, specificity reached 0.90 and 0.87, respectively. Notably, most cancer samples exhibited low tumor fraction (TF&#x2009;<&#x2009;3%), yet the model maintained robust performance in low-TF samples (AUCs: 0.952 and 0.941, respectively). Performance remained consistent across tumor stage, grade, and histological subtypes. The classifier also demonstrated potential clinical utility in diagnostically challenging settings, including lipid-poor angiomyolipoma and oncocytoma, with 12 of 13 oncocytoma samples correctly classified in an independent cohort. In addition, the model correctly identified 85.3% of benign masses&#x2009;>&#x2009;4&#xa0;cm, for which surgical intervention is more commonly considered, and 84.6% of malignant tumors&#x2009;&#x2264;&#x2009;4&#xa0;cm, for which management can be challenging. Collectively, these findings support cfDNA fragmentomics as a promising noninvasive liquid biopsy approach for renal mass evaluation and clinical decision-making.

Humans↗

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2↗

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans↗

Establishment of a prognostic model based on ER stress-related cell death genes and proposing a novel combination therapy in acute myeloid leukemia.

BACKGROUND: Acute myeloid leukemia (AML) is a highly heterogeneous malignancy, presenting significant challenges in accurately predicting patient prognosis. Dysregulation of endoplasmic reticulum (ER) stress and resistance to programmed cell death (PCD) are hallmarks of AML cells. However, the prognostic significance of the interplay between ER stress and cell death pathways in AML remains largely unexplored. METHODS: We analyzed RNA sequencing and clinical data from 887 AML patients across 4 cohorts to develop an ER stress-related cell death index (ERCDI) using 10 machine-learning algorithms with 117 unique combinations. Survival and time-dependent Receiver Operating Characteristic Curve (ROC) analyses were performed to assess the model's efficacy. Clinical characteristics, the tumor immune microenvironment, and drug sensitivity differences between the high- and low-risk groups were also analyzed. The CMap database was used to identify potential therapeutic drugs. In vitro and in vivo experiments, including CCK-8, colony formation, flow cytometry, Transwell assays, and xenograft mouse models, were conducted to evaluate the effects of the target genes and candidate drugs. RESULTS: The ERCDI demonstrated strong prognostic and predictive performance for prognosis in AML patients. Furthermore, the ERCDI effectively predicted immunotherapy and chemotherapy outcomes and was associated with the immune features of the different risk groups. DNA damage-inducible transcript 4 protein (DDIT4), a key gene associated with ERCDI, is related to poor prognosis in AML patients with high expression. Additionally, the knockdown of DDIT4 significantly inhibited AML cell proliferation, induced cell apoptosis, and promoted cell cycle arrest. Chaetocin was subsequently identified as a candidate compound for AML treatment. Subsequent experiments suggested that combining chaetocin and venetoclax is a potentially promising therapeutic strategy for AML. CONCLUSION: The ERCDI provides personalized risk assessment and treatment recommendations for individual AML patients. The combined use of chaetocin and venetoclax can potentially be repurposed for AML therapy.

Humans↗

Use of IR Biotyper as a feasible methodology to type Klebsiella pneumoniae.

UNLABELLED: Klebsiella pneumoniae is one of the most frequently reported healthcare-associated pathogens. The current gold standard approach to perform the epidemiological typing of these bacteria is Whole Genome Sequencing (WGS), which is an expensive and challenging procedure. IR Biotyper (Bruker Daltonics, GmbH) is a new equipment based on Fourier transform infrared spectroscopy, which allows a rapid, low-cost, and user-friendly method to type bacterial isolates. However, there is a need for studies that evaluate the efficacy of the IR Biotyper. The aim of this study was to evaluate the capability of IR Biotyper to type K. pneumoniae according to sequence type (ST) and capsular type-using K locus (KL)-as well as to develop a classifier using machine learning. Seventy-three isolates of K. pneumoniae previously characterized by WGS were selected for IR Biotyper analysis using principal component analysis for dimensionality reduction, Euclidean, and unweighted pair group method with arithmetic mean (UPGMA) for clustering method, and spectra were analyzed in the 1,300-800 cm&#x207b;&#xb9; wavenumber range. Among these, 54 isolates were used to create a classifier, and 19 were used to validate the classifier. When considering the ST, ST307 was grouped in the same cluster as ST11. When KL was considered for the analysis, the clusters were 100% correctly grouped according to their KL type. Furthermore, the classifier developed was able to classify the isolates according to KL with a high concordance. This study showed that KL correlates well with KL for typing K. pneumoniae isolates using the IR Biotyper. Additionally, IR Biotyper demonstrated to be a cost-effective method and a promising tool to classify isolates within minutes. IMPORTANCE: Klebsiella pneumoniae is a major cause of severe hospital infections, and controlling its spread requires quick identification and comparison of bacterial strains. WGS is accurate but expensive, slow, and technically demanding. In this study, we evaluated the IR Biotyper, a device that uses infrared light to analyze bacteria and group them by capsule type-a key feature linked to their spread. The IR Biotyper matched WGS results with high accuracy, delivering results in minutes instead of days. This fast, affordable method can help hospitals detect outbreaks earlier and respond more effectively. Our findings suggest that the IR Biotyper is a valuable tool for routine use in microbiology laboratories, supporting epidemiological surveillance and outbreak control.

Klebsiella pneumoniae↗

NMR metabolomics and glycomics for cancer detection in patients with non-specific symptoms: a prospective observational cohort study.

BACKGROUND: Early cancer diagnosis in patients with non-specific symptoms is limited by the lack of discriminatory tests. Within the Oxfordshire Suspected CANcer (SCAN) pathway, exploratory biomarker work showed that serum 1H NMR-based metabolomics can identify cancer with high accuracy. SCAN2 evaluated whether integrating metabolomics with glycomics provides complementary molecular information and improves discrimination in a clinically complex, real-world population. METHODS: Serum from 369 SCAN patients (59 cancers) was analysed using AXINON&#xae; System-derived NMR metabolomics and HPLC-MS glycomics. Machine-learning models were trained to predict cancer status, with performance assessed by receiver operating characteristic (ROC) analysis of pooled cross-validated predictions. To place cancer risk in a broader clinical context, a second classifier modelling alternative non-cancer diagnosis was incorporated, and mean predicted probabilities from both models were jointly projected into a two-dimensional space, maintaining strict separation of training and test data. FINDINGS: In the full cohort, integration of glycomics with metabolomics achieved an AUC of 0.814 (95% CI 0.808-0.820). In a refined sub-cohort excluding major comorbidities and selected cancer types (32 cancers, 277 non-cancers), performance improved to an AUC of 0.884 (95% CI 0.879-0.890). Discriminatory features included cancer-associated biantennary fucosylated glycans alongside amino acid metabolites (glutamate, histidine) and lipoprotein-related measures. A classifier distinguishing metastatic from non-metastatic disease (n = 29 vs. 30) achieved an AUC of 0.80. Joint probability analysis in the full cohort preserved cancer-associated signatures across comorbidity burden, with projection-based classification achieving an accuracy of 89.2% (95% CI 85.7-92.6). INTERPRETATION: These findings validate the SCAN1 metabolomic signature in a more clinically complex cohort and indicate that integrating glycomics with metabolomics provides complementary biological information for cancer discrimination. Joint probability analysis provides an interpretable framework for cancer risk stratification within multimorbid diagnostic pathways, supporting the clinical potential of scalable multi-omics blood testing. FUNDING: EPSRC, EU Horizon 2020, Wellcome/MLSTF, Novo Nordisk Foundation.

Humans↗

Liquid Biopsy-Multiomics Link Adhesion Pathway Dysregulation to Kidney Injury Severity.

INTRODUCTION: Severe acute kidney injury (AKI) is strongly associated with the risk of developing chronic kidney disease; however, little is known about the cell type-specific mechanisms driving kidney injury severity. METHODS: In this multicenter observational study, we used clinically obtained liquid biopsy proteomics and machine learning (ML) to predict severe outcomes in patients with COVID-associated and non-COVID AKI. Further, we orthogonally combined 169 urine proteomics with 437 plasma proteomics samples and 40 urine sediment single-cell transcriptomics samples to identify complementary dysregulated mechanisms. RESULTS: Using a 10-fold cross-validated random forest algorithm, we identified a set of urinary proteins that demonstrate predictive power for both discovery and validation set with AUC of 87% and 76%, respectively. These predictive proteomics features obtained demonstrate that cell adhesion and autophagy-associated pathways are uniquely impacted in severe AKI. Differentially abundant proteins (DAPSs) associated with these pathways are highly expressed in cells of the juxtamedullary nephron, endothelial cells (ECs), and podocytes, indicating that these kidney cell types could be potential targets. Single-cell transcriptomic analysis in the in vitro model of kidney organoids infected with SARS-CoV-2 reveal dysregulation of extracellular matrix (ECM) organization in multiple nephron segments, recapitulating the clinically observed fibrotic response across multiomics datasets. Ligand-receptor interaction analysis of the podocyte and tubule organoid clusters shows significant reduction and loss of interaction between integrins and basement membrane receptors in the infected kidney organoids. CONCLUSION: Collectively, these data suggest that ECM degradation and adhesion-associated mechanisms could be the main driver of severe kidney injury.

AKI↗