Search PubMedSearch

SEARCH · Search PubMed

Results for “Leave-one-out cross-validation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

7 recordsLinked to original sources

Machine Learning and Metabolomics to Characterize Warburg-Like Metabolic Subtypes in Human Retinal Endothelial Cells Exposed to Risk Factors Associated With Proliferative Diabetic Retinopathy.

PURPOSE: High glucose (HG), hypoxia (Hyp), and their combination are major risk factors for proliferative diabetic retinopathy (PDR). Although these conditions induce features of the Warburg-like metabolic reprogramming in human retinal endothelial cells (HRECs), it remains unclear whether they produce distinct metabolic and angiogenic subtypes. This study aimed to characterize the Warburg-like-associated metabolic heterogeneity induced by these PDR-related risk factors and evaluate the ability of supervised machine-learning models to distinguish these subtypes. METHODS: HRECs were cultured under normoglycemic, HG, Hyp (2% O2), and combined HG-Hyp conditions. Untargeted LC-MS/MS metabolomics quantified metabolites spanning carbohydrates, amino acids, nucleotides, and lipids. Principal component analysis (PCA) assessed overall metabolic variation, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis identified metabolic pathways associated with angiogenesis. In vitro angiogenesis assays measured endothelial tube formation and branching. Nine supervised classifiers (decision tree, logistic regression, naïve Bayes, random forest, K-Nearest Neighbors, neural network, gradient boosting, AdaBoost, and Support Vector Machine) were trained on the highest-ranked metabolites selected by the Information Gain Ratio feature-ranking approach. Model performance was evaluated using 10-fold cross-validation, leave-one-out cross-validation (LOOCV), permutation testing, and a classifier stability analysis under biologically meaningful distributional shift using an independent chemically induced hypoxia model (CoCl2). RESULTS: PCA revealed partial separation of metabolic profiles across conditions, indicating different Warburg-like metabolic subtypes. The combined HG-Hyp condition exhibited enhanced angiogenic potential relative to either HG or Hyp alone. KEGG pathway enrichment analysis identified fatty acid biosynthesis and elongation among the most significantly enriched pathways in HRECs under combined HG-Hyp conditions, alongside amino sugar and nucleotide sugar metabolism, glycerophospholipid metabolism, the pentose phosphate pathway, and glycolysis/gluconeogenesis. Supervised machine-learning classifiers distinguished these metabolic subtypes, with AdaBoost and gradient Boosting showing the most balanced, reproducible performance across 10-fold cross-validation, LOOCV, and permutation testing, and remaining the most reliable classifiers under domain-shift testing (area under the curve = 0.88, P = 0.0061). CONCLUSIONS: In this exploratory analysis, HG, Hyp, and their combination drive metabolically and functionally distinct subtypes of Warburg-like metabolic reprogramming in HRECs, with HG-Hyp in combination producing a highly angiogenic phenotype. Boosting-based ensemble classifiers provide a promising framework for detecting these subtypes even under domain-shift conditions, warranting validation in larger independent datasets. TRANSLATIONAL RELEVANCE: Integrating metabolomics with machine-learning classification offers a strategy to identify Warburg-like metabolic subtypes in retinal endothelial cells, providing insights into angiogenic mechanisms and guiding the development of targeted diagnostics or therapeutics for PDR.

Humans

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n = 26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette ≈ 0.16) that remained unassociated with overall survival (log-rank p = 0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) = 0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p = 0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans

Estimating the association of antimicrobial resistance genes with minimum inhibitory concentration in Escherichia coli: an observational study.

BACKGROUND: Surveillance and prediction of antibiotic resistance in Escherichia coli relies on curated databases of genes and mutations. We aimed to quantify the effect of acquiring specific genetic elements on minimum inhibitory concentrations (MICs) for particular antibiotic-species combinations, addressing the current scarcity of such data in existing databases. METHODS: For this observational study, we evaluated a collection of E coli isolates with linked whole-genome sequencing and MIC data, originating from human urinary or bloodstream infections obtained from the Oxford University Hospitals National Health Service Foundation Trust in Oxfordshire, UK. We used multivariable interval regression models to estimate the change in MIC (with 95% CIs) for specific antibiotics associated with the acquisition of antibiotic resistance genes and associated mutations in the National Center for Biotechnology Information AMRFinder database, with and without an adjustment for population structure. We then tested the ability of these models to predict MIC and binary resistance or susceptibility using leave-one-out cross-validation. FINDINGS: We evaluated 2875 E coli isolates obtained during 2013-2018 and 2020. Although most ARGs and resistance mutations (89 [80%] of 111) were associated with an increased MIC, a much smaller number (27 [24%] of 111) was found to be putatively independently resistance-conferring (ie, associated with an MIC above the European Committee on Antimicrobial Susceptibility Testing breakpoint) when acquired in isolation. We found evidence of differential effects of acquired ARGs and resistance mutations between different generations of cephalosporin antibiotics and showed that sub-breakpoint variation in MIC can be linked to genetic mechanisms of resistance. 20 697 (83·3%; range 52·9-97·7 across all antibiotics) of 24 858 MICs were correctly exactly predicted and 23 677 (95·2%; 87·3-97·7) of 24 858 MICs were predicted to within one doubling dilution. INTERPRETATION: Quantitative estimates of the independent effect of the acquisition of ARGs on MIC add to the interpretability and utility of existing databases. Compared with approaches using machine learning models, the use of these estimates yields similar or better performance in the prediction of antibiotic resistance phenotype with more readily interpretable results. The methods outlined here could be readily applied to other antibiotic-pathogen combinations. FUNDING: The National Institute for Health and Care Research (NIHR) and the Medical Research Council (MRC).

Escherichia coli

MicroRNAs signatures in small extracellular vesicles for psychological resilience in young adults using machine learning.

AIMS: Psychological resilience refers to an individual's capacity to adapt to adverse events. MicroRNAs (miRNAs) play a crucial role in regulating post-transcriptional processes, while small extracellular vesicles (sEVs) act as transport vehicles. This study aimed to employ genome-wide profiling to identify and validate differences in the expression of resilience-associated sEV-miRNAs between low resilience (LR) and high resilience (HR) in young adults. METHODS: Eighty participants were divided into LR or HR based on the Connor - Davidson Resilience Scale (CD-RISC). The expression levels of the target sEV-miRNAs in LR and HR were compared and analyzed. RESULTS: Expression analyses demonstrated significant differences in let-7b, miR-151b, miR-335, and miR-193a between LR and HR (p&#x2009;<&#x2009;0.01), with let-7b showing the highest discriminative ability. The AUC values for each sEV-miRNA ranged from 0.74 to 0.94, based on logistic regression and three machine learning models: random forest, support vector machine, and eXtreme gradient boosting. Based on leave-one-out cross-validation in different models, the combined four sEV-miRNAs demonstrated strong performance for detecting LR (AUC&#x2009;=&#x2009;0.87-0.90). Sex-specific differences were also observed, with female participants showing more pronounced resilience signatures in targeted sEV-miRNAs. CONCLUSIONS: These findings suggest that sEV-miRNAs hold potential as biomarkers for psychological resilience in young adults.

Humans

Metagenome-scale modeling to assess microbiome metabolic complementarity for precision microbiota transplantation therapies.

Fecal microbiota transplantation (FMT) holds therapeutic promise beyond recurrent Clostridioides difficile infection, but clinical outcomes remain unpredictable and donor-selection strategies remain limited, in part because the role of donor&#x2012;recipient metabolic interactions in shaping the post-FMT community remains poorly understood. Here, we leverage metagenome-scale metabolic modeling to quantify metabolic niche complementarity between donor and recipient microbiomes and predict post-FMT community composition. Using MICOM-derived metabolic models, we show that donor genomes whose metabolic flux profiles are more dissimilar from the recipient community colonize at significantly higher rates in a murine FMT model. In a human IBS trial, the same metric predicted post-FMT community composition via leave-one-out cross-validation and captured known disease-associated alterations in short-chain fatty acid, sulfur, and gas metabolism. We then performed 2,548 in silico FMT simulations between IBS-D/M patients and donors from the OpenBiome biobank to evaluate personalized donor screening, identifying super-donors characterized by high taxonomic diversity, broad metabolic niche coverage, and community interaction networks dominated by cross-feeding rather than competition. Together, these results support metabolic niche complementarity as a potential determinant of post-FMT community composition and provide a mechanistic basis for evaluating donor-recipient metabolic compatibility. This framework offers a scalable approach for generating testable hypotheses for personalized donor selection.

Fecal Microbiota Transplantation

DNA methylation-based ageing in a deuterostome invertebrate: an epigenetic clock for the crown-of-thorns seastar (Acanthaster cf. solaris).

Accurate and reliable ageing tools are essential for wildlife conservation and management. While DNA methylation has emerged as a promising tool for age estimation in vertebrates, its application to invertebrates remains contested and has been limited to arthropods. Here, we develop an epigenetic clock for the Pacific crown-of-thorns seastar (CoTS; Acanthaster cf. solaris), a destructive coral predator contributing to habitat degradation across Indo-Pacific reefs. Using Oxford Nanopore Technologies, we generated whole-genome DNA methylation profiles across five age groups and identified 1910 CpG sites with methylation patterns significantly associated with age. We then fitted age prediction models using elastic net regression and evaluated predictive performance with leave-one-out cross-validation (LOOCV), achieving a mean absolute error of 0.31 &#xb1; 0.22 years, corresponding to 4-6% of the CoTS lifespan (5-8 years). This accuracy suggests the potential to differentiate annual cohorts, supporting future management-relevant inference. To facilitate practical implementation, we constructed an optimized epigenetic clock from 14 CpG sites consistently selected across LOOCV iterations. Our results demonstrate that DNA methylation-based age estimation is feasible in a deuterostome invertebrate, extending epigenetic ageing approaches beyond arthropods and establishing their potential to advance age determination and management in invertebrates that lack reliable ageing methods.

Animals

Identification of candidate variants in plasma associated with early versus late disease progression under anti-PD-1 therapy in metastatic NSCLC.

BACKGROUND: Immune checkpoint inhibitors (ICIs), including anti-programmed cell death protein 1 (anti-PD-1) antibodies, have significantly improved outcomes in patients with metastatic non-small cell lung cancer (mNSCLC). However, substantial heterogeneity exists in clinical benefit, with some patients exhibiting early progression (EP) and others late progression (LP). To date, no biomarkers of EP versus LP disease have been implemented in clinical practice. Circulating tumor DNA (ctDNA) analysis represents a minimally invasive strategy for identifying such biomarkers. In this proof-of-concept study, we evaluated the performance of the TruSight Oncology 500 ctDNA (TSO500 ctDNA) panel and explored its feasibility to identify candidate variants associated with early and late disease progression under anti-PD-1 therapy. METHODS: Baseline ctDNA from eight mNSCLC patients treated with pembrolizumab was extracted and sequenced using the TSO500 ctDNA assay, a 523-gene targeted next-generation sequencing panel. Patients were classified according to their response as LP or EP. Variant calling was performed using the DRAGEN Bio-IT platform, and variants were annotated and clinically interpreted using the Clinical Genomics Workspace (CGW; PierianDx) according to Association for Molecular Pathology (AMP)/American Society of Clinical Oncology (ASCO)/College of American Pathologists (CAP) guidelines. Survival outcomes were assessed using Kaplan-Meier and log-rank tests. Performance of ctDNA variants was evaluated using receiver operating characteristic (ROC) curve analysis, and multi-gene models were assessed using leave-one-out cross-validation with penalized logistic regression. RESULTS: All patients harbored detectable variants, including SNVs (100%), MNVs (87.5%), deletions (75%), and insertions (62.5%). Tier I variants were identified in 37.5% of patients, while all cases showed tier II and multiple tier III alterations. TP53 variants were associated with poorer outcomes under anti-PD-1 therapy. Individual gene alterations in TP53, ERBB3, SMC1A or LATS1 showed moderate discriminatory performance between LP and EP patients; however, combination of mutated genes improved apparent discrimination. Notably, specific two-gene combinations (SMC1A + LATS1 or ERBB3 + LATS1) showed the highest discriminatory performance between LP and EP patients in this exploratory cohort. CONCLUSIONS: This study demonstrates the feasibility and analytical performance of the TSO500 ctDNA panel and provides hypothesis-generating evidence that plasma gene variants may be useful to evaluate early versus late disease progression in patients with mNSCLC receiving immunotherapy.

TruSight Oncology 500