Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

MetaFX: feature extraction from whole-genome metagenomic sequencing data.

MOTIVATION: Microbial communities consist of thousands of microorganisms and viruses and have a tight connection with an environment, such as gut microbiota modulation of host body metabolism. However, the direct relationship between the presence of certain microorganism and the host state often remains unknown. Toolkits using reference-based approaches are limited to microbes present in databases. Reference-free methods often require enormous resources for metagenomic assembly or results in many poorly interpretable features based on k-mers. RESULTS: Here we present MetaFX-an open-source library for feature extraction from whole-genome metagenomic sequencing data and classification of groups of samples. Using a large volume of metagenomic samples deposited in databases, MetaFX compares samples grouped by metadata criteria (e.g. disease, treatment, etc.) and constructs genomic features distinct for certain types of communities. Features constructed based on statistical k-mer analysis and de Bruijn graphs partition. Those features are used in machine learning models for classification of novel samples. Extracted features can be visualized on de Bruijn graphs and annotated for providing biological insights. We demonstrate the utility of MetaFX by building classification models for 590 human gut samples with inflammatory bowel disease. Our results outperform the previous research disease prediction accuracy up to 17%, and improves classification results compared to taxonomic analysis by 9±10% on average. AVAILABILITY AND IMPLEMENTATION: MetaFX is a feature extraction toolkit applicable for metagenomic datasets analysis and samples classification. The source code, test data, and relevant information for MetaFX are freely accessible at https://github.com/ctlab/metafx under the MIT License. Alternatively, MetaFX can be obtained via http://doi.org/10.5281/zenodo.16949369.

Metagenomics

An adjuvant database for preclinical evaluation of vaccines and immunotherapeutics.

Adjuvants are immunostimulators used to enhance vaccine efficacy against infectious diseases. However, current methods for evaluating their efficacy and safety are limited, hindering large-scale screening. To address this, we developed a prototype Adjuvant Database (ADB) containing transcriptome data, generated using the same protocols as the widely used Open TG-GATEs (OTG) toxicogenomics database, covering 25 adjuvants across multiple species, organs, time points, and doses. This enabled cross-database integration of ADB and OTG. Transcriptomic patterns successfully distinguished each adjuvant regardless of organs or species. Using both databases, we built machine learning models to predict adjuvanticity and hepatotoxicity. Notably, we identified colchicine's adjuvant activity and FK565's liver toxicity through data-driven analysis. Overall, ADB combined with OTG offers a framework for transcriptomics-based, data-driven screening of adjuvant candidates.

Animals

A machine learning approach to identify key epigenetic transcripts for ageing research in human blood (Epitage).

DNA methylation is an established biomarker of human ageing and is used by a variety of tools to identify meaningful epigenetic signals. We investigated whether analysing CpGs grouped by transcript as functional units could generate a ranked list of transcripts most correlated with age that might otherwise be overlooked in genome-wide CpG-based studies. Here we present Epitage ( https://github.com/a00s/epitage ), a continuously updated ranked list of transcripts built from the GSE87571 dataset (714 whole-blood samples, ages 14-94 years) through intensive testing with machine-learning models. To support reproducible analyses, we developed ugPlot ( https://github.com/a00s/ugplot ), an open-source R/Shiny tool with a graphical user interface that automates model training, testing, and comparison. Initially, we identified 48 transcripts across 13 genes, with some transcripts from the genes OBSCN, PRRT1, and SPTBN4 showing better predictive performance when multiple associated CpGs were analysed together rather than individually. In contrast, for the majority of transcripts, a dominant individual CpG still showed a higher Spearman correlation with age, as seen in established ageing genes such as ELOVL2, FHL2, and TRIM59. Epitage is a transcript-ranking list based on the methylation patterns observed in the analysed dataset. It provides a reproducible framework for prioritising transcripts associated with human ageing and for guiding future epigenetic studies.

Humans

Antimicrobial resistance analysis of Klebsiella pneumoniae bloodstream infections based on a random forest algorithm: a longitudinal study based on data from tertiary hospitals in China from 2012 to 2023.

BACKGROUND: Bloodstream infections (BSIs) caused by Klebsiella pneumoniae pose a significant global health burden, complicated by rising antimicrobial resistance (AMR). This study aimed to characterize resistance patterns, identify predictors of carbapenem resistance, and develop a machine learning model to predict patient outcomes. METHODS: In a retrospective analysis of 109 279 K. pneumoniae BSIs from tertiary hospitals in China (2012-2023), 11&#x2009;000 isolates underwent whole-genome sequencing (WGS) and antimicrobial susceptibility testing. Cox proportional hazards and logistic regression models identified predictors of 30-day mortality and carbapenem-resistant K. pneumoniae (CRKP), respectively. A random forest model predicted AMR trends and outcomes, evaluated by accuracy, precision, recall, and ROC-AUC using R Studio (R Studio, Inc., Boston, MA, USA). RESULTS: Carbapenem resistance occurred in 32.3% of isolates, with rates of 41.9% for third-generation cephalosporins and 41.2% for fluoroquinolones. Among sequenced isolates, ST11 with blaKPC was the dominant CRKP genotype (12.0%). blaKPC (OR 3.97, 95% CI 3.10-5.11) and blaNDM (OR 2.80, 95% CI 2.07-3.71) strongly predicted carbapenem resistance; ICU admission predicted 30-day mortality (HR 2.10, 95% CI 1.80-2.46, p<0.001). Mortality was higher in CRKP (40.2%) vs. susceptible cases (21.5%). The random forest model achieved 89.2% accuracy and 0.92 ROC-AUC, with drug share, age, and CRKP status as top predictors. CONCLUSIONS: CRKP, especially ST11-blaKPC, drives excess mortality. Key predictors highlight the urgency for enhanced AMR surveillance and targeted therapy.

Humans

Natural language processing-based model to predict radiation pneumonitis in patients with locally advanced non-small cell lung cancer undergoing chemoradiotherapy: a retrospective cohort study.

BACKGROUND: Radiation pneumonitis (RP) remains a significant treatment-related toxicity in patients with unresectable, locally advanced non-small cell lung cancer (NSCLC) undergoing chemoradiotherapy (CRT). Most existing predictive models rely on static baseline demographic or dosimetry variables and lack real-time clinical applicability. We developed a novel predictive framework that integrates longitudinal symptom data extracted from clinical notes using natural language processing (NLP) with clinical and dosimetry features to improve early RP prediction. METHODS: We retrospectively identified 227 patients with locally advanced NSCLC treated with definitive CRT at a high-volume cancer center in the United States. We included all patients older than 18 years who were diagnosed between Jan 1, 2006, and Dec 31, 2022 with histologically or cytologically confirmed unresectable Stage 2 or 3 NSCLC and treated with conformal radiotherapy to a minimum dose of &#x2265;45 Gy with or without chemotherapy. Of these, 31 RP events were identified through manual adjudication using radiologic criteria and chart review. NLP was used to extract the temporal relationship of 16 pre-specified symptoms with treatment from over 100,000 clinical notes spanning pre- and during-treatment intervals. We trained and validated machine learning models on combinations of baseline clinical data, radiation dosimetry, and NLP-derived symptom features. Model performance was evaluated using a nested cross-validation framework, with an outer cross-validation loop reserved for performance assessment and an inner cross-validation loop used for model training and integration, and summarized using area under the receiver operating characteristic curve (AUC) and partial AUC (pAUC) at high specificity thresholds. Clinical utility was evaluated using decision curve analysis (DCA). FINDINGS: The best-performing model incorporated longitudinal NLP features and achieved a median AUC of 0.759 (90% confidence interval 0.753-0.766), significantly outperforming baseline models using only dosimetry (AUC 0.613) or clinical variables (AUC 0.635). NLP-based features such as cough trajectory, shortness of breath, and wheezing were among the most important predictors. Inclusion of NLP-derived symptom data improved early identification of high-risk patients, particularly in the clinically relevant high-specificity range (pAUC 0.021 vs. 0.010 for dosimetry alone). DCA showed that the calibrated MLP model provided greater net benefit than default strategies of treating all or no patients across clinically relevant threshold possibilities. INTERPRETATION: In this early work, NLP-based extraction of longitudinal symptoms from routine clinical documentation meaningfully enhances RP prediction in patients undergoing CRT for NSCLC. This approach leverages existing electronic health record infrastructure to deliver real-time, scalable, and interpretable risk estimates, offering a pathway toward potential early intervention and personalized toxicity management. The model and DCA requires external and prospective validation before clinical deployment; as such, future work should focus on this validation and integration into clinical decision support systems. FUNDING: AstraZeneca.

Chemoradiotherapy

Whole-genome phenotype prediction with machine learning: open problems in bacterial genomics.

MOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas).

Machine Learning

DNA Methylation-Based Classification of Kidney Neoplasms.

Renal neoplasms are morphologically and molecularly heterogeneous, with their diagnosis often hindered by interobserver variability and overlapping microscopic features. A subset of cases is unclassifiable despite immunohistochemical, mutation, and cytogenetic-based diagnostic workup. Through examination of the genome-wide DNA methylation signatures of over 2000 renal neoplasms, we identified 23 coherent groups that correlate with known neoplasm types and identified novel clinically relevant subtypes of existing neoplasm types. We used machine learning models to develop and validate a classifier trained on DNA methylation profiles of 1284 samples. The classifier was tested on an external data set of 287 renal neoplasms with >90% concordance between expected neoplasm type and high-score DNA methylation-based classification. Discordance between the original histologic label and methylation class led to potential reclassification of some cases. This work demonstrates proof of principle for the feasibility of a DNA methylation classifier as a clinically useful tool to assist in the diagnosis of renal neoplasms.

Humans

Investigating the mechanisms of PhIP-induced colorectal cancer through network toxicology, machine learning, and molecular dynamics simulation.

BACKGROUND: Over the past few years, 2-amino-1-methyl-6-phenylimidazo[4,5-b]pyridine (PhIP)- a compound from grilled or processed meats-has emerged as a major player in cancer development, especially colorectal cancer (CRC). This work dives into its potential links to CRC and uncovers the key genes that bridge this connection. METHODS: We tapped into various databases to pinpoint target genes tied to PhIP and CRC, then ran protein-protein interaction (PPI) analyses for visualization. Next, we explored underlying mechanisms through Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment. To nail down predictions, we tested 107 machine learning pipelines and picked the best one, validating its accuracy and the core genes' prognostic value across datasets. Next, molecular docking and dynamics simulations probed the interactions between these genes and PhIP. Finally, cell proliferation was assessed using Cell Counting Kit-8 (CCK-8) and 5-ethynyl-2'-deoxyuridine (EdU) assays, and polymerase chain reaction (PCR) was performed to validate the expression levels of the hub genes. RESULTS: Our analysis identified 39 overlapping genes, from which a machine learning model (glmBoost + Enet) identified six candidate targets: CDK4, CEBPB, COMT, SOX9, TIMP1, and TOP2A. To prioritize these, a hierarchical screening framework was applied. Molecular docking and dynamics simulations identified CDK4, COMT, and TIMP1 as the most stable interactors with PhIP. Functional assays confirmed that PhIP treatment significantly enhanced the proliferation of CRC cells. Crucially, quantitative PCR (qPCR) validation in multiple CRC cell lines identified TIMP1 as the primary target, showing the most consistent and significant upregulation upon PhIP exposure. CONCLUSIONS: In essence, these genes drive PhIP is role in CRC, offering novel insights into its molecular pathways. This could reshape how we tackle food-related pollutants, paving the way for better prevention and targeted therapies.

Colorectal cancer (CRC)

Noninvasive detection and differentiation of gastric malignancy using cell-free DNA biomarkers.

INTRODUCTION: Gastric cancer remains a major global health burden, with high mortality driven by late-stage diagnoses that limit treatment options and reduce survival. Current diagnostic methods such as endoscopy and biopsy are invasive, resource-intensive, and impractical for large-scale early detection. OBJECTIVES: This study aimed to develop and validate an ensemble machine learning model integrating four cell-free DNA (cfDNA) fragmentomic feature classes derived from 5&#xa0;&#xd7;&#xa0;whole genome sequencing (WGS) data to non-invasively differentiate malignant gastric cancer from benign gastric lesions in high-risk or symptomatic patients. METHODS: A total of 681 plasma samples were prospectively collected, comprising 329 from patients with gastric cancer or high-grade intraepithelial neoplasia (HGIN) and 352 from individuals with benign gastric conditions. The dataset was divided into a training cohort (n&#xa0;=&#xa0;333) and a temporally independent validation cohort (n&#xa0;=&#xa0;348). An external validation cohort of 305 participants was also included. RESULTS: The ensemble model achieved an AUROC of 0.920 in cross-validation testing on the training cohort, 0.912 in the independent validation cohort, and 0.896 (95% CI 0.860-0.932) in the external cohort. At a pre-specified prediction threshold of 0.402, the model demonstrated 93.3% sensitivity and 71.9% specificity in the validation cohort, yielding a PPV of 71.3% and an NPV of 93.5%. In the external cohort, sensitivity and specificity were 91.7% and 69.1%, respectively (PPV 75.7%, NPV 88.8%). Model scores correlated with clinical stage, tumor grade, and histopathological subtype. Approximately 71% of non-cancer patients could have been spared unnecessary endoscopy. CONCLUSIONS: The cfDNA fragmentomics-based ensemble model enables accurate, non-invasive differentiation between gastric cancer and benign gastric lesions in high-risk or symptomatic patients. This approach demonstrates strong potential as a pre-endoscopy triage tool, supporting earlier detection and more efficient use of diagnostic resources.

Humans

A Multi-omics Regulated Cell Death Framework Defines Immune Phenotypes and Guides Precision Therapy in Colorectal Cancer.

Colorectal cancer (CRC) is molecularly and immunologically heterogeneous, contributing to variable treatment response. Because regulated cell death (RCD) intersects with tumor metabolism, immune regulation, and therapeutic susceptibility, we built an RCD-centered framework for CRC stratification. Multi-cohort transcriptomic data were used to infer RCD subtypes with non-negative matrix factorization (NMF) and non-negative least squares (NNLS). Genomic, bulk RNA-seq, single-cell RNA-seq, and spatial transcriptomic datasets were integrated to characterize subtype-associated biology. Machine-learning models were developed for immunotherapy response and survival-risk estimation. Candidate compounds were screened by GDSC2-based drug-sensitivity modeling and molecular docking, and FSTL3 was functionally assessed in vitro. The framework separated CRC samples into two RCD-related phenotypes resembling immune-hot and immune-cold states. RCD1 showed immune activation and higher mutational burden, whereas RCD2 showed immune-suppressed features, intratumoral heterogeneity, and aggressive biology. RCD-associated signatures showed potential for predicting immunotherapy response and survival risk. Dasatinib was prioritized for immune-cold, high-risk tumors, with preliminary evidence supporting its activity in CRC cells, while functional assays suggested a role for FSTL3 in growth, invasion, epithelial-mesenchymal transition, and apoptosis regulation. These findings suggest that RCD-based multi-omics analysis may refine CRC stratification and help generate therapeutic hypotheses.

Colorectal cancer

Machine learning-based analysis of the impact of 5'&#xa0;untranslated region on protein expression.

The 5' untranslated region (5'UTR) plays a crucial regulatory role in messenger RNA (mRNA), with modified 5'UTRs extensively utilized in vaccine production, gene therapy, etc. Nevertheless, manually optimizing 5'UTRs may encounter difficulties in balancing the effects of various cis-elements. Consequently, multiple 5'UTR libraries have been created, and machine learning models have been employed to analyze and predict translation efficiency (TE) and protein expression, providing insights into critical regulatory features. On the one hand, these screening libraries, based on TE and mean ribosome load, struggle to accurately quantify protein expression; on the other hand, a precise method for quantifying 5'UTRs necessitates a significantly costlier library. To resolve this dilemma, we constructed a library utilizing firefly luciferase as the reporter to measure accurate protein expression. In addition, we optimized the library construction method by clustering mRNA sequences to reduce redundant data and minimize the size of the dataset. This dual strategy by increasing accuracy and reducing dataset size was found to be effective in predicting the 5'UTRs from the PC3 cell line.

5' Untranslated Regions

Prediction of gene expression using histone modification patterns extracted by Particle Swarm Optimization.

MOTIVATION: Histone modifications play an important role in transcription regulation. Although the general importance of some histone modifications for transcription regulation has been previously established, the relevance of others and their interaction is subject to ongoing research. By training Machine Learning models to predict a gene's expression and explaining their decision making process, we can get hints on how histone modifications affect transcription. In previous studies, trained models were either hardly explainable or the models were trained solely on the abundance of histone modifications. Based on other studies, which used histone modification patterns, rather than their abundance, to identify potential regulatory elements, we hypothesize the histone modification pattern in a gene's promoter to be more predictive for gene expression. We used an optimization algorithm to extract predictive histone modification profiles. RESULTS: Our algorithm called PatternChrome achieved an average area under curve (AUC) score of 0.9029 over 56 samples for binary classification, outperforming all previous algorithms for the same task. We explained the models decisions to deduce the effect of specific features, certain histone modifications or promoter positions on transcription regulation. Although the predictive histone modification patterns were extracted for each sample separately, they can be used to predict gene expression in other samples, implying that the created patterns are largely generalizable. Interestingly, the impact of histone modifications on gene regulation appears predominantly indifferent to cellular specificity. Through explanation of the classifier's decisions, we substantiate established literature knowledge while concurrently revealing novel insights into the intricate landscape of transcriptional regulation via histone modification. AVAILABILITY AND IMPLEMENTATION: The code for the PatternChrome algorithm, the scripts for the analyses and the required data can be found at (https://gitlab.gwdg.de/MedBioinf/generegulation/patternchrome).

Humans

Spatial concordance metrics and related risk factors of brain-peripheral barrier axes: unveiling distinct concordance patterns for mental and neurological axes.

Numerous studies have documented bidirectional interactions between the central nervous system and barrier organs (skin, gut, and lung). While genome-wide association studies have revealed shared genetic factors across brain-peripheral barrier axes, investigating these connections from an environmental perspective in large populations remains difficult. Using data from the Global Burden of Disease (GBD) 2023, I extracted annual incidence rates for 56 diseases related to brain-peripheral barrier axes and exposure rates for the 70 most detailed risk factors across 204 countries and territories. By categorizing regional incidence rates into four quartiles for each disease, I pinpointed regions with concordance of these axes and constructed a spatial atlas of disease concordance within the brain-peripheral barrier axis from a macro-epidemiologic view. Subsequently, I calculated global spatial concordance percentages for each axis, which allowed the comparatively assessment of concordance patterns across different axes, specific diseases, and their variations over time, across the lifespan, and by gender. Finally, I applied machine learning models and Shapley additive explanations to identify risk factors related to the spatial concordance of each axis. From 1990 to 2023, the overall trend for most brain-peripheral barrier axis pairs remained stable. Spatial concordance patterns showed dynamic fluctuations across the lifespan, followed by a convergence toward stability in older age. Several risk factors are related to most brain-peripheral barrier axes. Notably, the mental and neurological axes exhibited distinct concordance patterns. Compared with neurological axes, concordance within mental axes showed a broader and more dispersed geographic distribution, with greater variation across sexes and over time. Furthermore, concordance percentages of mental and neurological axes exhibited opposing age-related trends, contrasting disease spectra for peripheral conditions, and inverse relationships with alcohol and sodium consumption. Those divergences suggest distinct mechanisms underlying the brain-peripheral barrier axes in mental and neurological diseases. Related risk factors offer population-based hypotheses for further investigation in individual-level studies.

Humans

Inferring Gene Regulatory Networks in Stem Cells: Methods and Applications.

Gene regulatory networks (GRNs) represent the complex interplay of transcription factors, regulatory elements, and target genes that orchestrate cellular identity and function, playing a crucial role in the differentiation and maintenance of stem cells. This chapter provides an overview of experimental and computational methodologies for inferring GRNs, with particular emphasis on single-cell approaches. We first review key experimental techniques for detecting transcription factor binding sites, chromatin accessibility, and DNA motifs, alongside essential databases that support GRN reconstruction. We then introduce computational inference methods that can be categorized into four principal frameworks: correlation-based approaches, regression and machine learning models, probabilistic and deep learning methods, and integrative or message-passing frameworks. To illustrate practical application, we present a case study applying the pySCENIC workflow to a peripheral blood mononuclear cell single-cell RNA sequencing dataset from mouse, demonstrating how regulon-based analysis can reveal cell-type-specific regulatory programs. This chapter aims to serve as a practical guide for researchers seeking to understand and implement GRN inference methodologies in stem cell biology and related fields.

Gene Regulatory Networks

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in&#xa0;vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date &#x2248;150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements

A comprehensive meta-analysis of tissue resident memory T cells and their roles in shaping immune microenvironment and patient prognosis in non-small cell lung cancer.

Tissue-resident memory T cells (TRM) are a specialized subset of long-lived memory T cells that reside in peripheral tissues. However, the impact of TRM-related immunosurveillance on the tumor-immune microenvironment (TIME) and tumor progression across various non-small-cell lung cancer (NSCLC) patient populations is yet to be elucidated. Our comprehensive analysis of multiple independent single-cell and bulk RNA-seq datasets of patient NSCLC samples generated reliable, unique TRM signatures, through which we inferred the abundance of TRM in NSCLC. We discovered that TRM abundance is consistently positively correlated with CD4+ T helper 1 cells, M1 macrophages, and resting dendritic cells in the TIME. In addition, TRM signatures are strongly associated with immune checkpoint and stimulatory genes and the prognosis of NSCLC patients. A TRM-based machine learning model to predict patient survival was validated and an 18-gene risk score was further developed to effectively stratify patients into low-risk and high-risk categories, wherein patients with high-risk scores had significantly lower overall survival than patients with low-risk. The prognostic value of the risk score was independently validated by the Cancer Genome Atlas Program (TCGA) dataset and multiple independent NSCLC patient datasets. Notably, low-risk NSCLC patients with higher TRM infiltration exhibited enhanced T-cell immunity, nature killer cell activation, and other TIME immune responses related pathways, indicating a more active immune profile benefitting from immunotherapy. However, the TRM signature revealed low TRM abundance and a lack of prognostic association among lung squamous cell carcinoma patients in contrast to adenocarcinoma, indicating that the two NSCLC subtypes are driven by distinct TIMEs. Altogether, this study provides valuable insights into the complex interactions between TRM and TIME and their impact on NSCLC patient prognosis. The development of a simplified 18-gene risk score provides a practical prognostic marker for risk stratification.

Humans

Estimating the association of antimicrobial resistance genes with minimum inhibitory concentration in Escherichia coli: an observational study.

BACKGROUND: Surveillance and prediction of antibiotic resistance in Escherichia coli relies on curated databases of genes and mutations. We aimed to quantify the effect of acquiring specific genetic elements on minimum inhibitory concentrations (MICs) for particular antibiotic-species combinations, addressing the current scarcity of such data in existing databases. METHODS: For this observational study, we evaluated a collection of E coli isolates with linked whole-genome sequencing and MIC data, originating from human urinary or bloodstream infections obtained from the Oxford University Hospitals National Health Service Foundation Trust in Oxfordshire, UK. We used multivariable interval regression models to estimate the change in MIC (with 95% CIs) for specific antibiotics associated with the acquisition of antibiotic resistance genes and associated mutations in the National Center for Biotechnology Information AMRFinder database, with and without an adjustment for population structure. We then tested the ability of these models to predict MIC and binary resistance or susceptibility using leave-one-out cross-validation. FINDINGS: We evaluated 2875 E coli isolates obtained during 2013-2018 and 2020. Although most ARGs and resistance mutations (89 [80%] of 111) were associated with an increased MIC, a much smaller number (27 [24%] of 111) was found to be putatively independently resistance-conferring (ie, associated with an MIC above the European Committee on Antimicrobial Susceptibility Testing breakpoint) when acquired in isolation. We found evidence of differential effects of acquired ARGs and resistance mutations between different generations of cephalosporin antibiotics and showed that sub-breakpoint variation in MIC can be linked to genetic mechanisms of resistance. 20&#x2009;697 (83&#xb7;3%; range 52&#xb7;9-97&#xb7;7 across all antibiotics) of 24&#x2009;858 MICs were correctly exactly predicted and 23&#x2009;677 (95&#xb7;2%; 87&#xb7;3-97&#xb7;7) of 24&#x2009;858 MICs were predicted to within one doubling dilution. INTERPRETATION: Quantitative estimates of the independent effect of the acquisition of ARGs on MIC add to the interpretability and utility of existing databases. Compared with approaches using machine learning models, the use of these estimates yields similar or better performance in the prediction of antibiotic resistance phenotype with more readily interpretable results. The methods outlined here could be readily applied to other antibiotic-pathogen combinations. FUNDING: The National Institute for Health and Care Research (NIHR) and the Medical Research Council (MRC).

Escherichia coli

NMR metabolomics and glycomics for cancer detection in patients with non-specific symptoms: a prospective observational cohort study.

BACKGROUND: Early cancer diagnosis in patients with non-specific symptoms is limited by the lack of discriminatory tests. Within the Oxfordshire Suspected CANcer (SCAN) pathway, exploratory biomarker work showed that serum 1H NMR-based metabolomics can identify cancer with high accuracy. SCAN2 evaluated whether integrating metabolomics with glycomics provides complementary molecular information and improves discrimination in a clinically complex, real-world population. METHODS: Serum from 369 SCAN patients (59 cancers) was analysed using AXINON&#xae; System-derived NMR metabolomics and HPLC-MS glycomics. Machine-learning models were trained to predict cancer status, with performance assessed by receiver operating characteristic (ROC) analysis of pooled cross-validated predictions. To place cancer risk in a broader clinical context, a second classifier modelling alternative non-cancer diagnosis was incorporated, and mean predicted probabilities from both models were jointly projected into a two-dimensional space, maintaining strict separation of training and test data. FINDINGS: In the full cohort, integration of glycomics with metabolomics achieved an AUC of 0.814 (95% CI 0.808-0.820). In a refined sub-cohort excluding major comorbidities and selected cancer types (32 cancers, 277 non-cancers), performance improved to an AUC of 0.884 (95% CI 0.879-0.890). Discriminatory features included cancer-associated biantennary fucosylated glycans alongside amino acid metabolites (glutamate, histidine) and lipoprotein-related measures. A classifier distinguishing metastatic from non-metastatic disease (n = 29 vs. 30) achieved an AUC of 0.80. Joint probability analysis in the full cohort preserved cancer-associated signatures across comorbidity burden, with projection-based classification achieving an accuracy of 89.2% (95% CI 85.7-92.6). INTERPRETATION: These findings validate the SCAN1 metabolomic signature in a more clinically complex cohort and indicate that integrating glycomics with metabolomics provides complementary biological information for cancer discrimination. Joint probability analysis provides an interpretable framework for cancer risk stratification within multimorbid diagnostic pathways, supporting the clinical potential of scalable multi-omics blood testing. FUNDING: EPSRC, EU Horizon 2020, Wellcome/MLSTF, Novo Nordisk Foundation.

Humans