Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning

Subphenogroups of acute heart failure with preserved ejection fraction: comprehensive proteomics and pathway analysis.

BACKGROUND: Heterogeneity of heart failure with preserved ejection fraction (HFpEF) results in significant challenges for treatment development. Identifying and characterising distinct HFpEF phenogroups may aid in tailoring therapeutic strategies for these patients. The objective of this study was to assess proteomic patterns of HFpEF phenogroups identified through a machine-learning-based clustering model, with the aim of uncovering specific biological pathways associated with each phenogroup. METHODS: This study represents a post-hoc analysis of the ongoing Prospective mUlticenteR obServational stUdy of patIenTs with Heart Failure with preserved Ejection Fraction (PURSUIT-HFpEF) study, which is a multicentre prospective observational study of hospitalised patients with acute decompensated HFpEF. Of the overall cohort (N=1238), this study analysed 198 patients with HFpEF with available proteomics data. These patients were classified into four phenogroups using the machine-learning-based clustering model. The SomaScan assay V.4.1 was used to measure levels of >7000 plasma proteins, and subsequent pathway analysis was conducted to determine the biological differences among the phenogroups. RESULTS: We identified four distinct phenogroups: Phenogroup 1 ('rhythm trouble'), Phenogroup 2 ('ventricular-arterial uncoupling'), Phenogroup 3 ('low output and systemic congestion') and Phenogroup 4 ('systemic failure'). The proteomics revealed distinct protein expression profiles among the phenogroups, with ribonuclease 4, tax1-binding protein 1, regenerating islet-derived protein 3-gamma and alpha-1-antichymotrypsin being the most significant markers to specific identified phenogroups. Pathway analysis suggested differences in immune response, autonomic activation, cellular homeostasis and tissue repair mechanisms across the phenogroups. CONCLUSIONS: Using a comprehensive plasma proteomics approach, our study identified distinct proteomic profiles of HFpEF phenogroups, which in turn suggest specific underlying biological processes. These profiles suggest the involvement of inflammatory activation, tissue injury and regenerative responses, immune modulation and systemic stress signalling as key components of HFpEF pathophysiology. TRIAL REGISTRATION NUMBER: UMIN-CTR ID: UMIN000021831.

Humans

Chromatin structures from integrated AI and polymer physics model.

The physical organization of the genome in three-dimensional space regulates many biological processes, including gene expression and cell differentiation. Three-dimensional characterization of genome structure is critical to understanding these biological processes. Direct experimental measurements of genome structure are challenging; computational models of chromatin structure are therefore necessary. We develop an approach that combines a particle-based chromatin polymer model, molecular simulation, and machine learning to efficiently and accurately estimate chromatin structure from indirect measures of genome structure. More specifically, we introduce a new approach where the interaction parameters of the polymer model are extracted from experimental Hi-C data using a graph neural network (GNN). We train the GNN on simulated data from the underlying polymer model, avoiding the need for large quantities of experimental data. The resulting approach accurately estimates chromatin structures across all chromosomes and across several experimental cell lines despite being trained almost exclusively on simulated data. The proposed approach can be viewed as a general framework for combining physical modeling with machine learning, and it could be extended to integrate additional biological data modalities. Ultimately, we achieve accurate and high-throughput estimations of chromatin structure from Hi-C data, which will be necessary as experimental methodologies, such as single-cell Hi-C, improve.

Chromatin

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Machine learning-guided risk stratification in elderly AML based on genomic, immunophenotypic and therapeutic profiles.

BACKGROUND: Elderly patients with acute myeloid leukemia (AML) exhibit considerable biological and clinical heterogeneity, hindering precise prognosis. Existing prognostic systems inadequately capture the complexity of elderly AML due to their reliance on data from younger cohorts and omission of key factors like immunophenotypic markers and therapeutic profiles. This study aimed to develop and internally validate a machine learning-based prognostic model specifically tailored to elderly AML patients. METHODS: A total of 156 patients were analyzed using a two-stage modeling strategy. Clinical and genomic variables were modeled first, followed by independent analysis of immunophenotypic features. Feature selection was performed using multilayer perceptron (MLP) and random forest (RF), while multivariate Cox regression was used for final model construction. Internal validation was conducted using 1000 bootstrap iterations to assess model stability and performance. RESULTS: The model demonstrated strong predictive performance, with a concordance index (C-index) of 0.702. Time-dependent area under the curve (AUC) and calibration plots confirmed accurate prediction of 1-, 3-, and 5-year overall survival. Decision curve analysis indicated favorable net benefit across a range of threshold probabilities. Key independent prognostic factors identified included TP53 mutations, high CD13 expression, and IDH2 mutations. CONCLUSION: This model provides a robust and interpretable tool for individualized risk stratification in elderly AML. By integrating genomic, immunophenotypic, and therapeutic variables, it may help optimize treatment decisions and improve outcomes for this vulnerable population. Future efforts should focus on external validation and integration of dynamic biomarkers.

Humans

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

AI-driven CRISPR screening: optimizing gene editing through automation and intelligent decision support.

BACKGROUND: CRISPR-based genetic screening has become a central methodology in functional genomics, enabling systematic interrogation of gene function, genetic interactions and context-dependent vulnerabilities at scale. However, the rapid expansion of screening modalities-including multi-condition designs, combinatorial perturbations, in vivo applications and single-cell readouts-has exposed fundamental limitations of heuristic-driven experimental design and post hoc statistical analysis. MAIN BODY: This Review synthesizes how artificial intelligence is reshaping CRISPR screening by introducing predictive, adaptive and system-level intelligence across the experimental lifecycle. We organize recent advances into two tightly coupled modules. First, machine learning and deep learning (ML/DL) methods optimize experimental design by learning context-dependent perturbation behavior, anticipating confounding effects and enabling iterative, information-efficient screening strategies. Second, large language model-agent (LLM-agent) systems complement these advances by externalizing scientific reasoning, integrating biological knowledge at scale and coordinating analysis and decision-making in human-in-the-loop workflows. CONCLUSIONS: Together, ML/DL and LLM-agent approaches reframe CRISPR screening from a static analytical pipeline into an intelligent experimental system, with important implications for robustness, scalability and biological discovery.

Artificial Intelligence

Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.

MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.

CRISPR-Cas Systems

Modelling the structure and function of enzymes by machine learning.

A machine learning program, GOLEM, has been applied to two problems: (1) the prediction of protein secondary structure from sequence and (2) modelling a quantitative structure-activity relationship in drug design. GOLEM takes as input observations and combines them with background knowledge of chemistry to yield rules expressed as stereochemical principles for prediction. The secondary structure prediction was explored on the alpha/alpha class of proteins; on an unrelated test set it yielded 81% accuracy. The rules from GOLEM defined patterns of residues forming alpha-helices. The system studied for drug design was the activities of trimethoprim analogues binding to E. coli dihydrofolate reductase. The GOLEM rules were a better model than standard regression approaches. More importantly, these rules described the chemical properties of the enzyme-binding site that were in broad agreement with the crystallographic structure.

Amino Acid Sequence

Enhancing detection of polygenic adaptation: a comparative study of machine learning and statistical approaches using simulated evolve-and-resequence data.

BACKGROUND: Detecting signals of polygenic adaptation remains a significant challenge in population genomics, as traditional methods often struggle to identify the associated subtle, multi-locus allele-frequency shifts. Here, we introduced and tested several novel approaches combining machine learning techniques with traditional statistical tests to detect polygenic adaptation patterns in time-series of allele frequency changes from whole genome data. We implemented a Naive Bayesian Classifier (NBC) and One-Class Support Vector Machines (OCSVM), and compared their performance against the classical Fisher's Exact Test (FET). Furthermore, we combined machine learning and statistical models (OCSVM-FET and NBC-FET), resulting in 5 competing approaches. The framework is mainly designed and validated for evolve-and-resequence (EaR) experimental designs, where defined selection pressures and temporal sampling are feasible, but might be applicable for certain natural experiments as well. RESULTS: Using a simulated dataset based on empirical C. riparius Pool-Seq data, we evaluated methods across evolutionary scenarios varying in generation, selection strength, and number of loci under selection. Our results demonstrate that the combined OCSVM-FET approach consistently outperformed competing methods, achieving the lowest false positive rate, highest area under the curve, and high accuracy. The performance peak aligned with what we term the 'late dynamic phase' of adaptation - the period after initial selection has occurred but before fixation - highlighting the method's sensitivity to ongoing selective processes. CONCLUSIONS: Furthermore, we emphasize the critical role of parameter tuning, balancing biological assumptions with methodological rigor. While broader applicability remains an important direction for future work, the present benchmarking is intentionally scoped to EaR experimental contexts.

Machine Learning

Integrating Imaging-Derived Clinical Endotypes with Plasma Proteomics and External Polygenic Risk Scores Enhances Coronary Microvascular Disease Risk Prediction.

Coronary microvascular disease (CMVD) is an underdiagnosed but significant contributor to the burden of ischemic heart disease, characterized by angina and myocardial infarction. The development of risk prediction models such as polygenic risk scores (PRS) for CMVD has been limited by a lack of large-scale genome-wide association studies (GWAS). However, there is significant overlap between CMVD and enrollment criteria for coronary artery disease (CAD) GWAS. In this study, we developed CMVD PRS models by selecting variants identified in a CMVD GWAS and applying weights from an external CAD GWAS, using CMVD-associated loci as proxies for the genetic risk. We integrated plasma proteomics, clinical measures from perfusion PET imaging, and PRS to evaluate their contributions to CMVD risk prediction in comprehensive machine and deep learning models. We then developed a novel unsupervised endotyping framework for CMVD from perfusion PET-derived myocardial blood flow data, revealing distinct patient subgroups beyond traditional case-control definitions. This imaging-based stratification substantially improved classification performance alongside plasma proteomics and PRS, achieving AUROCs between 0.65 and 0.73 per class, significantly outperforming binary classifiers and existing clinical models, highlighting the potential of this stratification approach to enable more precise and personalized diagnosis by capturing the underlying heterogeneity of CMVD. This work represents the first application of imaging-based endotyping and the integration of genetic and proteomic data for CMVD risk prediction, establishing a framework for multimodal modeling in complex diseases.

Cardiovascular Disease

Dissecting genetic architecture and improving machine learning&#x2011;based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC&#xa0;>&#xa0;0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

Sex Hormone Receptors, HBV Integrations and Their Prognostic Predictive Value Among Hepatocellular Carcinoma Patients.

Hepatocellular carcinoma (HCC) related to hepatitis B virus (HBV) infection predominantly affects males, yet few studies have investigated the association between sex hormones and HBV integrations, and their involvement in HCC prognosis. We assessed estrogen receptor alpha (ER&#x3b1;) and androgen receptor (AR) expression via immunohistochemistry on tissue microarrays constructed from 426 HBV-related HCC samples. HBV integration features were determined using HBV-captured sequencing data. Logistic regression models were utilized to evaluate the association between sex hormone receptor expression level and HBV integration features. Cox regression models, combined with machine learning (ML) methods, were implemented to investigate the prognostic value of sex hormone receptors and HBV integrations concerning overall survival. We found high AR expression level was significantly associated with higher HBV integration levels (adjusted odds ratio [aOR]&#x2009;=&#x2009;1.84, 95% confidence interval [CI]: 1.09-3.11, P for trend&#x2009;=&#x2009;0.012), TERT integration (aOR&#x2009;=&#x2009;2.34, 95% CI: 1.16-4.74, P for trend&#x2009;=&#x2009;0.047), intergenic integration (aOR&#x2009;=&#x2009;2.25, 95% CI: 1.20-4.24, P for trend&#x2009;=&#x2009;0.021), and promoter integration (aOR&#x2009;=&#x2009;1.81, 95% CI: 1.00-3.31, P for trend&#x2009;=&#x2009;0.034). The inclusion of sex hormone receptors and HBV integrations in the predictive models led to improvements across all performance metrics in the Cox regression analyses (AUC improvement: 0.014 [Training], 0.026 [Validation]) and the ML (AUC improvement: 0.022 [Training]), although a slight deterioration in performance was noted in the ML validation set. The results suggested a relationship between AR expression level and HBV integration events, as well as the potential utility of HBV integration biomarkers and sex hormone receptor profiles in assessing post-surgical prognosis among HCC patients.

Humans

A multi-scale fusion model based on multi-phase contrast-enhanced CT for predicting pancreatic cancer resectability.

Purpose.Develop a multi-scale fusion model (MSFM) based on multi-phase contrast-enhanced computed tomography (CECT) to predict pancreatic cancer (PC) resectability, thereby assisting expert decision-making.Methods.This retrospective study enrolled 280 patients with PC from four institutions, which were randomly divided into a training cohort (202 patients) and an independent test cohort (78 patients). Three-phase CECT images (arterial, venous, and delayed phases) were used for modeling. The MSFM comprises two sub-networks: (1) a multi-phase fusion network for extracting cross-phase shared fusion features, (2) a phase-specific branch network for capturing phase-specific features; and a post-fusion strategy to generate the final predictive score by integrating the shared fusion features and three groups of phase-specific features. Additionally, a human-machine fusion deep learning model (HMfDL) was constructed by fusing the predictive score of the MSFM with expert assessments.Results.In the independent test, the MSFM achieved an AUC (area under the receiver operating characteristic curve) of 0.8385 (95% CI: 0.7521-0.9249), accuracy of 84.62%, sensitivity of 72.00%, and specificity of 90.57%. This performance outperformed single-phase models (AUC range: 0.7638-0.7781), two-phase models (AUC range: 0.7826-0.7864), and ten states-of-the-art classifiers (AUC range: 0.7404-0.7796). The HMfDL further improved the performance, reaching an AUC of 0.8626 (95% CI: 0.7853-0.9400), accuracy of 91.03%, sensitivity of 80.00%, and specificity of 96.23%. Notably, the HMfDL corrected 58.82% of misdiagnosis made by experts.Conclusions. The MSFM effectively fuses multi-phase CECT to enable highly accurate predictions of PC resectability, and provides valuable support for expert decision-making through HMfDL.

Humans

Functional neuroimaging subtypes of obsessive-compulsive disorder: A systematic review and meta-analysis.

Obsessive-compulsive disorder (OCD) exhibits substantial clinical heterogeneity that may reflect underlying neurobiological diversity. Neuroimaging-based subtyping may advance precision psychiatry by identifying biologically distinct subgroups with differential treatment responses. This study systematically synthesized evidence from functional neuroimaging subtyping studies in OCD to identify reproducible neurobiological subtypes, characterize their clinical profiles, and establish a consensus-based classification framework. We reviewed 40 original studies employing machine learning, clustering, normative modeling, or classification approaches, encompassing approximately 8,150 patients. Consensus clustering identified three reproducible neurobiological subtypes. The Limbic-Hyperactive subtype, comprising approximately 40% of patients, exhibited amygdala and insula hyperconnectivity, elevated anxiety levels, predominant contamination and washing symptoms, and favorable response to cognitive-behavioral therapy. The Fronto-Striatal-Hypoconnected subtype, comprising approximately 35% of patients, demonstrated reduced orbitofrontal-striatal connectivity, cognitive inflexibility, predominant checking and ordering symptoms, and a favorable response to selective serotonin reuptake inhibitors. The Global-Disrupted subtype, comprising approximately 25% of patients, exhibited widespread connectivity disruption, greater symptom severity, and poor treatment response. Support vector machine classification achieved 81.5% accuracy for subtype assignment, though classification of OCD versus healthy controls showed limited generalizability in multisite settings (AUC 0.567-0.673). These findings support a neuroimaging-based framework for personalized treatment selection but require prospective validation.

Humans

Deciphering microbial and metabolic influences in gastrointestinal diseases-unveiling their roles in&#xa0;gastric cancer, colorectal cancer, and inflammatory bowel disease.

INTRODUCTION: Gastrointestinal disorders (GIDs) affect nearly 40% of the global population, with gut microbiome-metabolome interactions playing a crucial role in gastric cancer (GC), colorectal cancer (CRC), and inflammatory bowel disease (IBD). This study aims to investigate how microbial and metabolic alterations contribute to disease development and assess whether biomarkers identified in one disease could potentially be used to predict another, highlighting cross-disease applicability. METHODS: Microbiome and metabolome datasets from Erawijantari et al. (GC: n&#x2009;=&#x2009;42, Healthy: n&#x2009;=&#x2009;54), Franzosa et al. (IBD: n&#x2009;=&#x2009;164, Healthy: n&#x2009;=&#x2009;56), and Yachida et al. (CRC: n&#x2009;=&#x2009;150, Healthy: n = 127) were subjected to three machine learning algorithms, eXtreme gradient boosting (XGBoost), Random Forest, and Least Absolute Shrinkage and Selection Operator (LASSO). Feature selection identified microbial and metabolite biomarkers unique to each disease and shared across conditions. A microbial community (MICOM) model simulated gut microbial growth and metabolite fluxes, revealing metabolic differences between healthy and diseased states. Finally, network analysis uncovered metabolite clusters associated with disease traits. RESULTS: Combined machine learning models demonstrated strong predictive performance, with Random Forest achieving the highest Area Under the Curve(AUC) scores for GC(0.94[0.83-1.00]), CRC (0.75[0.62-0.86]), and IBD (0.93[0.86-0.98]). These models were then employed for cross-disease analysis, revealing that models trained on GC data successfully predicted IBD biomarkers, while CRC models predicted GC biomarkers with optimal performance scores. CONCLUSION: These findings emphasize the potential of microbial and metabolic profiling in cross-disease characterization particularly for GIDs, advancing biomarker discovery for improved diagnostics and targeted therapies.

Humans