Search PubMedSearch

SEARCH · Search PubMed

Results for “Deep Neural Network”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Transfer learning with multiomics integration and deep neural networks reveals drug resistance mechanisms in cancer.

Drug resistance remains one of the primary challenges in effective cancer therapy. In this study, we employed a deep neural network (DNN)-based transfer learning (TL) approach to predict drug response and uncover drug resistance mechanisms. We integrated gene expression, somatic mutation, and copy number aberration (CNA) data with drug response profiles using multi-omics integration (MI). We used the Genomics of Drug Sensitivity in Cancer (GDSC) data for training and incorporated drugs with same pathways into the training models. We then evaluated drug response predictions on independent in-vivo PDX Encyclopedia (PDX) and ex-vivo the Cancer Genome Atlas (TCGA) datasets. In addition, we conducted pathway enrichment analyses to elucidate the mechanisms underlying drug resistance for paclitaxel, 5-fluorouracil (5-FU), gemcitabine, and cetuximab. We also applied Fisher's exact test (FET) to assess potential associations between drug resistance and the presence of mutations or CNAs. Our pan-drug models outperformed other methods based on the area under the precision-recall curve (AUCPR). Our pathway enrichment analyses revealed LDHB-mediated pyruvate metabolism and FYN-mediated focal adhesion might have pivotal roles in paclitaxel resistance, while PINK1-mediated mitophagy might be critical in 5-FU resistance. In addition to transcriptional activation, FET suggested that CNAs in LDHB and PINK1 may also be associated with resistance to paclitaxel and 5-FU, respectively. Furthermore, enrichment results for paclitaxel and cetuximab indicated shared resistance mechanisms between the two drugs. Importantly, our findings are consistent with prior experimental studies, providing literature-based validation of our results. Overall, our DNN-based TL approach achieved strong predictive performance across PDX & TCGA datasets and enrichment analyses provided valuable biological insights into drug resistance mechanisms.

Humans

Deep generative neural network for accurate drug response imputation.

Drug response differs substantially in cancer patients due to inter- and intra-tumor heterogeneity. Particularly, transcriptome context, especially tumor microenvironment, has been shown playing a significant role in shaping the actual treatment outcome. In this study, we develop a deep variational autoencoder (VAE) model to compress thousands of genes into latent vectors in a low-dimensional space. We then demonstrate that these encoded vectors could accurately impute drug response, outperform standard signature-gene based approaches, and appropriately control the overfitting problem. We apply rigorous quality assessment and validation, including assessing the impact of cell line lineage, cross-validation, cross-panel evaluation, and application in independent clinical data sets, to warrant the accuracy of the imputed drug response in both cell lines and cancer samples. Specifically, the expression-regulated component (EReX) of the observed drug response achieves high correlation across panels. Using the well-trained models, we impute drug response of The Cancer Genome Atlas data and investigate the features and signatures associated with the imputed drug response, including cell line origins, somatic mutations and tumor mutation burdens, tumor microenvironment, and confounding factors. In summary, our deep learning method and the results are useful for the study of signatures and markers of drug response.

Antineoplastic Agents

Prediction of Atrial Fibrillation From the ECG in the Community Using Deep Learning: A Multinational Study.

BACKGROUND: We aimed to refine and validate a deep neural network model from the ECG to predict atrial fibrillation (AF) risk, using samples from diverse backgrounds: the Framingham Heart Study (FHS), UK Biobank, and Estudo Longitudinal da Saúde do Adulto (ELSA-Brasil). We compared the model's performance to the clinical Cohorts for Heart and Aging Research in Genomic Epidemiology consortium (CHARGE-AF) risk score and evaluated the association with other cardiovascular outcomes. METHODS: The ECG-derived deep-learning prediction of AF (ECG-AF) model was refined using 60% of FHS samples free of AF. Its performance was then tested in the remaining FHS samples, UK Biobank, and ELSA-Brasil, with discrimination assessed by the area under the receiver operating characteristic curve. The association of ECG-AF with cardiovascular outcomes was assessed using Cox proportional hazards models. RESULTS: The study sample included 10 097 FHS participants (mean age 53±12 years; 54.9% women), 49 280 participants from the UK Biobank (mean age 64±8 years, 47.9% women), and 12 284 participants from ELSA-Brasil (mean age 53±8 years, 54.7% women). The ECG-AF model showed moderate discrimination for incident AF (area under the curve, 0.82 [95% CI, 0.80-0.84]) in the FHS, comparable to the CHARGE-AF score (area under the curve, 0.83 [95% CI, 0.81-0.85]), and incremental when combined (area under the curve, 0.85 [95% CI, 0.83-0.87]). In UK Biobank and ELSA-Brasil, combining ECG-AF and CHARGE also improved prediction. Higher ECG-AF scores were associated with increased risks of heart failure, myocardial infarction, stroke, and all-cause mortality in all 3 cohorts. CONCLUSIONS: In multinational cohort studies, the single-input ECG-AF deep neural network model demonstrated good performance in predicting AF and other cardiovascular outcomes, comparable to a multivariable clinical risk score, with improved performance when combined.

Humans

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6 mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6 mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6 mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6 mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6 mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning

A multi-modal transformer for cell type-agnostic regulatory predictions.

Sequence-based deep learning models have emerged as powerful tools for deciphering the cis-regulatory grammar of the human genome but cannot generalize to unobserved cellular contexts. Here, we present EpiBERT, a multi-modal transformer that learns generalizable representations of genomic sequence and cell type-specific chromatin accessibility through a masked accessibility-based pre-training objective. Following pre-training, EpiBERT can be fine-tuned for gene expression prediction, achieving accuracy comparable to the sequence-only Enformer model, while also being able to generalize to unobserved cell states. The learned representations are interpretable and useful for predicting chromatin accessibility quantitative trait loci (caQTLs), regulatory motifs, and enhancer-gene links. Our work represents a step toward improving the generalization of sequence-based deep neural networks in regulatory genomics.

Humans

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis

Deep-Learning Model for Tumor-Type Prediction Using Targeted Clinical Genomic Sequencing Data.

UNLABELLED: Tumor type guides clinical treatment decisions in cancer, but histology-based diagnosis remains challenging. Genomic alterations are highly diagnostic of tumor type, and tumor-type classifiers trained on genomic features have been explored, but the most accurate methods are not clinically feasible, relying on features derived from whole-genome sequencing (WGS), or predicting across limited cancer types. We use genomic features from a data set of 39,787 solid tumors sequenced using a clinically targeted cancer gene panel to develop Genome-Derived-Diagnosis Ensemble (GDD-ENS): a hyperparameter ensemble for classifying tumor type using deep neural networks. GDD-ENS achieves 93% accuracy for high-confidence predictions across 38 cancer types, rivaling the performance of WGS-based methods. GDD-ENS can also guide diagnoses of rare type and cancers of unknown primary and incorporate patient-specific clinical information for improved predictions. Overall, integrating GDD-ENS into prospective clinical sequencing workflows could provide clinically relevant tumor-type predictions to guide treatment decisions in real time. SIGNIFICANCE: We describe a highly accurate tumor-type prediction model, designed specifically for clinical implementation. Our model relies only on widely used cancer gene panel sequencing data, predicts across 38 distinct cancer types, and supports integration of patient-specific nongenomic information for enhanced decision support in challenging diagnostic situations. See related commentary by Garg, p. 906. This article is featured in Selected Articles from This Issue, p. 897.

Humans

abCRISPR: deep learning-based design of abasic gRNA sequences for specific CRISPR-Cas9 genome editing.

SUMMARY: CRISPR-Cas9 has become a widely used tool for genome editing. However, its off-target cleavage caused by partial sequence matches with guide RNAs (gRNAs) remains a critical limitation. Recently, abasic gRNAs (ØXØ) have been developed to enhance target specificity, but their effects vary depending on the positional sequence context. Here, we present abCRISPR, a deep neural network (DNN) framework for the rational design of ØXØ sequences with minimized off-target activity. abCRISPR leverages informative few-shot training with paired datasets of abasic and unmodified gRNAs, using high-quality random mismatch target libraries, exhaustively sequenced for mismatched off-target substrates (n = 97583) in in vitro CRISPR-Cas9 cleavage experiments. Predicted off-target activities for both abasic and unmodified gRNAs showed strong correlation with experimental data (r ≥ 0.95, 10-fold cross-validation). Notably, these comprehensive training sets provide robust ground-truth negatives, enabling accurate and sensitive prediction of off-targets. For unmodified gRNAs, abCRISPR (AUC = 0.98) was validated to outperform existing deep learning-based methods (AUC = 0.45-0.68). When applied to the human genome, abCRISPR generated ØXØ sequences, covering 58 875 004 potent CRISPR-targetable sites with improved target specificity. Together, this work provides a comprehensive bioinformatics resource for safe and precise CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The source code for abCRISPR and training data are available at https://doi.org/10.5281/zenodo.20398246. abCRISPR results for the human genome are available at http://clip.korea.ac.kr/abCRISPR/.

Deep Learning

AI-driven multi-omics modeling of myalgic encephalomyelitis/chronic fatigue syndrome.

Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a chronic illness with a multifactorial etiology and heterogeneous symptomatology, posing major challenges for diagnosis and treatment. Here we present BioMapAI, a supervised deep neural network trained on a 4-year, longitudinal, multi-omics dataset from 249 participants, which integrates gut metagenomics, plasma metabolomics, immune cell profiling, blood laboratory data and detailed clinical symptoms. By simultaneously modeling these diverse data types to predict clinical severity, BioMapAI identifies disease- and symptom-specific biomarkers and classifies ME/CFS in both held-out and independent external cohorts. Using an explainable AI approach, we construct a unique connectivity map spanning the microbiome, immune system and plasma metabolome in health and ME/CFS adjusted for age, gender and additional clinical factors. This map uncovers altered associations between microbial metabolism (for example, short-chain fatty acids, branched-chain amino acids, tryptophan, benzoate), plasma lipids and bile acids, and heightened inflammatory responses in mucosal and inflammatory T cell subsets (MAIT, γδT) secreting IFN-γ and GzA. Overall, BioMapAI provides unprecedented systems-level insights into ME/CFS, refining existing hypotheses and hypothesizing unique mechanisms-specifically, how multi-omics dynamics are associated to the disease's heterogeneous symptoms.

Humans

Pretraining improves prediction of genomic datasets across species.

MOTIVATION: Recent studies suggest that deep neural network models trained on thousands of human genomic datasets can accurately predict genomic features, including gene expression and chromatin accessibility. However, training these models is computation- and time-intensive, and datasets of comparable size do not exist for most other organisms. RESULTS: Here, we identify modifications to an existing state-of-the-art model that improve model accuracy while reducing training time and computational cost. Using this streamlined model architecture, we investigate the ability of models pretrained on human genomic datasets to transfer performance to a variety of different tasks. Models pretrained on human data but fine-tuned on genomic datasets from diverse tissues and species achieved significantly higher prediction accuracy while significantly reducing training time compared to models trained from scratch, with Pearson correlation coefficients between experimental results and predictions as high as 0.8. Further, we found that including excessive training tasks decreased model performance and that this decrease could be partially but not completely rescued by fine-tuning. Thus, simplifying model architecture, applying pretrained models, and carefully considering the number of training tasks may be effective and economical techniques for building new models across data types, tissues, and species. AVAILABILITY AND IMPLEMENTATION: Code is available on GitHub and Figshare: https://github.com/optimizedlearning/genomicsML, https://doi.org/10.6084/m9.figshare.31796116.

Genomics

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans

Identification of MMP14 and MKLN1 as colorectal cancer susceptibility genes and drug-repositioning candidates from a genome-wide association study.

BACKGROUND: Genome-wide association studies (GWAS) and subsequent functional interpretation have been used to identify susceptible genes and potential drug-repositioning candidates. This study aimed to identify genes associated with colorectal cancer (CRC) and potential drug-repositioning candidates. METHODS: Patients with CRC at Seoul National University Hospital (SNUH, discovery study) and Chonnam National University Hospital (CNUH, replication study) were included as case groups. The Korean Genome and Epidemiology Study (KoGES) participants were included as a control group. Single-nucleotide polymorphisms (SNPs) were extracted from blood-derived DNA (N&#x2009;=&#x2009;409,063). A SNP-based logistic regression model was applied. Furthermore, post-GWAS analysis was conducted. Drug-repositioning candidates were identified using a pre-trained deep neural network and the druggability assessment tool. RESULTS: In the discovery study, we conducted a 1:3 age- and sex-matched case-control study that included 500 CRC cases (mean age 63.0&#x2009;&#xb1;&#x2009;7.15&#xa0;years) and 1,500 healthy controls (mean age 62.9&#x2009;&#xb1;&#x2009;7.07&#xa0;years), each group comprising 50% males and 50% females. The replication study enrolled 4,860 patients with CRC and 46,384 healthy controls. The two-stage GWAS revealed statistically significant associations among MKLN1 (rs75170436, 7q32.3, beta (log odds ratio)&#x2009;= -&#x2009;0.90, Pmeta&#x2009;=&#x2009;5.90&#x2009;&#xd7;&#x2009;10-13), MMP14 (rs3751489, 14q11.2, beta (log odds ratio)&#x2009;= -&#x2009;1.91, Pmeta&#x2009;=&#x2009;2.31&#x2009;&#xd7;&#x2009;10-12). Post-GWAS functional analysis revealed strong associations on two genes highlighting deleterious effects and increased gene expression. Drug-repositioning analysis identified GW0742 (PPAR&#x3b2;/&#x3b4; agonist) with the highest binding score and druggability score for MMP14 with a reference allele (12.06, 0.85). CONCLUSIONS: Using GWAS, MKLN1 and MMP14 were found to be associated with CRC development and we identified GW0742 (PPAR&#x3b2;/&#x3b4; agonist) as a potential drug-repositioning candidate for CRC based on MKLN1 and MMP14. These findings improve the understanding of CRC development and provide insights into novel therapeutic targets and candidates for CRC treatment.

Humans

Integration of single cell multiomics data by deep transfer hypergraph neural network.

Multi-omics characterization of individual cells offers remarkable potential for analyzing the dynamics and relationships of gene regulatory states across millions of cells. How to integrate multimodal data is an open problem, existing integration methods struggle with accuracy and modality-specific biological variation retention. In this paper, we present scHyper (scalable, interpretable machine learning for single cell integration), a low-code and data-efficient deep transfer model designed for integrating paired and unpaired single-cell multimodal data. We benchmark scHyper against datasets from different multimodal data. ScHyper learns a low-dimensional representation and aligns the covariance matrices of the measured modalities, achieving high accuracy even with large scale atlas-level datasets with low memory and computational time across different cell lines, shedding light on regulatory relationships between different types of omics. Altogether, we show that scHyper is a versatile and robust tool for cell-type label transfer and integration from multimodal single-cell datasets.

Single-Cell Analysis

HINN: Hierarchical Input Neural Network identifies multi-omics biomarker for cognitive decline.

Understanding complex diseases requires models that can integrate diverse layers of biological data while yielding insights that are biologically interpretable. Although multi-omics integration with machine learning (ML) has advanced disease prediction and biomarker discovery, most existing approaches overlook the hierarchical and regulatory relationships that connect these molecular layers. Here, we present the Hierarchical Input Neural Network (HINN), a deep learning framework that incorporates known cross-omics relationships directly into its architecture, capturing the flow of information from genomics to epigenomics, transcriptomics, and downstream biological processes. By embedding these relationships, HINN improves both predictive performance and biological interpretability. We applied HINN to blood-derived multi-omics data from individuals with Alzheimer's disease or mild cognitive impairment to predict cognitive scores from standardized assessments. HINN outperformed both baseline and state-of-the-art models and pinpointed multi-omics biomarkers-including SNPs and promoter-region CpG sites in ATP6V1C1 and RCHY1 -that were significantly correlated with plasma p-Tau181 levels. These features map to biologically relevant processes with potential implications for cognitive decline. Our findings demonstrate how combining deep learning with biological knowledge can uncover interpretable, blood-based biomarkers for cognitive decline due to complex diseases such as Alzheimer's. All code and data are openly available at https://github.com/bozdaglab/HINN.

Alzheimer&#x2019;s disease

Towards mechanistic models of mutational effects: Deep learning on Alzheimer's A&#x3b2; peptide.

Deep Mutational Scanning (DMS) has enabled multiplexed measurement of mutational effects on protein properties, including kinematics and self-organization, with unprecedented resolution. However, potential bottlenecks of DMS characterization include experimental design, data quality, and depth of mutational coverage. Here, we apply deep learning to comprehensively model the mutational effect of the Alzheimer's Disease associated peptide A&#x3b2;42 on aggregation-related biochemical traits from DMS measurements. Among tested neural network architectures, Convolutional Neural Networks and Recurrent Neural Networks are found to be the most cost-effective models with high performance even under insufficiently-sampled DMS studies. While sequence features are essential for satisfactory prediction from neural networks, geometric-structural features further enhance the prediction performance. Notably, we demonstrate how mechanistic insights into phenotype may be extracted from the neural networks themselves suitably designed. This methodological benefit is particularly relevant for biochemical systems displaying a strong coupling between structure and phenotype such as the conformation of A&#x3b2;42 aggregate and nucleation, as shown here using a Graph Convolutional Neural Network (GCN) developed from the protein atomic structure input. In addition to accurate imputation of missing values (which here ranged up to 55% of all phenotype values at key residues), the mutationally-defined nucleation phenotype generated from a GCN shows improved resolution for identifying known disease-causing mutations relative to the original DMS phenotype. Our study suggests that neural network derived sequence-phenotype mapping can be exploited not only to provide direct support for protein engineering or genome editing but also to facilitate therapeutic design with the gained perspectives from biological modeling.

Alzheimer's disease

Multimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases.

Plant diseases destroy 20-40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.

convolutional neural networks

GE-IA-NAM: gene-environment interaction analysis via imaging-assisted neural additive model.

MOTIVATION: Gene-environment (G-E) interaction analysis is crucial in cancer research, offering insights into how genetic and environmental factors jointly influence cancer outcomes. Most existing G-E interaction methods are regression-based, which may lack flexibility to capture complex data patterns. Recent advances have investigated deep neural network-based G-E models. However, these methods may be more vulnerable to information deficiency due to challenges such as limited sample size and high dimensionality. Apart from genetic and environmental data, pathological images have emerged as a widely accessible and informative resource for cancer modeling, presenting its potential to enhance G-E modeling. RESULTS: We propose the pathological imaging-assisted neural additive model for G-E analysis (GE-IA-NAM). The flexible and interpretable additive network architecture is adopted to account for individualized effects associated with genetic factors, environmental factors, and their interactions. To improve G-E modeling, an assisted-learning strategy is investigated, which adopts a joint analysis to integrate information from pathological images. Simulations and the analysis of lung and skin cancer datasets from The Cancer Genome Atlas demonstrate the competitive performance of the proposed method. AVAILABILITY AND IMPLEMENTATION: Python code implementing the proposed method is available at https://github.com/Mr-maoge/NAM-IA-GE. The data that support the findings in this article are openly available in TCGA (The Cancer Genome Atlas) at https://portal.gdc.cancer.gov/.

Gene-Environment Interaction