Search PubMedSearch

SEARCH · Search PubMed

Results for “multimodal integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources.

MOTIVATION: Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. RESULTS: We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is "task agnostic", in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer's disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. AVAILABILITY AND IMPLEMENTATION: miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Humans

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans

Artificial Intelligence for Colorectal Surgeons-Part II: Research Applications, Challenges in Adoption, and Practical Resources.

BACKGROUND: This is part II of a 2-part series examining artificial intelligence in colorectal surgery. Part I established foundational concepts and clinical applications. Implementation, however, requires understanding research methodologies, available resources, and the specific challenges currently limiting widespread adoption. These topics are the focus of part II. OBJECTIVE: To examine artificial intelligence's transformation of surgical research, provide practical implementation resources, address adoption challenges, and explore future directions in colorectal surgery. METHODS: Comprehensive literature review focusing on artificial intelligence research methodology, implementation barriers, educational resources, and emerging technologies relevant to colorectal surgeons. RESULTS: Artificial intelligence streamlines clinical trial design through predictive modeling and natural language processing, reducing enrollment challenges that contribute to failed or inadequate trial accrual. Machine learning enables heterogeneity analysis within clinical trials, identifying treatment-responsive subgroups. Foundation models unlock analysis of unstructured electronic health record data at scale. Professional societies and universities offer specialized artificial intelligence education programs, with open-access data sets facilitating research participation. However, implementation faces multifaceted challenges: technical infrastructure demands, with real-time processing requiring dedicated graphics processing unit clusters; regulatory frameworks struggling with continuously evolving algorithms; undefined liability distribution for artificial intelligence-assisted decisions; algorithmic bias risking health care disparities; and the "black box" problem limiting clinical trust. Economic barriers include substantial initial costs without clear reimbursement pathways. Future directions include multimodal artificial intelligence integrating imaging, genomics, and histopathology; cognitive robotic systems with real-time decision support; digital twin technology for patient-specific surgical simulation; and global surgical artificial intelligence networks enabling distributed learning across institutions. CONCLUSIONS: Although artificial intelligence offers transformative potential for colorectal surgery research and practice, successful implementation requires addressing technical, regulatory, ethical, and economic challenges. The surgeon's evolving role demands both traditional expertise and computational fluency. Future advances in multimodal integration, autonomous systems, and global collaboration will fundamentally reshape surgical practice but will require thoughtful implementation prioritizing patient benefit and clinical value.

Humans

Longitudinal Clinical, Physiological, and Molecular Profiling of Female Patients With Metastatic Cancer: Protocol and Feasibility of a Multicenter High-Definition Oncology Study.

PURPOSE: A substantial proportion of patients receiving genomically matched therapies do not achieve clinical benefit, underscoring the influence of nongenetic factors on cancer outcomes. High-Definition Oncology (HDO) proposes integrating longitudinal, multimodal patient data-spanning clinical, molecular, physiological, and behavioral domains-to enable truly individualized cancer care. This manuscript describes the HDO study design, framework, and feasibility results in women with metastatic cancer. METHODS: We initiated a prospective, multicenter observational study (HDO study; ClinicalTrials.gov identifier: NCT06590506) enrolling 300 female patients with newly diagnosed metastatic breast, lung, or colorectal cancer. Here, we report the study design, standardized workflows, prespecified feasibility criteria, and early internal pilot results. Eleven data modalities are collected longitudinally, including tumor and germline genomics, germline epigenomics, gut microbiome, blood and stool metabolomics and proteomics, exposome characterization, wearable-derived physiological monitoring, digital footprint assessment, medical imaging, and patient-reported outcomes. Standardized workflows govern clinical procedures, data acquisition, biospecimen processing, and quality control across all participating sites. RESULTS: Feasibility was evaluated in the first 30 participants (10% of planned accrual). Patients completed 100% of scheduled clinical visits, 97.4% of planned plasma collections, 80.7% of stool samples, and all tumor biopsies. Wearable devices captured activity, heart rate, sleep, and blood oxygen saturation data during 95.0%, 84.2%, 90.6%, and 70.7% of total patient-days, respectively. Biospecimens met predefined quality control metrics across all molecular modalities. Engagement with mobile applications for pain and emotion reporting exceeded 80%. CONCLUSION: The HDO study demonstrates the feasibility of comprehensive, longitudinal, multimodal data collection in women with metastatic cancer. This internal pilot establishes an integrated framework for future analyses aimed at characterizing disease trajectories, defining molecular and physiological determinants of outcomes, and developing patient-specific computational models.

Humans

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis

Diagnosing the undiagnosed: AI-enhanced multimodal modeling for placental mesenchymal dysplasia in high-risk pregnancies.

Placental mesenchymal dysplasia (PMD) is a rare vascular placental disorder that mimics molar pregnancy but often coexists with a viable fetus, making its misdiagnosis potentially devastating. In high-risk pregnancies, artificial intelligence (AI)-enhanced multimodal modeling - incorporating imaging, genomics, proteomics, and clinical features - offers a transformative diagnostic strategy. Leveraging Bayesian hyperparameter optimization for model refinement, this approach improves diagnostic accuracy while reducing uncertainty and clinician hesitation. Recent clinical studies support its efficacy and interpretability through SHAP and LIME models, while real-time surgical enhancements using Bayesian methods highlight its broader clinical utility. Despite current challenges such as data heterogeneity and integration barriers, multimodal AI provides unprecedented resolution in placental analysis, enabling precise differentiation between PMD and similar fetopathies. Ultimately, this advancement supports timely, non-invasive diagnosis, personalized management, and emotionally informed decision-making aligned with ethical AI implementation standards.

Bayesian optimization

CAUSAL artificial intelligence and data-driven decision intelligence in personalized medicine: a review of healthcare informatics systems.

This review examines the integration of causal artificial intelligence (AI) and data-driven decision intelligence within healthcare informatics systems to advance personalized medicine and clinical decision-making. A narrative review methodology was employed, synthesizing interdisciplinary literature from major databases, including PubMed, Scopus, Web of Science, IEEE Xplore, and ScienceDirect. Studies focusing on causal inference, decision intelligence, and healthcare informatics applications in personalized medicine were included. Data were extracted on methodological approaches, healthcare settings, analytical techniques, and clinical applications, followed by thematic synthesis. Findings indicate that causal AI enhances clinical decision support by enabling estimation of treatment effects and simulation of intervention outcomes at the individual patient level. Integration of multimodal health data such as electronic health records, genomic data, and real-time monitoring improves prediction accuracy and supports tailored treatment strategies. Additionally, causal models improve interpretability, fostering clinician trust and facilitating transparent decision-making. Robust healthcare informatics infrastructures, including interoperable systems and data warehouses, were identified as critical enablers of causal analytics. Overall, causal AI represents a transformative advancement in healthcare analytics, supporting more informed, individualized, and evidence-based clinical decisions. Its integration within healthcare informatics systems has significant potential to improve patient outcomes and guide the future of intelligent, personalized healthcare delivery.

Precision Medicine

JASMINE: A powerful representation learning method for enhanced analysis of incomplete multi-omics data.

Integrative analysis of multi-omics data provides a more comprehensive and nuanced view of a subject's biological state. However, high-dimensionality and ubiquitous modality missingness present significant analytical challenges. Existing methods for incomplete multi-omics data are scarce, do not fully leverage both modality-specific and shared information, and produce task-biased representations. We propose JASMINE, a self-supervised representation learning method for incomplete multi-omics data that preserves both modality-specific and joint information and enhances sample similarity structure. JASMINE produces embeddings that achieve superior performance across multiple tasks for two different incomplete multi-omics datasets while requiring only a single round of training per dataset.

missing data

Osteoarthritis phenotypes: advancing precision medicine through clinical, structural, and molecular stratification.

PURPOSE: Osteoarthritis (OA) is now understood as a heterogeneous syndrome driven by diverse biological, biomechanical, metabolic, genetic, and molecular mechanisms. This variability explains differences in disease progression and treatment response, challenging the traditional "one-size-fits-all" approach. This review highlights OA phenotyping as a key step toward precision medicine, focusing on clinical, structural, and molecular classifications that inform individualized care. METHODS: A narrative review was conducted using a non-systematic search of major databases and Osteoarthritis Research Society International sources (2010-2026). Evidence was thematically synthesized across clinical, imaging, and molecular domains to characterize OA phenotypes and their potential relevance to precision medicine. RESULTS: Multiple OA phenotypes were identified: inflammatory, metabolic, biomechanical, cartilage-subchondral, pain-sensitization, and aging/senescence. These exhibit distinct clinical features, risk factors, and therapeutic responses. Imaging-based phenotypes (e.g., inflammatory, meniscus-cartilage, subchondral bone, atrophic, hypertrophic) and molecular endotypes (low turnover, structural damage, systemic inflammation) further refine stratification. Pain-structure discordance is notable in sensitization phenotypes and may predict poorer surgical outcomes. Joint-specific variations and emerging genomic and epigenetic insights underscore disease complexity. Advances in imaging, biomarkers, and machine learning may enable earlier detection and patient clustering, though clinical application remains limited. CONCLUSION: Phenotype- and endotype-based classification represents a critical advancement toward precision OA management. Tailored interventions based on stratification hold promise for improving outcomes; however, clinical translation remains limited by overlapping phenotypes, lack of validated biomarkers, and inconsistent results from phenotype-driven trials. Wider clinical adoption requires standardized definitions, validation across joints, and integration of multimodal diagnostic tools into routine practice.

Humans

Artificial intelligence in kidney cancer: a review of clinical applications across the disease spectrum.

PURPOSE OF REVIEW: This review examines recent advances (2024-2025) in the application of artificial intelligence (AI) to kidney cancer diagnosis, prognosis, and treatment planning. It categorizes studies across 13 clinical scenarios to assess where AI offers the most clinical utility. RECENT FINDINGS: AI models have demonstrated strong performance in a range of tasks including tumor grading, subtype classification, survival prediction, and risk stratification. Integration of radiomics, genomics, and histopathology has enabled personalized, noninvasive, and timely decision-making. The highest-performing models used CT-based radiomics, particularly for predicting progression-free and recurrence-free survival. However, performance varies across tasks and tumor subtypes, with lower accuracy in detecting oncocytomas or benign vs. malignant differentiation. AI applications in metastatic and nonresected cases remain underexplored, and ultrasound remains a largely under researched modality. While some models improve diagnostic accuracy and workflow efficiency, broader validation across diverse populations is still needed. SUMMARY: AI is transforming kidney cancer care across multiple clinical stages. Although promising, real-world implementation demands ongoing validation and postdeployment monitoring to prevent performance degradation due to distributional drift. AI's integration with multimodal data offers substantial potential to improve outcomes and reduce overtreatment.

Humans

AI-driven diagnostic and prognostic models for metabolic dysfunction-associated steatotic liver disease: insights from clinical, imaging, and multi-omics studies-a scoping review.

Metabolic dysfunction-associated steatotic liver disease (MASLD), formerly known as non-alcoholic fatty liver disease (NAFLD), is the most common chronic liver disease around the world, affecting 33.6% of the adult population (95% CI: 28.1%-39.5%; I 2&#x2009;=&#x2009;99.9%), or roughly one in three. The extent of the liver damage is variable, from simple steatosis to metabolic dysfunction-associated steatohepatitis (MASH, formerly NASH), cirrhosis and hepatocellular carcinoma (HCC). Early diagnosis is essential to prevent serious liver damage. Traditional diagnostic techniques such as liver biopsy, imaging, and biomarker testing are all invasive, costly, reduced sensitive to early-stage disease, and they also have variability among observers. Modern diagnostic and prognostic approaches based on the principles of Artificial Intelligence (AI) and specifically on machine learning (ML) and deep learning (DL) have enabled multimodal approaches integrating clinical, imaging and molecular data. This scoping review conducted per PRISMA-ScR guidelines, synthesizes findings from 73 studies (search window 2020-2026) across three dimensions: clinical data driven models, imaging-based classifiers (ultrasound, CT and MRI), and multi-omics (genomics, transcriptomics and proteomics) techniques. Moreover, emergence of models such as U-Net and LiverNet 2.x, classification models like DeepLiverNet and BiLSTM models, as well as transformer frameworks and the identification of biomarkers models are also described. This study also investigates challenges such as data heterogeneity, data interpretability, fairness and real-world clinical application. Finally, important areas of research opportunities and future directions are highlighted to present a developing clinically applicable, explainable and ethical AI solutions to manage MASLD.

MASLD

Biological Foundation Models for Complex Disease Research and Clinical Translation.

Complex diseases, including cancer, rare genetic disorders, neurodevelopmental and psychiatric conditions, and neurodegenerative diseases, arise from interactions among genetic variation, gene regulation, and cellular states that are difficult to capture using a single data type or biological scale. Biological foundation models address this challenge by treating nucleotides and genes as tokens and learning representations that can be transferred to downstream biomedical and clinical tasks. In this review, we examine two major model classes, genomic sequence foundation models and cell foundation models, and compare their tokenization strategies, model architectures, pretraining objectives, and adaptation methods. We summarize their emerging applications in regulatory variant interpretation, disease-associated cell-state analysis, drug-response prediction, and therapeutic target discovery across complex diseases. We distinguish applications supported by experimental or retrospective validation from those that remain primarily computational or conceptual. We further discuss key challenges to clinical translation, including multimodal data integration, model interpretability, benchmarking, patient-specific prediction, and privacy protection. We highlight future opportunities to integrate biological foundation models with emerging frameworks of medical digital twins, agentic AI, and federated learning. By linking model design to translational goals, this review provides a practical framework for evaluating biological foundation models and their readiness for complex disease research and clinical use.

biological foundation model

Integrating metagenomic next-generation sequencing into a multimodal diagnostic framework for spinal infection: enhancing etiological identification and clinical prediction.

BACKGROUND: Spinal infection (SI) remains diagnostically challenging because of heterogeneous etiologies, nonspecific clinical manifestations, and the limited sensitivity of conventional microbiological approaches, particularly following empirical antimicrobial exposure. Although metagenomic next-generation sequencing (mNGS) enables unbiased pathogen detection, its incremental clinical value beyond pathogen identification and its role within integrated diagnostic strategies remain incompletely established. METHODS: We retrospectively analyzed 208 consecutive patients with suspected SI between August 2022 and August 2025. Final diagnoses were established using a multidisciplinary-adjudicated composite reference standard incorporating clinical, radiological, microbiological, and histopathological evidence. The diagnostic performance of mNGS was compared with conventional culture and histopathology. Furthermore, multimodal predictive models integrating clinical variables and microbiological information were developed using L1-regularized logistic regression. RESULTS: In the comparative cohort, mNGS achieved a significantly higher diagnostic yield than culture (66.5% vs. 27.41%, P < 0.001). Among confirmed SI cases, mNGS demonstrated higher sensitivity than conventional culture (91.67% vs. 40.15%, P < 0.001). mNGS identified a substantially broader pathogen spectrum, ranging from fastidious organisms such as Mycobacterium tuberculosis and Brucella to rare pathogens including Talaromyces marneffei and Coxiella burnetii, and maintained robust sensitivity (98.2%) despite prior antibiotic exposure. While an integrated clinical model achieved an AUC of 0.916, mNGS as a standalone modality provided superior discriminative power (AUC = 0.889) compared to histopathology (AUC = 0.836), the Conventional Biomarker Model (AUC = 0.742), and culture (AUC = 0.693). CONCLUSIONS: mNGS is a high-yield diagnostic tool for spinal infection, particularly in culture-negative and antibiotic-pretreated scenarios. Integrating mNGS into a multimodal clinical framework facilitates etiological clarity and precision antimicrobial therapy.

Humans

Scalable, generalizable and uncertainty-aware integration of spatial multiomics across diverse modalities and platforms with SCIGMA.

Recent advances in spatial omics technologies have enabled simultaneous profiling of transcriptomic, proteomic, epigenomic, metabolomic and imaging data at high spatial resolution, offering unprecedented opportunities to dissect tissue complexity. However, integrating these diverse and large-scale spatial multimodal datasets remains a major computational challenge. We present SCIGMA, a scalable and generalizable deep learning framework for spatial multiomics integration. SCIGMA introduces an uncertainty-aware contrastive learning objective and multiview graph neural networks to preserve modality-specific signals while learning biologically meaningful joint representations. Unlike previous methods, SCIGMA provides spatially resolved uncertainty estimates, interpretably identifying regions of biological or technical heterogeneity. SCIGMA supports integration of up to five modalities, and its modular framework is extensible to future technologies with even more modalities. It also scales to more than 1 million spatial locations, enabling analysis of high-resolution datasets such as Visium HD and Xenium Prime. We evaluated SCIGMA across 19 datasets spanning 8 modalities, 10 tissues and 9 platforms. On benchmarkable datasets, SCIGMA outperformed other methods in spatial domain detection, modality preservation, feature reconstruction and reproducibility. SCIGMA identifies biologically meaningful structures, refined spatial domains and modality-specific regulatory programs, providing a robust, flexible and future-ready solution for scalable spatial multimodal integration.

Multiomics

Multimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases.

Plant diseases destroy 20-40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.

convolutional neural networks