Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiomics integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Integrative multi-omics analysis unravels the metabolic landscape and reveals serum biomarkers for early diagnosis of hyperuricemia.

BACKGROUND: Hyperuricemia (HUA) is a major risk factor for gout and multiple metabolic disorders. Although serum uric acid (UA) is the gold standard for HUA diagnosis, it fails to reflect early metabolic disturbances and shows limited predictive value for asymptomatic HUA. This study sought to elucidate the pathological mechanisms underlying HUA and identify novel diagnostic biomarkers beyond UA. METHODS: This study enrolled 195 patients with HUA and 98 healthy controls. Global metabolomics and proteomics profiling were performed to characterize molecular alterations underlying HUA. Based on the biological relevance of the shared dysregulated pathways, a pathway correlation network was constructed to elucidate the pathological mechanisms driving HUA initiation and progression. Furthermore, diagnostic biomarkers for HUA were identified using machine learning algorithms, and were validated with an external cohort. RESULTS: HUA patients exhibited distinct metabolic and proteomic profiles compared with healthy controls. Integrated multi-omics pathway analysis revealed that peroxisome proliferators-activated receptor signaling pathway, arachidonic acid metabolism, purine metabolism, pyrimidine metabolism and sphingolipid signaling pathway were significantly dysregulated in HUA. Among them, arachidonic acid metabolism was identified as a hub pathway involved in HUA progression. Furthermore, a metabolite panel consisting of cysteine-S-sulfate, glycerophosphocholine and 4-hydroxyphenylpyruvic acid was screened by machine learning and validated in an independent cohort, which showed slightly higher diagnostic performance for HUA than UA. CONCLUSIONS: This study reveals the core metabolic and protein regulatory networks of HUA, and identifies a novel serum metabolite panel for the diagnosis of HUA. These findings provide new insights for improved clinical diagnosis and management.

Humans↗

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis↗

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software↗

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans↗

Multi-omics dynamic profiling reveals predictive biomarkers for first-line immunochemotherapy in extensive-stage small-cell lung cancer.

BACKGROUND: Extensive-stage small-cell lung cancer (ES-SCLC) is associated with a poor prognosis. Although first-line immunochemotherapy improves clinical outcomes, robust prognostic biomarkers for this treatment modality remain unavailable. The aim of this study was to identify non-invasive, easily accessible, and dynamically monitored biomarkers of ES-SCLC by machine learning integrating serum metabolomics, lipidomics, and proteomics at multiple time points. METHODS: A total of 816 serum samples were collected from ES-SCLC patients receiving first-line immunotherapy combined with chemotherapy or first-line chemotherapy for metabolomics, lipidomics, and proteomics analysis. The immunochemotherapy cohort was randomly divided into training and validation subsets at a 6:4 ratio. Biomarkers were identified using machine learning algorithms, and their prognostic significance was evaluated through receiver operating characteristic (ROC) analysis, Kaplan–Meier survival analysis, and multivariate Cox regression. Potential metabolic pathways and mechanisms were further explored via integrated multi-omic analysis. RESULTS: The immunochemotherapy exhibited a prolonged median progression-free survival (PFS) and higher objective response rate (ORR) compared to the chemotherapy group. A total of 5 serum metabolites (uric acid, L-aspartate-semialdehyde, dimethisterone, xanthine, L-cysteine), 6 lipids (Cer d18:1/26:0, Cer d18:2/25:0, SM d18:1/20:1, SM d17:1/25:1, DG O-18:1_16:0, PS 18:0_24:0), and 3 proteins (ACIN1, ACSL4, PHGDH) were identified and constructed into independent prognostic models. Among patients receiving immunochemotherapy, those categorized as low-risk based on the model demonstrated significantly longer PFS compared with those in the high-risk group. These prognostic signatures also retained predictive value in patients who underwent second-line treatment with anlotinib plus immunochemotherapy. Integrated analysis revealed that glycine, serine, and threonine metabolism was the commonly enriched pathway across all three omics layers. Notably, PHGDH (protein), L-aspartate-semialdehyde and L-cysteine (metabolites), and PS (18:0_24:0) (lipid), key elements in this pathway, were all incorporated in the predictive model. In addition, models of the composition of these substances after one cycle of treatment can still predict the prognosis of patients. CONCLUSION: In this study, we constructed and validated a set of non-invasive, dynamically monitorable prognostic models (containing 5 metabolites, 6 lipids, and 3 proteins) using machine learning by integrating multiple time point data from the serum metabolome, lipid panel, and proteome to accurately distinguish the prognostic risk of patients with ES-SCLC receiving immunochemotherapy. PFS was significantly prolonged in patients in the low-risk group, and this model remains predictive in the subsequent second-line treatment with anlotinib in combination with immunochemotherapy. Glycine-serine-threonine metabolic pathway may be the key mechanism, of which PHGDH, L-aspartate semialdehyde, L-cysteine and PS (18:0_24:0) are the core predictors. This study provides the first multi-omics dynamic prognostic tool for ES-SCLC immunochemotherapy and reveals potential therapeutic targets.

Humans↗

Integrated transcriptomic and metabolomic analysis reveals candidate regulatory networks associated with starch accumulation in tetraploid potato.

Potato (Solanum tuberosum L.) tuber starch is a major determinant of crop quality and industrial value, yet the regulatory mechanisms underlying starch accumulation in autotetraploid cultivars remain poorly resolved. Here, we performed integrated transcriptomic and metabolomic analyses using a segregating tetraploid population derived from parents with contrasting starch content. Extreme phenotypes were selected to systematically dissect the molecular basis of starch accumulation. Transcriptome profiling revealed extensive transcriptional reprogramming between high- and low-starch genotypes, with differentially expressed genes significantly enriched in carbohydrate metabolism, particularly the starch and sucrose metabolism pathway. Notably, multiple transcription factor families, including AP2/ERF, MYB, and bHLH, were prominently represented, suggesting coordinated regulatory control. Metabolomic analysis identified substantial metabolic divergence, with differentially accumulated metabolites predominantly enriched in starch and sucrose metabolism as well as secondary metabolic pathways. Most metabolites exhibited negative associations with starch content, indicating competitive carbon allocation between primary and secondary metabolism. Integrative multi-omics analysis further resolved a core regulatory module comprising key structural genes and transcription factors tightly associated with starch-related metabolites. In particular, genes involved in sucrose cleavage and ADP-glucose metabolism, together with trehalose-6-phosphate synthase (TPS) and UDP-glucose-associated pathways, emerged as critical nodes linking carbon flux to starch biosynthesis. Correlation network analysis suggested that AP2/ERF-, MYB-, and bHLH-type transcription factors modulate these pathways by coordinating structural gene expression and metabolic flux distribution. Collectively, our study establishes a transcriptional-metabolic framework for starch accumulation in tetraploid potato, highlighting the central role of carbon allocation and signaling intermediates in shaping starch content, and providing candidate targets for molecular breeding and genome editing.

Solanum tuberosum↗

Multi-omics integrative analysis provides insight into potential molecular responses to sustained high water flow in common carp (Cyprinus carpio) cultured in recirculating aquaculture.

To investigate the potential molecular responses by which water flow intensity affects the growth of common carp (Cyprinus carpio) in a recirculating aquaculture system (RAS), a control group (CG, actual water velocity 0.3&#xa0;cm/s) and three sustained flow treatment groups were established, including a low-flow group (LF, 1 body length per second, bl/s), a medium-flow group (MF, 2 bl/s), and a high-flow group (HF, 3 bl/s). After 12&#xa0;weeks of culture in the RAS, growth performance was compared among groups under different flow intensities. The best-performing group and the control group were then selected for the determination of intestinal digestive enzyme activities, as well as transcriptomic and whole-genome bisulfite sequencing analyses of muscle tissue. The results showed that the specific growth rate and feed intake of the HF group were significantly higher than those of the other groups (P&#xa0;<&#xa0;0.05), whereas no significant difference in feed conversion ratio was observed among groups. Compared with the CG group, lipase activity was significantly higher in the HF group (P&#xa0;<&#xa0;0.05), while &#x3b1;-amylase and trypsin activities showed increasing trends without significant differences. RNA-seq identified a total of 273 differentially expressed genes, including 72 upregulated genes and 201 downregulated genes in the HF group relative to the CG group. These genes were mainly enriched in glycolysis, pyruvate metabolism, ATP metabolism, the pentose phosphate pathway, the insulin signaling pathway, the PPAR signaling pathway, and the adipocytokine signaling pathway, indicating that sustained high water flow induced a muscle transcriptional response characterized by remodeling of energy metabolism and substrate utilization. Whole-genome bisulfite sequencing analysis showed that DNA methylation in common carp muscle occurred predominantly in the CpG context. Differentially methylated regions between the HF and CG groups were mainly distributed in transcription-related regulatory regions, including promoters, CpG islands, and CpG island shores. In promoter regions, the number of hypermethylated regions in the HF group relative to the CG group was markedly higher than that of hypomethylated regions. Integrated analysis further identified two candidate genes showing both promoter differential methylation and differential expression, namely LOC109094644 and bcorl1, suggesting that adaptation to high water flow may involve IGF-related growth regulation and remodeling of upstream transcriptional programs. The qPCR results were consistent with the transcriptomic data. Taken together, within the tested range, a sustained water flow of 3 bl/s was more conducive to the growth of common carp in the RAS, which may be associated with enhanced lipid digestion and utilization, remodeling of the muscle energy metabolic network, changes in promoter methylation, and the coordinated regulation of key candidate genes. This study provides a theoretical basis for clarifying the exercise adaptation mechanism of common carp in recirculating aquaculture and for optimizing flow velocity parameters.

Animals↗

From molecular responses to environmental monitoring: advances and translational gaps in omics approaches in fish environmental toxicology.

Fish occupy a central position in aquatic ecosystems and serve as important bioindicators for environmental monitoring, as well as powerful translational models for understanding toxic mechanisms conserved across higher vertebrates. In recent years, omics techniques have proven to be powerful tools to address complex environmental questions that conventional toxicology methods cannot answer. Despite this potential, a critical translational gap remains between molecular findings and their use in ecological risk assessment frameworks. This review critically synthesizes advances across omics techniques including epigenomics, transcriptomics, metabolomics and proteomics and their integration. Special emphasis is placed on methodological considerations and practical aspects of these techniques in fish environmental toxicology and environmental monitoring. Evidence from single-omics studies suggests conserved biomarker signatures across species while characterizing complex phenomena like non-monotonic dose-response relationships, mixture toxicity and transgenerational and stereoselective effects with implications for population level monitoring. Multi-omics studies, especially those involving triple omics, further enhance mechanistic resolution by reconstructing adverse outcome pathways. We further evaluate using case studies when additional molecular layers provide critical insight and when they offer limited advantage, a strategic distinction with direct implications in environmental monitoring programmes. Finally, current limitations and future directions that will ultimately bridge the translational gap and hold promise for advancing mechanistic ecotoxicology and predictive environmental monitoring are discussed.

Animals↗

Decoding the spatiotemporal patterns of food spoilage microbial communities: Integrating multi-omics and artificial intelligence to enable precision preservation.

In the global food supply chain, food wastage caused by spoilage has resulted in significant economic losses, food shortages, and environmental pressure. This process is fundamentally driven by the spatiotemporal dynamics of microbial communities. However, traditional research methods struggle to elucidate the complex mechanisms of spatial heterogeneity, interspecies interactions, and functional succession. This limits the development of effective preservation strategies. This review systematically reviews the cutting-edge progress of integrating multi-omics technologies and artificial intelligence (AI) to study food spoilage microbial communities, breaking through this bottleneck. We propose an intelligent theoretical framework that could potentially analyze microbial metabolic activities and predict dynamic shelf life if implemented. The conceptual framework integrates multidimensional data, including spatial metabolomics, temporal metatranscriptomics, single-cell transcriptomics, and longitudinal metagenomics. It can also be combined with AI models, such as graph neural networks. The article elaborates on the principles and applications of spatio-temporal monitoring technologies, such as nano secondary ion mass spectrometry, hyperspectral imaging, and the Internet of Things sensing. Through illustrative cases of typical perishable foods, it also explores how such a multi-omics - AI system might be applied to spoilage warning and precise intervention. Additionally, the article addresses the current challenges in data coverage, model generalization, and federated learning implementation. Then the research further explores emerging areas such as engineered probiotics, edge AI, and microfluidic sensing. These areas are targeted at transforming food preservation from an empirical control approach to a data-driven, precise regulatory framework. This transformation provides theoretical support and technical approaches for developing a smart, sustainable food preservation system.

Multiomics↗

scMGCL: accurate and efficient integration representation of single-cell multi-omics data.

MOTIVATION: Single-cell multi-omics data integration is essential for understanding cellular states and disease mechanisms, yet integrating heterogeneous data modalities remains a challenge. We present scMGCL, a graph contrastive learning framework for robust integration of single-cell ATAC-seq and RNA-seq data. Our approach leverages self-supervised learning on cell-cell similarity graphs, in which each modality's graph structure serves as an augmentation for the other. This cross-modality contrastive paradigm enables the learning of biologically meaningful, shared representations while preserving modality-specific features. RESULTS: Benchmarking against state-of-the-art methods demonstrates that scMGCL outperforms others in cell-type clustering, label transfer accuracy, and preservation of marker-gene correlations. Additionally, scMGCL significantly improves computational efficiency, reducing runtime and memory usage. The method's effectiveness is further validated through extensive analyses of cell-type similarity and functional consistency, providing a powerful tool for multi-omics data exploration. AVAILABILITY AND IMPLEMENTATION: Code and datasets are released at https://github.com/zlCreator/scMGCL.

Single-Cell Analysis↗

CancerOmicsStudio (CoS): a web server for integrative and interpretable analysis of multi-omics cancer data.

MOTIVATION: Large-scale omics resources, including The Cancer Genome Atlas, Genomics of Drug Sensitivity in Cancer, and the Cancer Dependency Map, have become essential for cancer research. However, these datasets are distributed across different platforms, formats and analysis frameworks, which limits their practical use by researchers without extensive computational expertise. RESULTS: We developed CancerOmicsStudio (CoS), a web server for integrative and interpretable analysis of multi-omics cancer data across 33 cancer types. CoS provides five major modules: CosAI, Traditional Analysis, Drug Sensitivity, CRISPR Dependency and Single-Cell Tumor Microenvironment. The Traditional Analysis module supports expression comparison, diagnostic evaluation, survival analysis, enrichment analysis and gene correlation. The Drug Sensitivity and CRISPR Dependency modules enable systematic evaluation of gene-drug response associations and gene essentiality in cancer cell lines. The Single-Cell Tumor Microenvironment module supports tumor microenvironment analysis at single-cell resolution. In total, approximately 1.23 million results have been precomputed to enable rapid retrieval. CosAI further allows users to submit natural-language queries and obtain results through a Real-time Analysis as Retrieval framework, with responses summarized by a lightweight language model. AVAILABILITY AND IMPLEMENTATION: CancerOmicsStudio is freely available at Zenodo (doi: 10.5281/zenodo.18744990) and https://cos.wanglab.bio.

Humans↗

Integrating multi-omics technologies to decipher microbiome functions.

Multi-omics approaches have revolutionized our understanding of microbial communities by enabling simultaneous interrogation of genomic, transcriptomic, proteomic, and metabolomic data. The systematic integration and analysis of these deep datasets help decipher the functional roles of microbiomes, providing critical insights into microbial activities, interactions, and dynamics across diverse environments. Biological complexity makes multi-omics analysis of a single, isolated organism demanding but highly informative, yet this complexity increases further when samples comprise hundreds to thousands of individual species. As microbiome research continues to expand into clinical, environmental, and engineered systems, standardized workflows, benchmarked datasets, and community-driven initiatives are essential to ensure reproducibility, standardization and interpretability. Establishing and disseminating best practices for experimental design, data processing, and integrative analyses will be critical for maximizing comparability and scientific rigor across studies. This perspective highlights recent advances in multi-omics microbiome research, outlines key obstacles in data integration and metadata harmonization, and proposes a collaborative roadmap for scalable, FAIR-compliant multi-omics investigations and potentially disruptive Artificial Intelligence (AI) advances comparable to those of AlphaFold in the field of microbiome science.

Multiomics↗

Multi-omics analyses provide insights into the molecular basis for salt tolerance of Phyla nodiflora.

The perennial herbaceous plant, Phyla nodiflora (Verbenaceae), which possesses natural resistance to multiple abiotic stresses, is widely used as a pioneer species in island ecological restoration. Due to the lack of information about its genome, the mechanism underlying its tolerance to environmental stresses, such as salinity, is almost entirely unknown. Here, we report on the high-quality genome of P. nodiflora that is 403.07&#x2009;Mb in size, and which was assembled and anchored onto 18 pseudo-chromosomes. Genomic synteny revealed that P. nodiflora underwent two whole genome duplication events, which promoted the expansion of genes related to environmental adaptation and the biosynthesis of secondary metabolites. An integrated genomic and transcriptomic analysis suggested that salt stress tolerance in P. nodiflora is associated with the expansion and activated expression of genes related to abscisic acid (ABA) homeostasis and signaling. The expansion of ZEP family genes may contribute to the consistent increase in ABA levels under salt stress. Lysine acetylomic analysis revealed that exposure to salt led to widespread protein deacetylation, with these proteins primarily involved in signal transduction, carbohydrate transport and metabolism, and transcription regulation. Deacetylation of glutathione S-transferase increased enzymatic activities in response to salt-induced oxidative stress. Collectively, the genomic, transcriptomic, and lysine acetylomic analyses provide profound insight into the molecular basis of the adaptation of P. nodiflora to salt stress, and will be helpful to engineer salt-tolerant plants for ecological restoration.

Salt Tolerance↗

Multi-Omics insights into OsZFP252-OsGA20ox5 mediated drought tolerance in rice through stomatal and vascular regulation.

Rice growth is highly dependent on water availability, and drought stress significantly impacts its entire life cycle. However, previous studies lack systematic investigations into drought-responsive candidate genes across the full life cycle of rice. This study integrates transcriptomic and phenotypic data from two rice lines, IR64 (drought-sensitive) and DK151 (drought-tolerant), under varied environmental conditions at distinct growth stages. Using k-means clustering, 13&#x2009;369 genes were categorized into 17 distinct expression patterns, revealing drought-responsive genes specifically upregulated or downregulated under drought stress. Weighted co-expression network analysis (WGCNA) further identified four gene modules strongly correlated with drought-related phenotypes, co-localizing 2859 drought-responsive genes through both approaches. Proteomics and metabolomics were supplemented at the booting stage, where phenotypic and transcriptomic differences under drought were most pronounced. Integrated omics results demonstrate gibberellin (GA) and abscisic acid (ABA) pathways play a key role during drought tolerance in rice, and 79 high-confidence drought-resistant candidate genes were prioritized from the 2859 drought-responsive genes. Among these, Gibberellin 20-oxidase 5 (OsGA20ox5) was identified as a key negative regulator of drought tolerance. Furthermore, the transcription factor zinc finger protein 252 (OsZFP252) directly binds to the OsGA20ox5 promoter, repressing its expression and enhancing ABA biosynthesis, thereby improving drought tolerance by increasing stomatal closure and expanding vascular bundle water transport capacity. Notably, the drought-tolerant haplotype 2-4 (Hap2-4) of OsGA20ox5 provides valuable insights for drought-resistant breeding.

Oryza↗

Next-Generation Disease Profiling by Integrating Histopathology with Spatial Multi-Omics Data.

The field of pathology has experienced several transformative changes in recent years with the advent of digital pathology and spatial multi-omics. These technologies have enhanced every aspect of pathology practice, from streamlining daily workflows to generating high-fidelity multi-omics data that provide pathologists with novel tools to refine disease profiling and clinical diagnosis. Each layer of multimodal data (genomic, metabolomic, proteomic, or transcriptomic) has uncovered a distinct facet of disease pathologies, and combined with machine learning/artificial intelligence-based data analysis and pattern recognition models, has provided holistic understanding of regulatory mechanisms underpinning them. However, high-dimensional data have far exceeded the volume, scale, and complexity of immunostaining methods implemented by pathologists and, thus, have generated significant challenges related to deconvolution, interpretation, and clinical translation. Furthermore, these multimodal studies have predominantly relied on computational methods to process data and extract disease-relevant insights, thus raising questions around relevance or role of a pathologist in this new era of multi-omics. This review will provide a perspective on the evolving fields of molecular histopathology and spatial -omics, leveraging them to approach disease profiling, and redefining the role of a pathologist during this process.

Humans↗

A generalized higher-order correlation analysis framework for multi-omics network inference.

Multiple -omics (genomics, proteomics, etc.) profiles are commonly generated to gain insight into a disease or physiological system. Constructing multi-omics networks with respect to the trait(s) of interest provides an opportunity to understand relationships between molecular features but integration is challenging due to multiple data sets with high dimensionality. One approach is to use canonical correlation to integrate one or two omics types and a single trait of interest. However, these types of methods may be limited due to (1) not accounting for higher-order correlations existing among features, (2) computational inefficiency when extending to more than two omics data when using a penalty term-based sparsity method, and (3) lack of flexibility for focusing on specific correlations (e.g., omics-to-phenotype correlation versus omics-to-omics correlations). In this work, we have developed a novel multi-omics network analysis pipeline called Sparse Generalized Tensor Canonical Correlation Analysis Network Inference (SGTCCA-Net) that can effectively overcome these limitations. We also introduce an implementation to improve the summarization of networks for downstream analyses. Simulation and real-data experiments demonstrate the effectiveness of our novel method for inferring omics networks and features of interest.

Genomics↗

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n&#x202f;=&#x202f;26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette &#x2248; 0.16) that remained unassociated with overall survival (log-rank p&#x202f;=&#x202f;0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) =&#x202f;0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p&#x202f;=&#x202f;0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans↗

Integrative multi-omics analyses suggest a candidate microbial metabolite-associated host gene network in ulcerative colitis.

Ulcerative colitis (UC) is associated with gut microbial dysbiosis, but the host molecular alterations potentially linked to microbially derived metabolites remain incompletely understood. We integrated Mendelian randomization (MR), microbial metabolite annotation, computational target prediction, colonic transcriptomics, network analysis, and machine learning. MiBioGen microbiome GWAS data were used as exposures and FinnGen Release 12 ULCERENTER as the outcome. Metabolites linked to MR-prioritized taxa were retrieved from GutMGene, and human targets were predicted using SwissTargetPrediction and SEA. UC-related genes were defined by integrating differential expression analysis and WGCNA and then intersected with predicted metabolite targets. MR prioritized one family and eight genera showing nominal genetically supported associations with UC, but none remained significant after Benjamini-Hochberg FDR correction. Three prioritized genera were linked to 15 microbe-metabolite records, corresponding to 13 unique metabolites; nine were retained for target prediction, yielding 277 unique predicted human targets. Transcriptomic analysis identified 1,530 DEGs and a 312-gene MEgrey60 module, with 273 overlapping genes, producing 1,569 unique UC-related genes. Their intersection with the 277 predicted targets yielded 47 candidate genes. Enrichment analyses highlighted mainly metabolic and lipid-related processes. Random Forest showed the highest mean AUC across the two independent external benchmarking cohorts, and SHAP prioritized EPHX1, HSD17B2, IGFBP5, and MMP10. IBDome analysis showed inflammation-associated expression differences in these genes. This study provides a genomics-informed, hypothesis-generating framework that prioritizes candidate microbe-metabolite-host relationships in UC for future experimental validation.

Humans↗