Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multi-omics data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

A regulatory network underlying idiopathic pulmonary fibrosis.

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease in which genetic susceptibility interacts with epithelial, immune, and mesenchymal remodeling. Although the chromosome 11p15.5 locus contains established IPF susceptibility signals near MUC5B and TOLLIP, the broader regulatory architecture of this region remains incompletely resolved. METHODS: We integrated IPF genome-wide association study summary statistics with methylation, expression, and protein quantitative trait loci using summary-data-based Mendelian randomization (SMR). SMR-prioritized candidates were evaluated in independent transcriptomic and methylation cohorts and further contextualized using microRNA, transcription-factor, protein-interaction, machine-learning, single-cell, and spatial transcriptomic analyses. Fibrosis-associated expression patterns were assessed in a bleomycin-induced pulmonary fibrosis rat model. RESULTS: The analyses recovered the established MUC5B and TOLLIP signals and prioritized BRSK2 as a comparatively underexplored candidate supported by eQTL-based SMR and independent molecular evidence. The BRSK2 pQTL association did not pass the HEIDI test and was therefore not interpreted as convergent protein-level genetic evidence. Network analyses linked BRSK2 to cell-cycle, metabolic-stress, and senescence-related programs, while cross-cohort machine learning prioritized FOXA2, CDC25B, and NFE2 as informative network features. Single-cell and spatial analyses localized BRSK2 preferentially to fibroblast and myofibroblast compartments and to regions with greater histological fibrosis severity. In fibrotic rat lungs, BRSK2 expression increased, whereas FOXA2 and CDC25B decreased at the transcript and protein levels. CONCLUSIONS: These findings refine the molecular landscape of the chromosome 11p15.5 IPF susceptibility locus and prioritize BRSK2 as a candidate component of an IPF-associated profibrotic fibroblast state. Its causal contribution, direct regulatory relationships, and therapeutic tractability require targeted mechanistic validation.

Idiopathic Pulmonary Fibrosis↗

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50 K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40 kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals↗

Multi-omics Mendelian Randomization Prioritizes Neutrophil Extracellular Trap-related Genes Associated with Atrial Fibrillation Risk.

BACKGROUND: Neutrophil extracellular traps (NETs) participate in thrombosis, inflammation, and cardiovascular remodeling, yet whether NET-related genes (NRGs) are associated with atrial fibrillation (AF) risk across multiple molecular layers remains unclear. This study used a multiomics Mendelian randomization framework to prioritize NRGs supported by methylation, expression, and protein quantitative trait loci (QTL) data. METHODS: Genome-wide significant cis instruments (P < 5 &#xd7; 10-8) were obtained for 90 methylation QTLs (mQTLs), 100 expression QTLs (eQTLs), and 38 protein QTLs (pQTLs) mapped to 137 literature- curated NRG entries. Summary-data-based Mendelian randomization (SMR) coupled with the heterogeneity in dependent instruments (HEIDI) test was applied using whole-blood mQTL data (n = 1,980), eQTLGen blood eQTL data (n = 31,684), and deCODE plasma pQTL data (n = 35,559). AF outcome data were obtained from a meta-analysis including 60,620 cases and 970,216 controls of European ancestry. RESULTS: At the methylation level, 21 CpG-feature associations across 13 genes remained significant after HEIDI filtering and false discovery rate (FDR) correction. Expression-level analysis identified eight significant gene-AF associations, whereas protein-level analysis identified seven significant features representing five unique proteins. Cross-omics integration prioritized C3, MAPK3, and STAT3 as Tier 1 genes, CTSC, LPAR3, and THBD as Tier 2 genes, and fourteen additional genes as Tier 3 candidates. C3 showed risk-increasing protein-level associations together with multiple significant CpG signals, whereas MAPK3 and STAT3 showed directionally protective expression/protein or methylation/protein patterns. DISCUSSION: The cross-omics convergence on C3, MAPK3, and STAT3 is consistent with complement activation, immune-fibrotic signaling, and cytokine-regulatory pathways implicated in AF biology, but the findings should be interpreted as genetic prioritization rather than definitive intervention-ready causality. CpG-level heterogeneity at the C3 locus and the blood/plasma origin of the QTL resources further support a cautious interpretation. Modest colocalization support and the unresolved possibility of pQTL sample overlap further support this cautious, hypothesis-generating interpretation. CONCLUSION: Multi-omics SMR prioritizes C3, MAPK3, and STAT3 as the most consistently supported NET-related genes associated with AF risk. These findings provide a framework for atrialtissue replication and mechanistic validation of NET-related pathways in AF.

Atrial fibrillation↗

A distinct effector B cell population drives autoantibody production in SARS-CoV-2 infection.

Autoantibodies (autoAbs) are linked to mortality and Long COVID, yet their cellular origins remain unclear. We analyzed the INCOV cohort and identified 12 age- and sex-matched participants with varying autoAb abundance and integrated single-cell RNA-seq and ATAC-seq data from B cells, plasma proteomics, proteome-wide autoAb profiling, clinical data, and in vitro assays. AutoAb abundance inversely correlated with neutralizing IgG and declined as infection resolved, paralleling the contraction of atypical memory B cells (AtMs). In vitro, AtMs preferentially differentiated into autoAb-producing antibody-secreting cells upon TLR7/8 stimulation. CD11c+ AtMs (double-negative 2, DN2s) in autoAb-high individuals exhibited increased TLR7 signaling, oxidative stress, and isotype switching, regulated by transcription factors T-bet and XBP1. Integrated genetic and genomic analyses showed that DN2s had the strongest enrichment for autoimmune trait heritability and inferred regulatory effects of autoimmune risk variants among B cell subsets. These findings identify DN2s as key precursors of autoAb-producing cells during SARS-CoV-2 infection.

B cell↗

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗

Integrated analysis of plasma metabolomics and proteomics reveals the biological characteristics of damp-heat and stasis-toxin syndrome in colorectal cancer.

OBJECTIVE: To investigate the biological attributes of core syndromes in colorectal cancer, namely, the damp-heat and stasis-toxin syndrome (SRYD). METHODS: Between October 2021 and October 2022, a cohort comprising 40 patients with colorectal cancer (CRC) diagnosed with damp-heat and stasis-toxin syndrome (SRYD group), 40 patients with CRC without this syndrome (non-SRYD group), and 40 healthy controls (Normal group) was recruited at Jiangsu Province Hospital of Chinese Medicine. Untargeted metabolomics analysis was conducted on plasma samples from all 120 participants, while differential protein analysis using four-dimensional data-independent acquisition proteomics was performed on 20 randomly selected samples per group. A combined analysis of proteomics and metabolomics data followed, and the identified potential diagnostic biomarkers were subsequently used to train and validate multiple machine learning models. RESULTS: Proteomic analysis revealed 130 differential proteins in the colorectal cancer with damp-heat and stasis-toxin syndrome (CRC-SRYD) group, enriched in pathways including complement and coagulation cascades, as well as nuclear factor kappa-B (NF-&#x3ba;B) signaling. Metabolomic analysis identified 584 differential metabolites within the same group, showing enrichment in pathways such as primary bile acid biosynthesis, central carbon metabolism in cancer, and glucagon signaling. Integrated pathway analysis indicated heightened activity of the NF-&#x3ba;B signaling pathway in the CRC-SRYD group. A biomarker panel, comprising 6 proteins and 9 metabolites selected through the ReliefF algorithm, was used to construct a diagnostic model with random forest, achieving an accuracy of 93.33%, sensitivity of 80.00%, and specificity of 100%. CONCLUSION: This study systematically elucidates plasma metabolomic and proteomic alterations in patients with CRC, establishing a robust diagnostic model for CRC syndrome (CRC-SRYD). Further investigation is warranted to clarify the underlying molecular mechanisms and biological foundations.

Humans↗

De novo chromatin remodelling variants in sporadic Chiari 1 malformation.

Chiari 1 malformation (CM1) is the most common congenital malformation of the human hindbrain. Although prior studies have implicated chromatin-remodeling genes in CM1, the de novo genetic architecture and underlying neurodevelopmental mechanisms remain incompletely defined. To investigate the molecular genetics of a novel familial form of CM1 linked with syringomyelia and tethered cord and determine whether rare, damaging de novo variants (DNVs) contribute to sporadic CM1 risk with gene- and pathway-level resolution, we performed whole-exome sequencing in an ultra-rare multigenerational family with CM1 and associated spinal pathology, and in the largest assembled trio-based cohort to date, comprising 1,585 proband-parent trios with sporadic, idiopathic CM1 (2017-2025). The comparison cohort included 1,798 unaffected control siblings. Clinical phenotyping was by systematic medical record review. Structural domain mapping, in silico modeling, and integration with single-cell transcriptomic data from developing human cerebellum was conducted to assess biological plausibility. A heterozygous loss-of-function variant in CHD3 segregated with CM1 and syringomyelia in a multigenerational family. In the trio-based cohort, rare protein-altering DNVs were significantly enriched across multiple chromodomain helicase DNA-binding (CHD) genes, including CHD1, CHD3, CHD4, and CHD8, exceeding gene-specific mutation expectations (protein-damaging variants: P = 1.3 &#xd7; 10-9; predicted loss-of-function variants: P = 8.6 &#xd7; 10-5). CHD1 contained two pathogenic DNVs (p.A999D and p.E984K). CHD4 (p.D744N, p.T1813P, and p.I1102T) and CHD8 (p.R1402X, p.R1472X, and p.R2035X) each contained three new DNVs. Variants clustered within conserved ATPase, helicase, and chromodomain regions essential for chromatin remodeling, and these patients frequently had comorbid developmental delay and related neurodevelopmental features. Single-cell transcriptomic analyses demonstrated enrichment in Purkinje cells and inhibitory neurons of midgestational cerebellum, where CHD gene products form a coherent chromatin-regulatory network. Rare, large-effect DNVs that disrupt chromatin-remodeling programs contribute to sporadic CM1, implicating genetically encoded dysregulation of cerebellar development as a central disease mechanism. Exome sequencing may complement surgical evaluation of children with sporadic CM1, particularly when accompanied by neurodevelopmental concerns, informing prognosis and family counseling.

de novo variants↗

Multi-Omics Integration Identifies a Five-Gene Metabolic Signature With Experimental Validation in Clear Cell Renal Cell Carcinoma.

BACKGROUND: Clear cell renal cell carcinoma (ccRCC) is hallmarked by profound metabolic reprogramming; however, its intricate crosstalk with the tumor immune microenvironment (TIME) and its clinical ramifications remain inadequately elucidated. This study aims to systematically decipher the metabolic-immune interplay in ccRCC through multi-omics integration, with the goal of identifying robust prognostic biomarkers and actionable therapeutic vulnerabilities. AIMS: This study aims to systematically decipher the metabolic-immune interplay in clear cell renal cell carcinoma (ccRCC) through multi&#x2011;omics integration, and to identify robust prognostic biomarkers and actionable therapeutic vulnerabilities that can inform precision risk stratification and individualized treatment strategies. METHODS: We integrated bulk transcriptomic, genomic, and clinical data from multiple ccRCC cohorts. Differential expression and functional enrichment analyses were performed to characterize metabolic pathway alterations. Mendelian randomization (MR) was employed to infer causal relationships between metabolic disorders and ccRCC risk. A machine learning-based prognostic framework, incorporating SHAP (SHapley Additive exPlanations) for feature interpretability, was constructed and rigorously validated. TIME heterogeneity was dissected using deconvolution algorithms, while drug sensitivity, tumor mutation burden (TMB), and TIDE scores were utilized to assess therapeutic responses and immune evasion. Candidate gene function was evaluated through in&#xa0;vitro gain- and loss-of-function assays, with expression validated via TCGA, HPA, western blot, and qRT-PCR. RESULTS: Enrichment analysis identified coordinated dysregulation in lipid metabolism, energy homeostasis, and hypoxia response pathways. MR analysis confirmed lipid metabolism disorders as a causal risk factor for ccRCC. Our machine-learning model, centered on five core SHAP-identified features (SUCLA2, ACAT1, PC, SUCLG1, and HMGCS2), demonstrated superior predictive accuracy over conventional clinical staging. Immune profiling unveiled dichotomous TIME states: the low-risk group retained active immune surveillance, whereas the high-risk group was enriched with immunosuppressive subsets. Drug sensitivity screening pinpointed LY2109761 and carmustine as high-risk-specific candidate agents. Furthermore, TMB and TIDE analyses stratified high-risk patients displaying genomic instability and immune evasion phenotypes. Functionally, SUCLA2 knockdown significantly enhanced ccRCC cell proliferation and invasion, while its overexpression suppressed these malignant phenotypes, corroborating its tumor-suppressive role. Expression patterns of the hub genes were consistently validated across multi-level datasets and experimental assays. CONCLUSION: This study establishes a precision oncology framework for ccRCC by functionally linking metabolic biomarkers, immunophenotypes, and stratified therapeutic strategies. Importantly, we identify SUCLA2 as a potential functional tumor suppressor and a promising target for further mechanistic and translational investigation.

Humans↗

Microbial partnerships and molecular mechanisms in plant stress physiology for climate-resilient and sustainable farming.

Plant-microbial partnerships and their underlying molecular mechanisms are indispensable, natural drivers of improved nutrient acquisition and stress tolerance in the face of climate-driven environmental challenges. Modern multi-omics tools, when coupled with artificial intelligence and synthetic biology, enable the precise design of targeted bioinoculants and synthetic microbial consortia. Translating these advanced microbiome-based strategies into scalable, field-level agricultural applications provides a sustainable path toward securing global food production while maintaining soil health. Global climate change imposes multifaceted abiotic and biotic stresses on crops, disrupting physiological and molecular processes and threatening agricultural productivity. Plant-associated microbes represent an underexplored yet powerful ally in enhancing crop resilience. This review presents current knowledge of plant-microbe interactions and the molecular mechanisms governing plant stress physiology, with an emphasis on climate-resilient and sustainable farming. Hence, ever-changing environmental cues pose a significant burden on agricultural productivity, and plant-associated microbial communities modulate a cascade of physiological and molecular responses, including production of phytohormones, signaling, regulation of reactive oxygen species homeostasis, and activation of plant immune responses to help plants withstand stress and enhance productivity. Moreover, root exudates, phytohormones, and quorum sensing mediate the central communication networks, facilitating plant-microbe cross talk. Additionally, the advances in OMICs approaches aid in disentangling the molecular underpinnings of these interactions by providing mechanistic insights and potential candidate gene targets for crop improvement and stress resilience. In the post-genomic era, integrating artificial intelligence and big data analysis to optimize microbiome-based strategies for sustainable agriculture is a new frontier for disentangling plant-microbe symbiosis to improve soil health, enhance crop yields, and improve stress tolerance. Thus, by integrating the ecological, physiological, and molecular perspectives, this review highlights the transformative potential of harnessing plant-microbe symbiosis for climate-resilient and sustainable agriculture.

Stress, Physiological↗

Multi-omic biomarkers in cardiovascular disease: Discovery to clinical translation.

Cardiovascular disease (CVD) remains the leading cause of mortality worldwide, necessitating improved risk stratification and early detection strategies. Multiomics approaches that integrate genomics, transcriptomics, proteomics, metabolomics, and epigenomics offer unprecedented opportunities for biomarker discovery and precision medicine in cardiovascular care. This narrative review examines the current landscape of multiomics biomarkers for CVD, tracing their evolution from discovery to clinical translation. We synthesize evidence from recent studies evaluating the clinical utility of integrated omics approaches across diverse cardiovascular conditions, including atherosclerotic cardiovascular disease, heart failure, and atrial fibrillation. High-throughput proteomics has identified novel protein signatures that enhance cardiovascular risk prediction beyond traditional risk factors. Metabolomics has revealed pathway-specific biomarkers, including trimethylamine N-oxide and lipid species, associated with atherogenesis. Polygenic risk scores derived from genomic data demonstrate incremental value when combined with clinical risk scores. Multiomics biomarkers represent a transformative approach to cardiovascular risk assessment and disease management.

Humans↗

Beyond ion channel dysfunction: Integration of the transcriptome and proteome from patient-specific re-engineered cardiac cells, and population-level QT genome-wide association study reveals broad cellular dysfunction.

BACKGROUND: Congenital long QT syndrome (LQTS) is a cardiac channelopathy with increased risk of cardiac-triggered syncope/seizures, sudden cardiac arrest, and sudden cardiac death. OBJECTIVE: This study aimed to describe the transcriptomic and proteomic profiles in patient-derived inducible pluripotent stem cell-derived cardiomyocyte (iPSC-CM) models of the 3 canonical genotypes of congenital LQTS: LQT1, LQT2, and LQT3 and integrate these omics-level findings with each other and with population/clinical level QT-genome-wide association study (GWAS) data. METHODS: LQT1, LQT2, LQT3 and respective isogenic control iPSC-CMs were cultured, and RNA and protein samples were collected. RNA sequencing and mass spectrometry-enabled proteomic analysis was performed. PrediXcan analysis was performed using QT GWAS summary statistics and transcriptome expression data. Differential gene and protein expression and ingenuity pathway analysis (IPA) was performed comparing each LQT genotype with its respective isogenic control. RESULTS: 1645 differentially expressed genes (DEGs) were identified; 13 were altered in all 3 LQTS genotypes. IPA analysis of DEGs revealed 301 altered pathways; 47 were altered in all LQTS genotypes. Proteomic analysis identified 2561 differentially expressed proteins (DEPs); 30 were altered in all 3 genotypes. IPA analysis of DEPs identified 646 altered pathways. 306 genes/proteins were identified as significantly altered in both the transcriptome and proteome; pathway analysis of these 301 genes identified 201 altered pathways. 7 pathways were altered in all 3 LQTS genotypes in both the transcriptome and proteome. Integration of the population-level PrediXcan results and the cardiomyocyte-derived omics results identified multiple shared pathways. CONCLUSION: Multi-omics analysis of LQTS and integration of omics results with QT GWAS data reveals that primary LQTS-causative ion channel defects precipitate secondary alterations in a wide range of cellular pathways. Our findings suggest more broad molecular level changes throughout the cell. This study lays the foundation for further exploration of broad cellular changes resulting from ion channel disturbances and how they contribute to disease mechanism.

Humans↗

Multi-omics panorama of glaucoma: Pathogenesis, biomarkers, and novel therapeutic strategies.

Glaucoma is a group of irreversible, blinding eye diseases characterized by progressive loss of retinal ganglion cells, leading to gradual visual field defects that severely impact patients' quality of life. Its complex pathophysiological mechanisms remain incompletely understood, limiting the development of early diagnostic and effective therapeutic strategies. Advances in omics technologies have provided new insights into elucidating the pathophysiology of glaucoma. We summarize specific alterations in genomics, transcriptomics, proteomics, metabolomics, epigenomics, and microbiomics associated with glaucoma. We emphasize the systematic analysis of disease mechanisms, identification of clinically applicable biomarkers, and discovery of novel therapeutic targets through the integration of these data. This approach paves new pathways for glaucoma subtype diagnosis and personalized treatment, while also outlining future research directions and challenges.

Humans↗

Decoding the molecular basis of blue grain color codominance in Qingke: Integrative analysis of RNA-seq, DNA methylation, and miRNA-seq.

The grains on single spike of the F1 generation from the cross between blue- and white-grained Qingke (Hordeum vulgare L. var. nudum Hook. f.) are randomly distributed in blue and white colors. This study integrated data from RNA-seq, DNA methylation, and miRNA-seq to analyze this trait. The results showed that the HvF3'5'H gene is likely central to the development of this codominant phenotype. Through cross-validation of three omics approaches, it was found that the HvMYB gene targeted by miR858-z, as well as the WRKY24 and At3g44326 genes targeted by novel-m0152-5p, novel-m0153-5p, and novel-m0154-5p, are correlated with DNA methylation. qRT-PCR analysis confirmed that the four aforementioned genes exhibited variety-specific and developmental stage-specific expression patterns. This study dissects the regulatory network underlying the codominant blue and white grain color divergence on a single Qingke spike from a multi-omics perspective.

DNA Methylation↗

PLSKO: a robust knockoff generator to control false discovery rate in omics variable selection.

MOTIVATION: Integrating the knockoff framework with any variable-selection method delivers stringent false discovery rate (FDR) control without recourse to p-values, offering a powerful alternative for differential expression analysis of high-throughput omics datasets. However, existing knockoff generators rely on restrictive modelling assumptions or coarse approximations that often inflate the FDR when applied to real-world data. RESULTS: We introduce Partial Least Squares Knockoff (PLSKO), an efficient, assumption-free generator that remains robust across diverse omics platforms. Our extensive simulations show that PLSKO is the only method to maintain FDR control with sufficient power in complex non-linear settings. Our semi-simulation studies drawn from RNA-seq, proteomics, metabolomics, and microbiome experiments confirm PLSKO generates valid knockoff variables. In pre-eclampsia multi-omics case studies, we combine PLSKO with Aggregation Knockoff to address the randomness of knockoffs and improve power, and demonstrate the method's ability to recover biologically meaningful features. AVAILABILITY AND IMPLEMENTATION: Our proposed algorithm is available on Github (https://github.com/guannan-yang/PLSKO) and Zenodo (https://doi.org/10.5281/zenodo.16879594).

Algorithms↗

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software↗

Integrative metabolomic and proteomic analysis of diabetic kidney disease progression with younger-onset type 2 diabetes.

AIM: Younger-onset type 2 diabetes (YT2D) confers a disproportionately high risk of diabetic kidney disease (DKD), yet early biomarkers and underlying mechanisms remain poorly defined. We aimed to identify metabolites associated with DKD progression and integrate metabolomic and proteomic data to elucidate pathways involved in a multi-ethnic Asian cohort. MATERIALS AND METHODS: In this prospective study, 787 YT2D patients (diagnosed at &#x2264; age 40) were followed for a median of 5.7&#x2009;years. DKD progression was defined as an annual decline in estimated glomerular filtration rate (eGFR) of &#x2265;3&#x2009;mL/min/1.73&#x2009;m2 or&#x2009;&#x2265;&#x2009;40% reduction in eGFR from baseline. Plasma metabolites were measured by nuclear magnetic resonance spectroscopy. Multivariable regression analysis was performed in a discovery (N&#x2009;=&#x2009;550) and internal validation cohort (N&#x2009;=&#x2009;237). Integrative metabolomic-proteomic analysis (N&#x2009;=&#x2009;428) was performed using sparse partial least squares discriminant analysis (sPLS-DA). RESULTS: Ninety-eight metabolites were differentially expressed between DKD progressors and non-progressors, of which total branched-chain amino acids (BCAAs) (OR&#x2009;=&#x2009;0.60, 95% CI 0.46-0.79), valine (OR&#x2009;=&#x2009;0.62, 95% CI 0.48-0.81), and leucine (OR&#x2009;=&#x2009;0.56, 95% CI 0.43-0.74) associated with DKD progression, independent of metabolic risk factors. Integrative analysis identified three components comprising 23 proteins and 30 metabolites, involved in the citrate cycle and apoptosis, which improved prediction of DKD progression beyond clinical risk factors (AUC 0.69-0.83). CONCLUSION: Lower plasma BCAA levels are independently associated with DKD progression in YT2D. Integrative multi-omics analysis highlights disruptions in metabolic and apoptotic pathways, providing insights into DKD pathophysiology and potential biomarkers for early risk stratification.

Humans↗

Radiogenomics predicts immune microenvironment heterogeneity and response to combination immunotherapy in hepatocellular carcinoma.

BACKGROUND: The combination of immune checkpoint inhibitors (ICIs) with anti-angiogenic agents is the preferred first-line therapy option for patients with advanced hepatocellular carcinoma (HCC), yet only a subset of patients responds, urging the quest for prediction biomarkers. We aimed to integrate genomics with radiology to propose an immune-derived radiogenomics biomarker of response to such combination immunotherapy and evaluate its added value in clinical context. METHODS: We integrated bulk RNA sequencing (RNA-seq) and proteomics data of 994 HCC patients with single-cell RNA-seq data of 11 samples across multiple datasets to identify an immune-related signature (IRS) that may influence sensitivity or resistance to such combined immunotherapy strategy, followed by verification of selected marker genes using immunohistochemistry and cytological experiments. We then trained/validated a cross-modality radiogenomics biomarker using machine learning based on TCIA database that was further tested in multi-scale independent cohorts covering 754 HCC patients. RESULTS: Integrative multi-omics analysis identifed a parsimonious 2-gene prognostic signature including KPNA2 and SMG5 that was significantly associated with immune heterogeneity and response to combination immunotherapy. Machine-learning pipeline exported the optimal 4-feature radiogenomics biomarker using support vector machine that significantly discriminated prognosis (hazard ratio 1.415&#x2013;1.890; p&#x2009;<&#x2009;0.05 for all) and modestly predicted response to ICI plus anti-angiogenic therapy (area under the curve 0.720&#x2013;0.829) in independent retrospective series across major imaging modalities (computed tomography/magnetic resonance imaging). In a prospective neoadjuvant cohort, this biomarker also showed favorable performance for predicting pathological response and tumor recurrence, accompanied by biological validation through single-cell RNA-seq analysis of pre-treatment biopsies. CONCLUSIONS: Our study provides a cross-device-cross-modal radiogenomics biomarker that can improve patient selection for emerging ICI plus anti-angiogenic therapy with novel potential therapeutic targets in HCC.

Humans↗

Bioinformatics in crop research: using genomic data for crop improvement.

Sustainable crop development aims to maintain or increase yields while reducing environmental impact and managing the challenges imposed by climate change. As the global population grows and arable land becomes scarcer, the integration of molecular breeding with bioinformatics has emerged as an effective strategy for long-term crop improvement. Bioinformatics enables researchers to analyze and interpret the vast quantities of genetic data generated by high-throughput sequencing, making it possible to identify molecular markers, candidate genes, and regulatory networks linked to specific agronomic traits, which breeders then translate into focused, ecologically sustainable breeding programs. This approach has enabled major progress across several fronts: the identification of genes conferring resistance to biotic stressors (pests, pathogens) and abiotic stressors (drought, salinity, heat); the development of nutrient-efficient, low-input crop varieties; the improvement of agronomic performance and nutritional quality through identification of yield- and quality-related genes; and the conservation and deployment of genetic diversity to safeguard long-term breeding sustainability. By combining genomic data with precision breeding techniques, researchers are developing crops that are better adapted to a growing population and a changing climate, positioning the integration of molecular breeding and bioinformatics as a central pillar of future global food security.

bioinformatics↗