Search PubMedSearch

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

OmicsQ: a user-friendly platform for interactive quantitative omics data analysis.

MOTIVATION: High-throughput omics technologies generate complex datasets with thousands of features that are quantified across multiple experimental conditions, but often suffer from incomplete measurements, missing values, and individually fluctuating variances. This requires analytical tools for accurate, deep and insightful biological interpretation, capable of dealing with a large variety of data properties and different amounts of completeness. Software capable of handling such data complexity and integrating with external applications for downstream analysis remains rare and mostly relies on programming-based environments, limiting accessibility for researchers without computational expertise. RESULTS: We present OmicsQ, an interactive, web-based platform designed to streamline quantitative omics data analysis. OmicsQ provides an intuitive, browser-based visualization interface that integrates established statistical processing tools. Those include robust batch correction, automated experimental design annotation, and handling of missing data without imputation, which maintains data integrity and avoids artifacts from a priori assumptions. OmicsQ seamlessly interacts with external applications (e.g. PolySTest, VSClust, ComplexBrowser) for statistical testing, clustering, analysis of protein complex behavior, and pathway enrichment, offering a comprehensive and flexible workflow from data import to biological interpretation that is broadly applicable across domains. AVAILABILITY AND IMPLEMENTATION: OmicsQ is implemented in R and Shiny and is available at https://computproteomics.bmb.sdu.dk/app_direct/OmicsQ. Source code and installation instructions: https://github.com/computproteomics/OmicsQ, DOI: 10.5281/zenodo.17778420.

Software

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software

Cell cycle-dependent protein dynamics in budding yeast resolved by deconvolution of bulk proteomics.

The cell division cycle is characterised by oscillatory dynamics in regulatory mechanisms and biosynthesis, coordinated with genome replication and segregation. To understand these dynamics, quantitative cell cycle-dependent protein concentration data are essential. Unfortunately, accurately resolving cell cycle-dependent protein dynamics is challenging because single-cell proteomics is currently infeasible and bulk proteomics requires - inherently imperfect - cell synchronisation. Here, we developed a computational method to deconvolve cell cycle-dependent protein concentration dynamics and applied it to new budding yeast bulk proteome data. Key to this method was a yeast population model, parameterised with experimental cell cycle progression and volume growth data, for quantifying the desynchronisation in sampled populations. We performed deconvolution on 3272 proteins, using cross-validation to determine regularisation parameters, and identified 539 proteins with cell cycle-dependent dynamics. Many of these dynamics were consistent with known yeast biology and dynamic proteins were enriched for several metabolic process, extending previous observations and supporting the emerging picture of metabolic activity as varying substantially over cell cycle phases. We consider the generated cell cycle-resolved budding yeast proteome data a key resource.

Journal Article

Proteomic discovery analysis of quantitatively assessed emphysema in the general population. The MESA Lung Study.

BACKGROUND: Pulmonary emphysema occurs frequently in older adults, often without airflow limitation. Its presence predicts symptoms, respiratory hospitalizations and deaths, and all-cause mortality. Proteomics may provide further insights into emphysema pathogenesis and inform therapeutic targets. OBJECTIVE: We performed a proteomic discovery analysis of percent emphysema on computed tomography (CT) in a population-based, multiethnic sample from the Multi-Ethnic Study of Atherosclerosis (MESA) Lung Study. Replication was performed in two chronic obstructive pulmonary disease (COPD)-based studies, the SubPopulations and InteRmediate Outcome Measures in COPD Study (SPIROMICS) and the Genetic Epidemiology of COPD (COPDGene) Study. METHODS: MESA recruited participants from the general population in 2000-02. The MESA Lung Study performed full-lung CT scans in 2010-12. Percent emphysema was defined as the percentage of lung voxels&#x2009;<&#x2009;-950 Hounsfield units. Over 7,200 plasma aptamers were measured via SomaScan. Cross-sectional linear and least absolute shrinkage and selection operator (LASSO) regression models were adjusted for demographics, anthropometrics, smoking, renal function, and scanner parameters. Statistical significance was defined as a false discovery rate p-value&#x2009;<&#x2009;0.05. Gene Ontology (GO)/Reactome enrichment analyses were performed. LASSO-selected proteins' predictive performance was evaluated. RESULTS: Among 2,504 participants in the MESA Lung Study, mean age was 69.4&#xa0;years, 1,291 had ever smoked, and median percent emphysema-like lung was 1.4%. In total, 1,234 aptamers were significantly associated with percent emphysema in the MESA Lung Study, and 35 replicated in the SPIROMICS and COPDGene Studies. Novel associations included protein family with sequence similarity (FAM) 177A1, syntenin-2, ubiquitin carboxyl-terminal hydrolase 25, and uncharacterized protein C20orf173. Previously identified emphysema-associated proteins included soluble advanced glycosylation end product-specific receptor (sRAGE), protein S100-A12, high mobility group protein B1, and roundabout homolog 2. Enrichment analyses identified 40 GO biological processes, including chemokine production and regulation and cell-cell adhesion and regulation, and two Reactome pathways, including RAGE signaling. In tenfold cross-validation, novel proteins were largely retained by LASSO (R2&#x2009;=&#x2009;5.4%), improved overall model performance (R2&#x2009;=&#x2009;24.8%), and uniquely explained greater variance in percent emphysema. CONCLUSIONS: This analysis in a general population sample identified novel and previously characterized proteins whose functional roles were validated by GO/Reactome enriched pathways, offering new insights into emphysema pathophysiology and therapeutics.

Humans

Hepatic metabolic adaptation to endurance exercise: temporal and sex differences by multiomics integration and validation.

BACKGROUND: Although endurance exercise benefits liver health, sex-specific adaptive trajectories remain unclear. This study mapped dynamic liver adaptation in males and females during prolonged training and identified underlying molecular programs. METHODS: Using publicly available time-resolved liver multi-omics data generated by the Molecular Transducers of Physical Activity Consortium (MoTrPAC), we established a computational pipeline for differential analysis of transcriptomic, proteomic, phosphoproteomic, and metabolomic data with FDR correction, followed by FGSEA pathway enrichment. Kinase activities were inferred through ortholog mapping and PhosphoSitePlus. Cross-omics co-expression networks were constructed using WGCNA and topological overlap to link omics features with physiological phenotypes. For experimental validation, liver tissues were collected from endurance-trained Sprague-Dawley rats, and key nodes were confirmed by Western blotting, qRT-PCR, and immunofluorescence/immunohistochemical staining. Public scRNA-seq data were further integrated to map multi-omics signals to single-cell resolution and assess functional changes in specific cell types. RESULTS: The hepatic response to exercise stress was stage-specific, shifting from early transcriptional activation to later proteomic and metabolic remodeling. Multi-omics integration revealed distinct sex-associated adaptive trajectories: males were more strongly associated with energy metabolism, redox-related programs, and amino acid/organic acid catabolism, whereas females showed prominent membrane lipid remodeling, proteostasis -related programs, and mitochondrial/ribosomal translational features. Single-cell analysis showed that tissue remodeling occurred without major lineage turnover, instead involving altered communication among pre-existing cell communities. Validation of PPP1R3G identified a protein-dominant exercise-responsive marker, supporting the contribution of post-transcriptional or protein-level regulation. CONCLUSIONS: Hepatic adaptation to endurance stress follows a cross-omics evolutionary pattern with sex-specific reprogramming of energy supply and homeostatic maintenance. This time-resolved framework clarifies how exercise improves liver function and supports sex-oriented metabolic interventions and therapeutic target discovery.

Animals

NExON-Bayes: a Bayesian approach to network estimation informed by ordinal covariates.

MOTIVATION: In heterogeneous disease settings, accounting for intrinsic sample variability is crucial for obtaining reliable and interpretable omic network estimates. However, most graphical model analyses of biomedical data assume homogeneous conditional dependence structures, potentially leading to misleading conclusions. To address this, we propose a joint Gaussian graphical model that leverages sample-level ordinal covariates (e.g. disease stage) to account for heterogeneity and improve the estimation of partial correlation structures. RESULTS: Our modelling framework, called NExON-Bayes, extends the graphical spike-and-slab framework to account for ordinal covariates, jointly estimating their relevance to the graph structure and leveraging them to improve the accuracy of network estimation. To scale to high-dimensional omic settings, we develop an efficient variational inference algorithm tailored to our model. Through simulations, we demonstrate that our method outperforms the vanilla graphical spike-and-slab (with no covariate information), as well as other state-of-the-art network approaches which exploit covariate information. Applying our method to reverse phase protein array data from patients diagnosed with stage I, II or III breast carcinoma, we estimate the behaviour of proteomic networks as cancer progresses. Our model provides insights not only through inspection of the estimated proteomic networks, but also of the estimated ordinal covariate dependencies of key groups of proteins within those networks, offering a comprehensive understanding of how biological pathways shift across disease stages. AVAILABILITY AND IMPLEMENTATION: A user-friendly R package for NExON-Bayes with tutorials is available on Github at github.com/jf687/NExON, and archived at https://doi.org/10.5281/zenodo.20312938. The source of the dataset used is cited in the relevant section.

Bayes Theorem

Dissecting spatial patterning and signaling with directional diffusion in spatial multi-omics.

Spatial multi-omics sequencing enables the simultaneous profiling of transcriptomics, proteomics, and epigenomics at a spatial resolution, offering insights into complex tissue organization and molecular regulation. However, the effective integration of multiple omics modalities in a spatial context remains a major challenge. Here, we present SpaDDM, a spatial multi-omics integration framework based on directional diffusion models (DDMs), which supports spatial pattern identification, cross-omics alignment, and inter-and intracellular signaling flow analysis. SpaDDM employs DDM-based graph networks to learn omics-specific representations by jointly incorporating spatial coordinates and molecular measurements within each modality, followed by an attention mechanism to align features across modalities. We benchmarked SpaDDM on diverse spatial multi-omics datasets, including transcriptomics-epigenomics and transcriptomics-proteomics combinations across multiple tissues and species. SpaDDM consistently outperformed existing methods by more accurately deciphering spatial tissue patterns and effectively reducing the boundary noise between spatial regions. Moreover, the learned low-dimensional coembedded representations of individual cells serve as integral mediators for inferring the signaling flows that underlie spatial patterning. Finally, we demonstrated that SpaDDM alignment of complementary information across multi-omics layers facilitates cross-omics translation and significantly improves the prediction of cell state alignments.

Multiomics

SpatioMark: quantifying the impact of spatial proximity on cell phenotype.

MOTIVATION: As research advances in spatially resolving the biological archetype of various diseases, technologies that capture the spatial relationships between cells are demonstrating increasing value. Whilst there are an increasing number of analytical methods being developed to identify the complex web of interactions between cells, the downstream impacts of these cell-cell relationships are under explored. RESULTS: We present SpatioMark, a statistical framework that simplifies the assessment of gene or protein expression changes within a cell type that are associated with the spatial proximity to other cell types. We demonstrate its performance across spatial proteomics and transcriptomics datasets. We link identified relationships with differences in patient survival. We highlight key challenges in identifying changes in molecular markers associated with the localization of cells. We propose correction strategies that reduce artefact-induced relationships. AVAILABILITY AND IMPLEMENTATION: SpatioMark is implemented in the Statial R package on Bioconductor: https://bioconductor.org/packages/release/bioc/html/Statial.html.

Humans

DARKIN: a zero-shot benchmark for phosphosite-dark kinase association using protein language models.

MOTIVATION: Protein language models (pLMs) have emerged as powerful tools for capturing the intricate information encoded in protein sequences, facilitating various downstream protein prediction tasks. With numerous pLMs available, there is a critical need for diverse benchmarks to systematically evaluate their performance across biologically relevant tasks. Here, we introduce DARKIN, a zero-shot classification benchmark designed to assign phosphosites to understudied kinases, termed dark kinases. Kinases, which catalyze phosphorylation, are central to cellular signaling pathways. While phosphoproteomics enables the large-scale identification of phosphosites, determining the cognate kinase responsible for the phosphorylation event remains an experimental challenge. RESULTS: In DARKIN, we prepared training, validation, and test folds that respect the zero-shot nature of this classification problem, incorporating stratification based on kinase groups and sequence similarity. We evaluated multiple pLMs using two zero-shot classifiers: a novel, training-free k-NN-based method, and a bilinear classifier. Our findings indicate that ESM, ProtT5-XL, and SaProt exhibit superior performance on this task. DARKIN provides a challenging benchmark for assessing pLM efficacy and fosters deeper exploration of under-characterized (dark) kinases by offering a biologically relevant test bed. AVAILABILITY AND IMPLEMENTATION: The DARKIN benchmark data and the scripts for generating additional splits are publicly available at: https://github.com/tastanlab/darkin.

Protein Kinases

IgStrand: A universal residue numbering scheme for the immunoglobulin-fold (Ig-fold) to study Ig-proteomes and Ig-interactomes.

The Immunoglobulin fold (Ig-fold) is found in proteins from all domains of life and represents the most populous fold in the human genome, with current estimates ranging from 2 to 3% of protein coding regions. That proportion is much higher in the surfaceome where Ig and Ig-like domains orchestrate cell-cell recognition, adhesion and signaling. The ability of Ig-domains to reliably fold and self-assemble through highly specific interfaces represents a remarkable property of these domains, making them key elements of molecular interaction systems: the immune system, the nervous system, the vascular system and the muscular system. We define a universal residue numbering scheme, common to all domains sharing the Ig-fold in order to study the wide spectrum of Ig-domain variants constituting the Ig-proteome and Ig-Ig interactomes at the heart of these systems. The "IgStrand numbering scheme" enables the identification of Ig structural proteomes and interactomes in and between any species, and comparative structural, functional, and evolutionary analyses. We review how Ig-domains are classified today as topological and structural variants and highlight the "Ig-fold irreducible structural signature" shared by all of them. The IgStrand numbering scheme lays the foundation for the systematic annotation of structural proteomes by detecting and accurately labeling Ig-, Ig-like and Ig-extended domains in proteins, which are poorly annotated in current databases and opens the door to accurate machine learning. Importantly, it sheds light on the robust Ig protein folding algorithm used by nature to form beta sandwich supersecondary structures. The numbering scheme powers an algorithm implemented in the interactive structural analysis software iCn3D to systematically recognize Ig-domains, annotate them and perform detailed analyses comparing any domain sharing the Ig-fold in sequence, topology and structure, regardless of their diverse topologies or origin. The scheme provides a robust fold detection and labeling mechanism that reveals unsuspected structural homologies among protein structures beyond currently identified Ig- and Ig-like domain variants. Indeed, multiple folds classified independently contain a common structural signature, in particular jelly-rolls. Examples of folds that harbor an "Ig-extended" architecture are given. Applications in protein engineering around the Ig-architecture are straightforward based on the universal numbering.

Humans

Bioinformatics pipeline for the systematic mining genomic and proteomic variation linked to rare diseases: The example of monogenic diabetes.

Monogenic diabetes is characterized as a group of diseases caused by rare variants in single genes. Like for other rare diseases, multiple genes have been linked to monogenic diabetes with different measures of pathogenicity, but the information on the genes and variants is not unified among different resources, making it challenging to process them informatically. We have developed an automated pipeline for collecting and harmonizing data on genetic variants linked to monogenic diabetes. Furthermore, we have translated variant genetic sequences into protein sequences accounting for all protein isoforms and their variants. This allows researchers to consolidate information on variant genes and proteins linked to monogenic diabetes and facilitates their study using proteomics or structural biology. Our open and flexible implementation using Jupyter notebooks enables tailoring and modifying the pipeline and its application to other rare diseases.

Humans

Associations of High Attenuation Area-Related Proteomic Biomarkers with Fibrotic or Subpleural Interstitial Lung Abnormalities.

Rationale: High-attenuation area (HAA) is a computed tomography (CT) tool that correlates with lung inflammation and fibrosis. Systemic molecular correlates of HAA (e.g., plasma proteins) may inform biological processes involved in interstitial lung disease. Objectives: To identify plasma proteins that associate with HAA and correlate with a higher probability of developing new-onset fibrotic or subpleural interstitial lung abnormalities (ILAs). Methods: Plasma protein levels were measured using a semiquantitative aptamer-based platform in MESA (the Multi-Ethnic Study of Atherosclerosis; N&#x2009;=&#x2009;5,486) and SPIROMICS (Subpopulations and Intermediate Outcome Measures in COPD Study; N&#x2009;=&#x2009;1,781). Linear regression models identified HAA-associated proteins after adjustment for demographic and socioeconomic factors, CT scanner parameters, study center, and batch. Associations of HAA-related proteins with new-onset fibrotic or subpleural ILAs were examined in MESA participants with ILA assessments on full-lung CT 10 years later. Immunohistochemical staining of select proteins was performed in lung tissue from pulmonary fibrosis cases. Measurements and Main Results: There were 75 proteins detected that were significantly associated with HAA in MESA and SPIROMICS. Gene Ontology analysis of these proteins identified processes involved in immune cell chemotaxis and cellular growth and apoptosis. Seven proteins were associated with a higher probability of new-onset fibrotic or subpleural ILAs in MESA, and two of these, junctional adhesion molecule-like protein and GTP cyclohydrolase 1 feedback regulatory protein, stained in areas of fibrosis in lung tissue from patients with interstitial lung disease. Conclusions: Plasma proteins associated with more HAA are involved in immune and cellular processes and associate with new-onset fibrotic-subpleural ILA.

Humans

Identifying potential biomarkers in the hippocampus of chronic fatigue syndrome rats treated with moxibustion at Zusanli (ST36): a proteomics study.

OBJECTIVE: To observe the effects of moxibustion at Zusanli (ST36) on rats with chronic fatigue syndrome (CFS) and to analyze the mechanisms of moxibustion through hippocampal Proteomics. METHODS: Male Sprague-Dawley (SD) rats were randomly divided into three groups: control group (CON), model group (MOD), and moxibustion group (MOX), with 12 rats in each group. The MOD and MOX groups underwent chronic multi-factor stress stimulation for 35 d to establish the CFS model. After modeling, the rats in the MOX group received mild moxibustion at Zusanli (ST36) (bilateral) for 10 minutes daily for 28 d. During the treatment period, rats in both the MOD and MOX groups continued modeling, while the CON group was kept under normal breeding conditions. The general condition of the rats was monitored, and behaviors were assessed using the Open Field Test (OFT), Exhaustion Treadmill Test, and Morris Water Maze (MWM). Hematoxylin and eosin (HE) staining and transmission electron microscopy (TEM) were employed to observe morphological changes in the hippocampus. Label-free Proteomics were utilized to identify differentially expressed proteins (DEPs) in the hippocampus, followed by bioinformatics analysis. The reliability of the Proteomics results was verified using Parallel Reaction Monitoring. RESULTS: A: Moxibustion at Zusanli (ST36) significantly reduced the general condition score of CFS rats, improved their behavioral performance in OFT, treadmill and MWM, and repaired the pathological and synaptic structural damage in the hippocampus.B: We identified DEPs by applying a fold change threshold of 1.2 and a significance level of P < 0.05. In the comparison between the CON and the MOD, we identified a total of 72 DEPs (31 up-regulated and 41 down-regulated) associated with the development of CFS. In the comparison between the MOX and the MOD group, we identified a total of 103 DEPs (40 up-regulated and 63 down-regulated) related to the therapeutic effects of moxibustion. Gene Ontology (GO) enrichment analysis showed that CFS and moxibustion treatment were related to multiple biological processes, molecular functions, and cellular components. Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis revealed that CFS pathogenesis was linked to base excision repair, steroid biosynthesis, and systemic lupus erythematosus, Furthermore, the treatment of CFS with moxibustion was relevant to terpenoid skeleton biosynthesis.C: Compared with the two comparison groups, we identified 16 potential biomarkers, noting that moxibustion reversed the up-regulation of 14 DEPs and the down-regulation of 2 DEPs in CFS. These proteins are mainly associated with synaptic plasticity, ribosomal function, neurotransmitter secretion, glycine metabolism, and mitochondrial function. CONCLUSION: Moxibustion at Zusanli (ST36) is effective in treating CFS, the potential biomarkers identified by Proteomics confirm that the mechanisms of moxibustion involve multiple targets and pathways, which may be key to regulating the structural and functional damage in the hippocampus associated with CFS, highlighting their significant value for future research.

Animals

Proteomic profiling of bone for the estimation of post-mortem interval and post-mortem submersion interval: a systematic review.

Accurate estimation of the Post-Mortem Interval (PMI) and Post-Mortem Submersion Interval (PMSI) remains a persistent challenge in forensic science, especially when traditional morphological and entomological methods fail due to advanced decomposition or in aquatic environments. Proteomic profiling of bone tissues has recently emerged as a promising approach, leveraging the predictable degradation patterns of bone proteins to estimate time since death more reliably. This systematic review, conducted in accordance with PRISMA guidelines, analyzed 24 peer-reviewed studies focusing on the application of proteomic techniques to bone tissue for PMI and PMSI estimation. The included studies were evaluated based on sample type, analytical techniques used, identified biomarkers, environmental conditions assessed, and the overall reliability and reproducibility of the findings. The review found that specific bone proteins, particularly collagen, osteocalcin, fetuin-A, etc. exhibited consistent degradation patterns that correlated strongly with elapsed post-mortem time. Cortical bone was identified as a more stable and informative matrix compared to trabecular bone. Mass spectrometry, especially LC-MS/MS, emerged as the predominant analytical technique due to its high sensitivity and accuracy in detecting low-abundance proteins over extended PMIs and PMSIs. However, protein degradation rates were significantly influenced by environmental variables such as temperature, humidity, soil pH, and microbial activity. This review also emphasizes the transformative role of bone proteomics in advancing forensic science while identifying key gaps that must be addressed to achieve global standardization and practical implementation in diverse forensic contexts. The integration of proteomics with other emerging technologies, such as machine learning algorithms and computational modeling, may further enhance the precision of PMI and PMSI estimation in future applications.

Postmortem Changes

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem

Using Large Genomic Biobanks to Generate Insights into Genetic Kidney Disease.

Chronic kidney disease (CKD) affects approximately 9% of the global population, leading to increased risks of end-stage kidney disease (ESKD), cardiovascular disease (CVD), and mortality. Patients with CKD are a huge burden on health care resources globally. CKD is a complex condition influenced by a combination of genetic, environmental, and traditional risk factors. Family studies have suggested heritability rates for CKD ranging from 30% to 75%, and large genomic biobank studies have proven essential in identifying genes with substantial effects on CKD risk and in capturing cumulative genetic risk through polygenic risk scores. These biobanks are crucial for discovering new genes associated with kidney health and disease, and their growing size enhances the power to detect novel genetic associations. Integrating multi-omics technologies such as transcriptomics, metabolomics, and proteomics further enriches our understanding of CKD, while advanced computational tools continue to expand our insights into genetic data. Polygenic risk scores, derived from hundreds of genetic variants with small effect sizes, can help identify individuals at high risk of CKD. Genomic biobanks offer valuable opportunities for early identification and personalized treatment of monogenic kidney disorders, such as autosomal dominant polycystic kidney disease and Alport syndrome. These biobanks help fill knowledge gaps, particularly in individuals with milder or asymptomatic presentations who are often underrepresented in traditional studies. Expanding genomic biobank efforts globally, especially in diverse populations, is vital to enhancing our understanding of the genetic underpinnings of kidney disease. This review highlights the significant contributions of genomic biobanks to advancing our comprehension of the genetics of CKD.

Humans

The Multi-Omics Landscape of Enzymatic Alterations in Systemic Lupus Erythematosus.

OBJECTIVE: Systemic lupus erythematosus (SLE) is an autoimmune disease closely associated with enzyme dysfunction, yet its underlying molecular mechanisms remain incompletely understood. This study aims to characterize enzyme-network alterations associated with SLE status and disease activity and to identify candidate molecules with potential clinical relevance. METHODS: We integrated proteomic and phosphoproteomic data from peripheral blood mononuclear cells (PBMCs) of 130 SLE patients and 90 healthy controls (HC), along with transcriptomic data from 1461 SLE patients. Through systematic analysis of key enzyme phosphorylation sites, upstream transcription factors (TFs), and computationally prioritized candidate compounds, we sought to characterize enzyme-centered regulatory associations. RESULTS: Integrated proteomic and phosphoproteomic analyses revealed significant metabolic and signaling pathway disturbances, along with distinct phosphorylation patterns in SLE immune cells. Multiple SLE-associated and disease-activity-associated candidate molecules were identified. Regulatory network analysis uncovered an upstream transcription factor cluster centered around STAT1. Computational drug screening identified computationally prioritized candidate compounds with multi-gene DSigDB associations, which require further clinical safety evaluation and experimental validation. CONCLUSIONS: This study constructs a molecular map of SLE, highlighting associations between enzyme-network alterations, catalytic dysregulation, and SLE-related immune molecular signatures, and identifies candidate molecules for future clinical and functional evaluation.

Humans