Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational↗

The inner nuclear membrane protein Lem2 is critical for normal nuclear envelope morphology.

The inner nuclear membrane (INM) of eukaryotic cells is characterized by a unique set of transmembrane proteins which interact with chromatin and/or the nuclear lamina. The number of identified INM proteins is steadily increasing, mainly as a result of proteomic and computational approaches. However, despite a link between mutation of several of these proteins and disease, the function of most transmembrane proteins of the INM remains unknown and depletion of many of these proteins from a variety of systems did not produce an obvious phenotype in the affected cells. Here, we report that depletion of the conserved INM protein Lem2 from human cell lines leads to abnormally shaped nuclei and severely reduces cell survival. We suggest that interactions of Lem2 with lamins or chromatin are critical for maintaining the integrity of the nuclear envelope.

Cell Survival↗

Advanced techniques in placental biology -- workshop report.

Major advances in placental biology have been realized as new technologies have been developed and existing methods have been refined in many areas of biological research. Classical anatomy and whole-organ physiology tools once used to analyze placental structure and function have been supplanted by more sophisticated techniques adapted from molecular biology, proteomics, and computational biology and bioinformatics. In addition, significant refinements in morphological study of the placenta and its constituent cell types have improved our ability to assess form and function in highly integrated manner. To offer an overview of modern technologies used by investigators to study the placenta, this workshop: Advanced techniques in placental biology, assembled experts who discussed fundamental principles and real time examples of four separate methodologies. Y. Sadovsky presented the principles of microRNA function as an endogenous mechanism of gene regulation. J. Robinson demonstrated the utility of correlative microscopy in which light-level and transmission electron microscopy are combined to provide cellular and subcellular views of placental cells. A. Croy provided a lecture on the use of microdissection techniques which are invaluable for isolating very small subsets of cell types for molecular analysis. Finally, G. Rice presented an overview methods on profiling of complex protein mixtures within tissue and/or fluid samples that, when refined, will offer databases that will underpin a systems approach to modern trophoblast biology.

Animals↗

The future of education in the molecular life sciences.

The changing landscape of education in biochemistry and molecular biology presents many challenges for the future, for students and educators alike. The exponential increase in knowledge, the genomics, proteomics and computing revolutions, and the merging of once separate fields in biology, chemistry, physics and mathematics, mean that we need to rethink how we should be preparing today's science undergraduates for the future. What do we need to change, and how will we implement it?

Biochemistry↗

engGNN: a dual-graph neural network for omics-based disease classification and feature selection.

Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph neural networks offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated graph, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive representations, thereby improving predictive performance and interpretability. Through extensive simulation studies and real-world applications to three independent gene expression datasets, engGNN consistently demonstrates strong classification performance compared with competitive baselines. Beyond classification, engGNN provides feature- and source-level interpretability, enabling biologically meaningful analyses such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.

Graph Neural Networks↗

SpectroPipeR-a streamlining post Spectronaut® DIA-MS data analysis R package.

SUMMARY: Proteome studies frequently encounter challenges in down-stream data analysis due to limited bioinformatics resources, rapid data generation, and variations in analytical methods. To address these issues, we developed SpectroPipeR, an R package designed to streamline data analysis tasks and provide a comprehensive, standardized pipeline for Spectronaut® DIA-MS data. This novel package automates various analytical processes, including XIC plots, ID rate summary, normalization, batch and covariate adjustment, relative protein quantification, multivariate analysis, and statistical analysis, while generating interactive HTML reports for e.g. ELN systems. AVAILABILITY AND IMPLEMENTATION: The SpectroPipeR package (manual: https://stemicha.github.io/SpectroPipeR/) was written in R and is freely available on GitHub (https://github.com/stemicha/SpectroPipeR).

Software↗

Predicting coarse-grained representations of biogeochemical cycles from metabarcoding data.

MOTIVATION: Taxonomic analysis of environmental microbial communities is now routinely performed thanks to advances in DNA sequencing. Determining the role of these communities in global biogeochemical cycles requires the identification of their metabolic functions, such as hydrogen oxidation, sulfur reduction, and carbon fixation. These functions can be directly inferred from metagenomics data, but in many environmental applications metabarcoding is still the method of choice. The reconstruction of metabolic functions from metabarcoding data and their integration into coarse-grained representations of biogeochemical cycles remains a difficult bioinformatics problem today. RESULTS: We developed a pipeline, called Tabigecy, which exploits taxonomic affiliations to predict metabolic functions constituting biogeochemical cycles. In a first step, Tabigecy uses the tool EsMeCaTa to predict consensus proteomes from input affiliations. To optimize this process, we generated a precomputed database containing information about 2404 taxa from UniProt. The consensus proteomes are searched using bigecyhmm, a newly developed Python package relying on Hidden Markov Models to identify key enzymes involved in metabolic function of biogeochemical cycles. The metabolic functions are then projected on coarse-grained representation of the cycles. We applied Tabigecy to two salt cavern datasets and validated its predictions with microbial activity and hydrochemistry measurements performed on the samples. The results highlight the utility of the approach to investigate the impact of microbial communities on biogeochemical processes. AVAILABILITY AND IMPLEMENTATION: The Tabigecy pipeline is available at https://github.com/ArnaudBelcour/tabigecy. The Python package bigecyhmm and the precomputed EsMeCaTa database are also separately available at https://github.com/ArnaudBelcour/bigecyhmm and https://doi.org/10.5281/zenodo.13354073, respectively.

Metagenomics↗

iModMix: integrative module analysis for multi-omics data.

SUMMARY: Integrative Module Analysis for Multi-omics Data (iModMix) is a biology-agnostic framework that enables the discovery of novel associations across any type of quantitative abundance data, including but not limited to transcriptomics, proteomics, and metabolomics. Instead of relying on pathway annotations or prior biological knowledge, iModMix constructs data-driven modules using graphical lasso to estimate sparse networks from omics features. These modules are summarized into eigenfeatures and correlated across datasets for horizontal integration, while preserving the distinct feature sets and interpretability of each omics type. iModMix operates directly on matrices containing expression or abundances for a wide range of features, including but not limited to genes, proteins, and metabolites. Because it does not rely on annotations (e.g., KEGG identifiers), it can seamlessly incorporate both identified and unidentified metabolites, addressing a key limitation of many existing metabolomics tools. iModMix is available as a user-friendly R Shiny application requiring no programming expertise (https://imodmix.moffitt.org), and as a Bioconductor R package for advanced users (https://bioconductor.org/packages/release/bioc/html/iModMix.html). The tool includes several public and in-house datasets to illustrate its utility in identifying novel multi-omics relationships in diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: iModMix is freely available from Bioconductor (https://bioconductor.org/packages/release/bioc/html/iModMix.html), and the example dataset package (iModMixData) is also available from Bioconductor (https://bioconductor.org/packages/release/ data/experiment/html/iModMixData.html). The R package source code and Docker are available from GitHub: https://github.com/biodatalab/iModMix. Shiny application can be accessed at: https://imodmix.moffitt.org.

Multiomics↗

Qualitatively modelling and analysing genetic regulatory networks: a Petri net approach.

MOTIVATION: New developments in post-genomic technology now provide researchers with the data necessary to study regulatory processes in a holistic fashion at multiple levels of biological organization. One of the major challenges for the biologist is to integrate and interpret these vast data resources to gain a greater understanding of the structure and function of the molecular processes that mediate adaptive and cell cycle driven changes in gene expression. In order to achieve this biologists require new tools and techniques to allow pathway related data to be modelled and analysed as network structures, providing valuable insights which can then be validated and investigated in the laboratory. RESULTS: We propose a new technique for constructing and analysing qualitative models of genetic regulatory networks based on the Petri net formalism. We take as our starting point the Boolean network approach of treating genes as binary switches and develop a new Petri net model which uses logic minimization to automate the construction of compact qualitative models. Our approach addresses the shortcomings of Boolean networks by providing access to the wide range of existing Petri net analysis techniques and by using non-determinism to cope with incomplete and inconsistent data. The ideas we present are illustrated by a case study in which the genetic regulatory network controlling sporulation in the bacterium Bacillus subtilis is modelled and analysed. AVAILABILITY: The Petri net model construction tool and the data files for the B. subtilis sporulation case study are available at http://bioinf.ncl.ac.uk/gnapn.

Algorithms↗

26th Lauriston S. Taylor Lecture: developing mechanistic data for incorporation into cancer and genetic risk assessments: old problems and new approaches.

The theme that runs through this 26th Taylor Lecture is the question of how can data on the mechanism of induction of genetic alterations by radiations and chemicals be used to support the development of risk estimates, particularly at low exposure levels. The premise is that chromosomal alterations are involved in the development of tumors and birth defects, and that data generated for genetic alterations can be interpreted in terms of these adverse health outcomes. The general conclusions are that chromosomal alterations can be induced by ionizing radiations by a single energy loss event in a target of the size of a DNA molecule and that aberrations generally result from misrepair or failure to repair the induced lesions (generally assumed to be double-strand breaks). Chromosomal alterations induced by chemicals are produced almost exclusively by replication errors on a damaged DNA template. Thus, cell cycle stage and DNA repair and replication fidelity will be influential on overall sensitivity to aberration induction. These same features are also important in considerations of genetic susceptibility-alterations in cell cycle control or DNA repair or replication fidelity can alter sensitivity. The differences in mechanism of induction of chromosomal aberrations by ionizing radiation and chemicals is most important when considering cells at risk and comparative sensitivities among species and cell types. Models of cancer induction have gradually evolved from initiation, promotion, and progression models to multistep genetic models to the most recent one of six acquired characteristics. This evolution has passed the level of concentration of research from single gene, single cell to multiple genes (pathways), and whole tissues. The latter areas of concentration are ideal for addressing with the new genomics, proteomics, and computational modeling approaches. The attention is still on the role of genetic alterations in cancer and hereditary effects and the mechanism of their formation--it is the approaches to address these that are changing.

Animals↗

Pheromone signaling mechanisms in yeast: a prototypical sex machine.

The actions of many extracellular stimuli are elicited by complexes of cell surface receptors, heterotrimeric guanine nucleotide-binding proteins (G proteins), and mitogen-activated protein (MAP) kinase complexes. Analysis of haploid yeast cells and their response to peptide mating pheromones has produced important advances in our understanding of G protein and MAP kinase signaling mechanisms. Many of the components, their interrelationships, and their regulators were first identified in yeast. Current analysis of the pheromone response pathway (see the Connections Maps at Science's Signal Transduction Knowledge Environment) will benefit from new and powerful genomic, proteomic, and computational approaches that will likely reveal additional general principles that are applicable to more complex organisms.

Cell Cycle↗

Pheromone signaling pathways in yeast.

The actions of many extracellular stimuli are elicited by complexes of cell surface receptors, heterotrimeric guanine nucleotide-binding proteins (G proteins), and mitogen-activated protein kinase (MAPK) complexes. Analysis of haploid yeast cells and their response to peptide mating pheromones has produced important advances in the understanding of G protein and MAPK signaling mechanisms. Many of the components, their interrelationships, and their regulators were first identified in yeast. Examples include definitive demonstration of a positive signaling role for G protein betagamma subunits, the discovery of a three-tiered structure of the MAPK module, development of the concept of a kinase-scaffold protein, and the discovery of the first regulator of G protein signaling protein. New and powerful genomic, proteomic, and computational approaches available in yeast are beginning to uncover new pathway components and interactions and have revealed their presence in unexpected locations within the cell. This updated Connections Map in the Database of Cell Signaling includes several major revisions to this prototypical signal response pathway.

GTP-Binding Proteins↗

Structural similarity between the bone marrow extracellular matrix protein and neurokinin 1 could be the limiting factor in the hematopoietic effects of substance P.

In the adult bone marrow (BM), immune cells are replenished through the process of definitive hematopoiesis, which is regulated by a complex process of cellular and humoral interactions. The latter include substance P (SP), a neurotransmitter that is produced by neural and nonneural cells. Neurokinin-1 (NK-1), the high-affinity SP receptor, shares structural similarity with fibronectin, a component of the BM extracellular matrix proteins. This study examines how such similarity could alter the effects of SP on the proliferation of the immature BM progenitors. In vitro studies show that 1 ng fibronectin/mL enhanced the stimulatory effect of SP on the proliferation of primitive BM progenitors. This finding was studied by computational studies: proteomics and three-dimensional molecular modeling. Use of surface-enhanced laser desorption/ionization ProteinChip technology showed that despite the induction of neutral endopeptidase, exogenous fibronectin hindered the degradation of SP to SP(1-4). These findings support a protective role for fibronectin in the digestion of SP. Since SP(1-4) is a negative regulator of hematopoiesis, this report indicates that the structural similarity between fibronectin and NK-1 could be important for maintaining hematopoietic stimulation. These studies could be extrapolated to hematological disorders that are associated with SP-fibronectin complexes.

Bone Marrow↗

A modular approach to the ECVAM principles on test validity.

The European Centre for the Validation of Alternative Methods (ECVAM) proposes to make the validation process more flexible, while maintaining its high standards. The various aspects of validation are broken down into independent modules, and the information necessary to complete each module is defined. The data required to assess test validity in an independent peer review, not the process, are thus emphasised. Once the information to satisfy all the modules is complete, the test can enter the peer-review process. In this way, the between-laboratory variability and predictive capacity of a test can be assessed independently. Thinking in terms of validity principles will broaden the applicability of the validation process to a variety of tests and procedures, including the generation of new tests, new technologies (for example, genomics, proteomics), computer-based models (for example, quantitative structure-activity relationship models), and expert systems. This proposal also aims to take into account existing information, defining this as retrospective validation, in contrast to a prospective validation study, which has been the predominant approach to date. This will permit the assessment of test validity by completing the missing information via the relevant validation procedure: prospective validation, retrospective validation, catch-up validation, or a combination of these procedures.

Animal Testing Alternatives↗

The need for standardization and improved open (meta)data practices in metaproteomics.

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices. Video Abstract.

Proteomics↗

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning↗

Enzyme kinetics shapes the growth response of metabolic networks.

Microbes adjust their metabolism to environmental challenges by changing protein expression levels, metabolite concentrations, and reaction rates. Average expression levels in large proteome sectors change coherently, while individual proteins show divergent shifts even within the same pathway. Here, we establish a metabolic model that integrates local enzyme kinetics and global network architecture to predict the joint growth response of proteins and metabolites. Under nutrient limitation, we predict a remarkably simple pattern of proteome reallocation with growth rate: protein expression levels change linearly but heterogeneously. For a given enzyme, the direction of change is determined by its local kinetic constants - catalytic rate and substrate affinity - and by the degree of nutrient restriction affecting its embedding pathway. This double-graded growth response of the proteome is mediated by restriction-dependent metabolite levels, which are predicted to decrease with growth rate in a nonlinear way. The model establishes three specific growth laws: protein expression changes of individual enzymes are negatively correlated with their expression and with their substrate saturation at high growth; average changes of pathways and larger functional sectors are correlated with their internal variance. These predictions are in quantitative agreement with measured system-wide proteomics and metabolomics data of E. coli. Enzyme-specific response patterns are a starting point for model-guided interventions into bacterial metabolism.

Kinetics↗