Search PubMedSearch

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Pathway-driven target prioritisation in drug discovery.

Genome-scale association studies and functional screens routinely implicate hundreds of candidate genes per disease, yet only a few will be clinically validated as drug targets. Choosing which to pursue is a central drug-discovery decision that depends on interpreting each candidate in its biological context. Curated pathway databases provide this context, while enrichment analysis applies it at scale, turning gene-level signals from genome-wide association, transcriptomic, proteomic and CRISPR studies into mechanistic hypotheses for prioritisation. This review examines how pathway-based methods inform target prioritisation, the databases and tools available for this purpose, and why pathway co-membership should be viewed as a starting point for validation rather than as evidence of causal involvement.

CRISPR

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus

The clinical promise of mass spectrometry-based single-cell proteomics: from bedside to bench.

INTRODUCTION: Single-cell proteomics (SCP) is entering into a transformative phase, moving beyond technically demanding benchmarking studies toward robust and reproducible workflows capable of quantifying thousands of proteins per cell. These advances highlight SCP's potential to address clinically relevant questions by resolving cellular and pathological heterogeneity that remains obscured in bulk proteomics. AREAS COVERED: This review discusses current advances, challenges, and clinical applications of SCP based on literature identified through searches in major scientific databases. Many clinically relevant samples remain underexplored in SCP studies, in part because their application requires careful evaluation of pre-analytical variables that can strongly influence proteomic readouts. Current SCP methodologies vary according to sample type, experimental conditions, and available resources. Compared with single-cell RNA sequencing, SCP remains limited in cellular throughput, making it challenging to define optimal sample sizes and to reliably detect both abundant and rare cell populations. These limitations also make dataset integration difficult, as reduced cellular coverage and sampling depth increase data sparsity. Moreover, implementing quality control strategies across sequential SCP experiments is essential to ensure data robustness, comparability, and accurate biological interpretation. EXPERT OPINION: Applying SCP to clinical samples advances our understanding of biological complexity and holds potential to drive progress in translational and precision medicine.

Humans

The proteogenomic landscape of the human kidney and implications for cardio-kidney-metabolic health.

Nearly one-third of the global population is affected by cardio-kidney-metabolic (CKM) diseases; however, the molecular mechanisms underlying CKM diseases are poorly understood. Here we show that tissue proteomics provide critical insights not captured by tissue gene expression or blood proteomics information by performing whole-genome and RNA sequencing and proteomics analysis of human kidney samples (n = 337), and we generated a publicly available database. Via Bayesian co-localization and Mendelian randomization analyses of kidney protein quantitative trait loci and 36 CKM genome-wide association studies, we prioritized 89 proteins for CKM traits. We prioritized relationships that could underlie the interconnectedness of CKM traits and discovered multiple and targetable mechanisms for CKM diseases, including the potential role of kidney angiopoietin-like protein 3 (ANGPTL3) in serum lipid levels and kidney function as well as the role of charged multivesicular body protein 1A in kidney function and hypertension. Notably, we identify pathways with confluence of evidence from genetic loci, tissue gene expression and protein levels for CKM traits. In summary, our large-scale kidney proteomics study uncovers proteins and targetable mechanisms prioritized for CKM diseases.

Humans

Mass Spectrometry-Based Profiling of Personalized Immunopeptidomes in Thai Renal Cell Carcinoma.

This study profiles the personalized immunopeptidomes of 13 Thai patients with renal cell carcinoma (RCC), addressing a critical knowledge gap in Southeast Asian populations characterized by distinct HLA allele distributions. We combined whole-exome sequencing (WES)-based personalized proteome construction with liquid chromatography-tandem mass spectrometry (LC-MS/MS), using both database-driven searches and de novo peptide sequencing. HLA typing identified several class I allotypes that are underrepresented in publicly available immunopeptidome resources, including seven alleles not previously represented in the databases examined; HLA-A*11:01 was the most frequent allele in this cohort. Database-based analysis identified a single tumor-specific neoantigen derived from a mutant JADE2 peptide in the patient with the highest tumor mutational burden, which was validated by a mutant-specific ELISPOT response. In contrast, de novo sequencing revealed numerous noncanonical peptides, a subset of which were supported by proteogenomic validation using PepQuery and detected exclusively in cancer proteomes but not in normal tissue data sets, indicating their potential as tumor-associated antigen candidates. Together, these results establish an integrated and scalable framework for identifying HLA-presented tumor-derived peptides and provide a foundational immunopeptidome resource to support personalized cancer immunotherapy development in Southeast Asia.

Humans

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein

Polyethylene transformation by a psychrotolerant Rhodococcus strain assessed by transcriptomics and 13C-isotope tracing.

Polyethylene is increasingly accumulating in nature, including remote places like the Arctic. While abiotic processes fragment polyethylene in situ, biotic transformation by microorganisms is assumed to occur. However, the enzymes and pathways involved remain poorly characterized. In this study, we used an in-house biobank from cold environments to screen for potential bacteria capable of degrading polyethylene by screening the strains in silico using the database PlasticDB and in vivo using a fluorescence-based assay. Using transcriptomic and proteomic analyses to identify genes in promising candidate strains that encode extracellular enzymes potentially capable of degrading PE, we selected a Rhodococcus erythropolis strain and two of its enzymes: a hypothetical protein (Hypr1) and a lipase family protein (Lip2). Expressing the candidate genes heterologously in Escherichia coli resulted in positive results in the fluorescence-based assay for polyethylene transformation. Applying 13C-labelled polyethylene for assessing and estimating polyethylene transformation and carbon assimilation, we found that R. erythropolis and both untransformed and recombinant E. coli extracellularly transformed the initially added polyethylene after 70 days. In addition, untransformed E. coli and R. erythropolis converted small, but significant amounts of polyethylene-derived carbon to carbon dioxide. The 13C-label was also traced into the bacterial biomass of R. erythropolis. Overall, our results provide evidence for biotic transformation of untreated polyethylene and suggests a hypothetical protein and a lipase family protein as two novel enzyme candidates associated with PE transformation.

Rhodococcus

NovoBoard: A Comprehensive Framework for Evaluating the False Discovery Rate and Accuracy of De Novo Peptide Sequencing.

De novo peptide sequencing is one of the most fundamental research areas in mass spectrometry-based proteomics. Many methods have often been evaluated using a couple of simple metrics that do not fully reflect their overall performance. Moreover, there has not been an established method to estimate the false discovery rate (FDR) of de novo peptide-spectrum matches. Here we propose NovoBoard, a comprehensive framework to evaluate the performance of de novo peptide-sequencing methods. The framework consists of diverse benchmark datasets (including tryptic, nontryptic, immunopeptidomics, and different species) and a standard set of accuracy metrics to evaluate the fragment ions, amino acids, and peptides of the de novo results. More importantly, a new approach is designed to evaluate de novo peptide-sequencing methods on target-decoy spectra and to estimate and validate their FDRs. Our FDR estimation provides valuable information to assess the reliability of new peptides identified by de novo sequencing tools, especially when no ground-truth information is available to evaluate their accuracy. The FDR estimation can also be used to evaluate the capability of de novo peptide sequencing tools to distinguish between de novo peptide-spectrum matches and random matches. Our results thoroughly reveal the strengths and weaknesses of different de novo peptide-sequencing methods and how their performances depend on specific applications and the types of data.

Peptides

Omics in hereditary optic neuropathies: A systematic review of clinical studies with an integrated point of view.

Hereditary optic neuropathies are characterized by bilateral visual loss due to the degeneration of retinal ganglion cells, resulting in optic nerve degeneration and atrophy. Although the genetic origin of the main isolated and syndromic hereditary optic neuropathies has been characterized, the clinical phenotypes exhibit significant and poorly understood variability in both penetrance and expressivity. Additionally, the genetic and environmental factors that influence the onset of these optic neuropathies remain poorly understood, with limited biomarkers to predict disease progression or as readouts for therapeutic trials. Data-driven omics strategies allow deep phenotyping to improve our understanding of pathophysiological mechanisms and to search for new biomarkers and therapeutic targets. We explore whether the omics strategies applied to patients with hereditary optic neuropathies have provided such new insights. MEDLINE, Web of Science and EMBASE databases were screened for studies with terms relating to hereditary optic neuropathies, transcriptomics, epigenomics, proteomics, metabolomics and lipidomics in clinical studies exploring patients' samples. Out of 1244 references identified, 22 articles were included after double-masked data curation. These articles focused only on the 3 main forms of hereditary optic neuropathies, namely, OPA1-related dominant optic atrophy (n = 4), Leber hereditary optic neuropathy (n = 13), and Wolfram syndrome (n = 5). While the methodological designs and results of these studies were highly heterogeneous, they revealed molecular alterations that we have attempted to discuss at the integrated multi-omics level. This data integration highlighted several common pathophysiological mechanisms such as energetic impairment, endoplasmic reticulum stress, proteotoxic and oxidative stresses, lipid remodeling and altered amino acid and purine metabolisms, while suggesting potential new biomarkers and therapeutic targets. These findings underscore the potential of integrated multi-omics approaches to deepen our understanding of the phenotypic complexity of hereditary optic neuropathies and to support the development of innovative diagnostic and therapeutic strategies.

Humans

Functional Analysis of MS-Based Proteomics Data: From Protein Groups to Networks.

Mass spectrometry-based proteomics allows the quantification of thousands of proteins, protein variants, and their modifications, in many biological samples. These are derived from the measurement of peptide relative quantities, and it is not always possible to distinguish proteins with similar sequences due to the absence of protein-specific peptides. In such cases, peptide signals are reported in protein groups that can correspond to several genes. Here, we show that multi-gene protein groups have a limited impact on GO-term enrichment, but selecting only one gene per group affects network analysis. We thus present the Cytoscape app Proteo Visualizer (https://apps.cytoscape.org/apps/ProteoVisualizer) that is designed for retrieving protein interaction networks from STRING using protein groups as input and thus allows visualization and network analysis of bottom-up MS-based proteomics data sets.

Proteomics

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning

IgStrand: A universal residue numbering scheme for the immunoglobulin-fold (Ig-fold) to study Ig-proteomes and Ig-interactomes.

The Immunoglobulin fold (Ig-fold) is found in proteins from all domains of life and represents the most populous fold in the human genome, with current estimates ranging from 2 to 3% of protein coding regions. That proportion is much higher in the surfaceome where Ig and Ig-like domains orchestrate cell-cell recognition, adhesion and signaling. The ability of Ig-domains to reliably fold and self-assemble through highly specific interfaces represents a remarkable property of these domains, making them key elements of molecular interaction systems: the immune system, the nervous system, the vascular system and the muscular system. We define a universal residue numbering scheme, common to all domains sharing the Ig-fold in order to study the wide spectrum of Ig-domain variants constituting the Ig-proteome and Ig-Ig interactomes at the heart of these systems. The "IgStrand numbering scheme" enables the identification of Ig structural proteomes and interactomes in and between any species, and comparative structural, functional, and evolutionary analyses. We review how Ig-domains are classified today as topological and structural variants and highlight the "Ig-fold irreducible structural signature" shared by all of them. The IgStrand numbering scheme lays the foundation for the systematic annotation of structural proteomes by detecting and accurately labeling Ig-, Ig-like and Ig-extended domains in proteins, which are poorly annotated in current databases and opens the door to accurate machine learning. Importantly, it sheds light on the robust Ig protein folding algorithm used by nature to form beta sandwich supersecondary structures. The numbering scheme powers an algorithm implemented in the interactive structural analysis software iCn3D to systematically recognize Ig-domains, annotate them and perform detailed analyses comparing any domain sharing the Ig-fold in sequence, topology and structure, regardless of their diverse topologies or origin. The scheme provides a robust fold detection and labeling mechanism that reveals unsuspected structural homologies among protein structures beyond currently identified Ig- and Ig-like domain variants. Indeed, multiple folds classified independently contain a common structural signature, in particular jelly-rolls. Examples of folds that harbor an "Ig-extended" architecture are given. Applications in protein engineering around the Ig-architecture are straightforward based on the universal numbering.

Humans

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational

Influence of nicotine on protein expression around hydrophilic osseointegrated implants: A proteomic study in male rats.

OBJECTIVE: To ensure the success of dental implant treatment, various factors must be considered, including osseointegration and systemic conditions. There is evidence in the literature that smokers may exhibit alterations in tissue healing, which can compromise the success of implant rehabilitation. Therefore, this study aimed to investigate the influence of nicotine on the protein profile of bone tissue around hydrophilic implants during the osseointegration process in rats. DESIGN: Bone tissue samples from the control and nicotine groups (n = 3 per group) were subjected to protein extraction, mass spectrometry, and bioinformatic analyses. Protein identification was performed using Proteome Discoverer 2.1 software and the SEQUEST algorithm, and the protein data were compared with those of a protein database of Rattus norvegicus obtained from UniProt. RESULTS: A total of 740 proteins were detected in both the control group and the nicotine-exposed group. Among them, the proteins biglycan, periostin and histone H4 were highlighted because of their higher abundance in the healthy implant group, while they were reduced in the nicotine-exposed group. CONCLUSIONS: Nicotine has the potential to alter the protein profile of bone tissue around hydrophilic implants during osseointegration, which may impair tissue remodeling and healing.

Animals