Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Computational methods for comparison of large genomic and proteomic datasets reveal protein markers of metastatic cancer.

Large-scale genomic and proteomic analysis has provided a wealth of information on biologically relevant systems, and the ability to analyze this information is crucial to uncovering important biological relationships. However, it has proven difficult to compare large datasets from different sources due to different gene and protein identifiers assigned by individual laboratories and database systems. Here, we describe the design of a fully automated blast program (BlastPro) that facilitates rapid comparison of large protein-protein, nucleotide--nucleotide, or nucleotide--protein datasets from numerous, independent studies. Using this system, we compared several published genomic and proteomic databases for proteins that are upregulated in highly motile, metastatic tumor cells. Analysis of five independent studies comprised of greater than 1 x 10(6) genomic sequences and greater than 1,000 proteins revealed that the cytoskeletal-associated protein alpha-actinin is increased at both the mRNA and protein level in metastatic breast, prostate, and skin cancer cells. Interestingly, spatial analysis of alpha-actinin expression revealed that it is amplified 8-fold in the leading pseudopodium compared to the cell body compartment of migrating cells. These findings indicate that amplification of alpha-actinin and its localization to the leading pseudopodium are potential biomarkers of cancer progression to a more metastatic phenotype. Together, our results demonstrate that the BlastPro system can be used to compare large genomic and proteomic datasets to reveal important biological relationships including those associated with cancer progression.

Actinin↗

InSilicoSpectro: an open-source proteomics library.

We present a new proteomics open-source project, InSilicoSpectro, aimed at implementing recurrent computations that are necessary for proteomics data analysis. Illustrative examples are mass list file format conversions, protein sequence digestion, theoretical peptide and fragment mass computations, graphical display, matching with experimental data, isoelectric point estimation, and peptide retention time prediction. The project library is written in Perl, a widely used scripting language in bioinformatics, and it offers a unique framework of integrated objects to implement complex proteomics data analyses. For instance, only a few lines of code are required to digest a protein with fixed and variable modifications, label peptides with 18O, compute the fragmentation spectra and display their match with experimental spectra. We believe that InSilicoSpectro will be of great help to bioinformaticians, without detailed knowledge of proteomics specifics, and to mass spectrometrists with computer programming interest as well.

Amino Acid Sequence↗

Fragmentation pathways of protonated peptides.

The fragmentation pathways of protonated peptides are reviewed in the present paper paying special attention to classification of the known fragmentation channels into a simple hierarchy defined according to the chemistry involved. It is shown that the 'mobile proton' model of peptide fragmentation can be used to understand the MS/MS spectra of protonated peptides only in a qualitative manner rationalizing differences observed for low-energy collision induced dissociation of peptide ions having or lacking a mobile proton. To overcome this limitation, a deeper understanding of the dissociation chemistry of protonated peptides is needed. To this end use of the 'pathways in competition' (PIC) model that involves a detailed energetic and kinetic characterization of the major peptide fragmentation pathways (PFPs) is proposed. The known PFPs are described in detail including all the pre-dissociation, dissociation, and post-dissociation events. It is our hope that studies to further extend PIC will lead to semi-quantative understanding of the MS/MS spectra of protonated peptides which could be used to develop refined bioinformatics algorithms for MS/MS based proteomics. Experimental and computational data on the fragmentation of protonated peptides are reevaluated from the point of view of the PIC model considering the mechanism, energetics, and kinetics of the major PFPs. Evidence proving semi-quantitative predictability of some of the ion intensity relationships (IIRs) of the MS/MS spectra of protonated peptides is presented.

Amino Acid Sequence↗

From proteins to systems: the British Society for Proteomics Research (BSPR) meeting 2005.

This report describes the highlights of the second scientific meeting of the British Society for Proteome Research (BSPR), jointly organised with the European Bioinformatics Institute (EBI), and held at The Genome Centre, Cambridge UK in July 2005. The theme of the meeting was "From Proteins to Systems" covering many diverse aspects of proteomics, bioinformatics and systems biology.

Computational Biology↗

Genomics to fluxomics and physiomics - pathway engineering.

Developments in microanalytical methods are enabling quantitative measurement of multiple metabolic fluxes and, in conjunction with transcript and proteomic profiling, are revolutionizing the ability of researchers to manipulate metabolism through pathway engineering in a variety of species. We review recent literature on the advances in genomics, proteomics, fluxomics and computational modeling focused on metabolic pathway engineering applications.

Bacteria↗

Engineering challenges of BioNEMS: the integration of microfluidics, micro- and nanodevices, models and external control for systems biology.

Systems biology, i.e. quantitative, postgenomic, postproteomic, dynamic, multiscale physiology, addresses in an integrative, quantitative manner the shockwave of genetic and proteomic information using computer models that may eventually have 10(6) dynamic variables with non-linear interactions. Historically, single biological measurements are made over minutes, suggesting the challenge of specifying 10(6) model parameters. Except for fluorescence and micro-electrode recordings, most cellular measurements have inadequate bandwidth to discern the time course of critical intracellular biochemical events. Micro-array expression profiles of thousands of genes cannot determine quantitative dynamic cellular signalling and metabolic variables. Major gaps must be bridged between the computational vision and experimental reality. The analysis of cellular signalling dynamics and control requires, first, micro- and nano-instruments that measure simultaneously multiple extracellular and intracellular variables with sufficient bandwidth; secondly, the ability to open existing internal control and signalling loops; thirdly, external BioMEMS micro-actuators that provide high bandwidth feedback and externally addressable intracellular nano-actuators; and, fourthly, real-time, closed-loop, single-cell control algorithms. The unravelling of the nested and coupled nature of cellular control loops requires simultaneous recording of multiple single-cell signatures. Externally controlled nano-actuators, needed to effect changes in the biochemical, mechanical and electrical environment both outside and inside the cell, will provide a major impetus for nanoscience.

Animals↗

Crystal structure of yeast YHR049W/FSH1, a member of the serine hydrolase family.

Yhr049w/FSH1 was recently identified in a combined computational and experimental proteomics analysis for the detection of active serine hydrolases in yeast. This analysis suggested that FSH1 might be a serine-type hydrolase belonging to the broad functional alphabeta-hydrolase superfamily. In order to get insight into the molecular function of this gene, it was targeted in our yeast structural genomics project. The crystal structure of the protein confirms that it contains a Ser/His/Asp catalytic triad that is part of a minimal alpha/beta-hydrolase fold. The architecture of the putative active site and analogies with other protein structures suggest that FSH1 may be an esterase. This finding was further strengthened by the unexpected presence of a compound covalently bound to the catalytic serine in the active site. Apparently, the enzyme was trapped with a reactive compound during the purification process.

Amino Acid Sequence↗

Genome annotating proteomics pipelines: available tools.

Proteomics based on tandem mass spectrometry is a powerful tool for identifying novel biomarkers and drug targets. Previously, a major bottleneck in high-throughput proteomics has been that the computational techniques needed to reliably identify proteins from proteomic data lagged behind the ability to collect the immense quantity of data generated. This is no longer the case, as fully automated pipelines for peptide and protein identification exist, and these are publicly and privately accessible. Such pipelines can automatically and rapidly generate high-confidence protein identifications from large datasets in a searchable format covering multiple experimental runs. However, the main challenge for the community now is to use these resources as they are, by taking full advantage of the pooling of information, so that the next barrier in our understanding of biology may be broken. There are currently two pipelines in the public domain that provide such potential: PeptideAtlas and the Genome Annotating Proteomic Pipeline. This review will introduce their features in the context of high-throughput proteomics, and provide indicative results as to their usefulness and usability through a side-by-side comparison of results obtained when processing a set of human plasma samples.

Animals↗

Prediction of liquid chromatographic retention times of peptides generated by protease digestion of the Escherichia coli proteome using artificial neural networks.

We developed a computational method to predict the retention times of peptides in HPLC using artificial neural networks (ANN). We performed stepwise multiple linear regressions and selected for ANN input amino acids that significantly affected the LC retention time. Unlike conventional linear models, the trained ANN accurately predicted the retention time of peptides containing up to 50 amino acid residues. In 834 peptides, there was a strong correlation (R2 = 0.928) between measured and predicted retention times. We demonstrated the utility of our method by the prediction of the retention time of 121,273 peptides resulting from LysC-digestion of the Escherichia coli proteome. Our approach is useful for the proteome-wide characterization of peptides and the identification of unknown peptide peaks obtained in proteome analysis.

Chromatography, High Pressure Liquid↗

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome↗

A computational approach for ordering signal transduction pathway components from genomics and proteomics Data.

BACKGROUND: Signal transduction is one of the most important biological processes by which cells convert an external signal into a response. Novel computational approaches to mapping proteins onto signaling pathways are needed to fully take advantage of the rapid accumulation of genomic and proteomics information. However, despite their importance, research on signaling pathways reconstruction utilizing large-scale genomics and proteomics information has been limited. RESULTS: We have developed an approach for predicting the order of signaling pathway components, assuming all the components on the pathways are known. Our method is built on a score function that integrates protein-protein interaction data and microarray gene expression data. Compared to the individual datasets, either protein interactions or gene transcript abundance measurements, the integrated approach leads to better identification of the order of the pathway components. CONCLUSIONS: As demonstrated in our study on the yeast MAPK signaling pathways, the integration analysis of high-throughput genomics and proteomics data can be a powerful means to infer the order of pathway components, enabling the transformation from molecular data into knowledge of cellular mechanisms.

Computational Biology↗

Community-driven advances in computational mass spectrometry: The perspective of EuBIC-MS members.

Advances in data acquisition, artificial intelligence, and integrative bioinformatics are driving the rapid evolution of computational mass spectrometry, and in turn, transforming modern proteomics, metabolomics, and lipidomics. These developments have greatly increased the scale and complexity of mass spectrometry data, underscoring the importance of evolving accurate, transparent, efficient and reproducible data processing workflows. Addressing these challenges requires collaborative innovation that brings together expertise in software engineering, statistics, and biology. The European Bioinformatics Community for Mass Spectrometry (EuBIC-MS), an initiative of the European Proteomics Association (EuPA), fosters a culture of open, community-driven development through its biennial Developers Meetings and Winter Schools. This commentary summarizes the scientific background and outcomes of the EuBIC-MS Developers Meeting 2025, which took place in Novacella, Italy. Three keynote presentations highlighted major frontiers in the field: deep proteome and phosphoproteome profiling, text mining for protein-protein interaction extraction, and scalable proteomics for AI-driven drug discovery. Seven community-selected hackathons addressed emerging challenges such as single-cell proteomics data analysis, FAIR metadata extraction, deep learning frameworks, R-Python interoperability, and DIA validation. Together, these efforts demonstrate the potential for scientific and technical innovation to arise from open collaboration, and highlight how community-driven initiatives can accelerate progress in computational mass spectrometry. SIGNIFICANCE: Modern proteomics increasingly depends on computational advances to translate complex, high-dimensional data into biological knowledge. The EuBIC-MS Developers Meeting 2025 exemplifies how community-driven collaboration can directly accelerate this process by bringing together experts from bioinformatics, statistics, and experimental proteomics to co-develop open, interoperable, and reproducible analytical tools. By fostering shared software frameworks, transparent benchmarking, and collaborative problem solving, the EuBIC-MS community helps ensure that technological innovation translates into reliable biological insights. This collaborative model strengthens the foundation for quantitative, system-level understanding of proteomes and establishes a sustainable path for integrating artificial intelligence and next-generation data acquisition into routine biological discovery. This commentary shows some current highlights in the field of computational mass spectrometry and community-based approaches undertaken during the most recent Developers Meeting to solve these challenges. The approaches discussed and initiated during the meeting - ranging from deep proteome profiling and phosphosite mapping to text mining, single-cell data analysis, and FAIR metadata extraction - address key bottlenecks that currently limit the biological interpretability and comparability of proteomics data.

Mass Spectrometry↗

Bioinformatics Resources for In Silico Proteome Analysis.

In the growing field of proteomics, tools for the in silico analysis of proteins and even of whole proteomes are of crucial importance to make best use of the accumulating amount of data. To utilise this data for healthcare and drug development, first the characteristics of proteomes of entire species-mainly the human-have to be understood, before secondly differentiation between individuals can be surveyed. Specialised databases about nucleic acid sequences, protein sequences, protein tertiary structure, genome analysis, and proteome analysis represent useful resources for analysis, characterisation, and classification of protein sequences. Different from most proteomics tools focusing on similarity searches, structure analysis and prediction, detection of specific regions, alignments, data mining, 2D PAGE analysis, or protein modelling, respectively, comprehensive databases like the proteome analysis database benefit from the information stored in different databases and make use of different protein analysis tools to provide computational analysis of whole proteomes.

Journal Article↗

Understanding the adaptation of Halobacterium species NRC-1 to its extreme environment through computational analysis of its genome sequence.

The genome of the halophilic archaeon Halobacterium sp. NRC-1 and predicted proteome have been analyzed by computational methods and reveal characteristics relevant to life in an extreme environment distinguished by hypersalinity and high solar radiation: (1) The proteome is highly acidic, with a median pI of 4.9 and mostly lacking basic proteins. This characteristic correlates with high surface negative charge, determined through homology modeling, as the major adaptive mechanism of halophilic proteins to function in nearly saturating salinity. (2) Codon usage displays the expected GC bias in the wobble position and is consistent with a highly acidic proteome. (3) Distinct genomic domains of NRC-1 with bacterial character are apparent by whole proteome BLAST analysis, including two gene clusters coding for a bacterial-type aerobic respiratory chain. This result indicates that the capacity of halophiles for aerobic respiration may have been acquired through lateral gene transfer. (4) Two regions of the large chromosome were found with relatively lower GC composition and overrepresentation of IS elements, similar to the minichromosomes. These IS-element-rich regions of the genome may serve to exchange DNA between the three replicons and promote genome evolution. (5) GC-skew analysis showed evidence for the existence of two replication origins in the large chromosome. This finding and the occurrence of multiple chromosomes indicate a dynamic genome organization with eukaryotic character.

Adaptation, Biological↗

Sub-proteome differential display: single gel comparison by 2D electrophoresis and mass spectrometry.

Two-dimensional (2D) gel electrophoresis and mass spectrometry (MS) have been used in comparative proteomics but inherent problems of the 2D electrophoresis technique lead to difficulties when comparing two samples. We describe a method (sub-proteome differential display) for comparing the proteins from two sources simultaneously. Proteins from one source are mixed with radiolabelled proteins from a second source in a ratio of 100:1. These combined proteomes are fractionated simultaneously using column chromatographic methods, followed by analysis of the pre-fractionated proteomes (designated sub-proteomes) using 2D gel electrophoresis. Silver staining and (35)S autoradiography of a single gel allows precise discrimination between members of each sub-proteome, using commonly available computer software. This is followed by MS identification of individual proteins. We have demonstrated the utility of the technology by identifying the product of a transfected gene and several proteins expressed differentially between two renal carcinoma proteomes. The procedure has the capacity to enrich proteins prior to 2D electrophoresis and provides a simple, inexpensive approach to compare proteomes. The single gel approach eliminates differences that might arise if separate proteome fractionations or 2D gels are employed.

Animals↗

Novel strategies for the early detection and prevention of lung cancer.

Lung cancer is the leading cause of cancer death in the United States. Despite evidence of molecular abnormalities in biological specimens, progress in this disease is hampered by the lack of diagnostic markers useful for clinical practice. The majority of patients with lung cancer are still diagnosed at an advanced stage, when prognosis is poor. This article reviews new strategies being studied for the early detection of lung cancer. These strategies involve new methods of imaging (including low-dose computed tomography [CT] scanning), DNA analysis, and proteomic-based techniques. These strategies have not only improved our understanding of lung cancer but show promise in offering better survival to patients with this deadly disease. Of paramount importance in the search for methods of early detection is the need for the identification of the ideal population to screen, a multidisciplinary approach, and validation of promising techniques.

Anticarcinogenic Agents↗

narrowPASEF: A Sample-Aware diaPASEF Method Optimization Strategy Improving Differential Proteomics Performance on Low-Abundance Proteins.

Recent instrumental and computational innovations in mass-spectrometry-based proteomics offer new promise in biomarker discovery, thanks to unprecedented proteome coverage and depth. Data-independent acquisition (DIA) methods are very promising in this context as they allow improved proteome coverage, reduced missing value rates, and enhanced quantification precision. However, DIA methods also suffer from their own challenges, such as increased data complexity, cycle times, and background noise. In this work, we propose a sample-aware diaPASEF method optimization strategy for a timsTOF platform. Thorough method optimizations have first been conducted on standard HeLa lysates. Then, a ground-truth calibrated sample series, consisting of a range of UPS amounts spiked into a complex Arabidopsis background, was used to mimic differential analyses under controlled conditions. These benchmark experiments demonstrate clear benefits of using narrowPASEF for differential protein discovery. Finally, our strategy was applied to real use case biological samples to conduct a differential analysis of purified mouse astrocyte cells across two different conditions. narrowPASEF improved the proteome depth by 13%, considering proteins quantified with a coefficient of variation (CV) of <20%, and led to a 68% (435 vs 729) increase in differentially expressed proteins. These results provide an opportunity for a more precise and comprehensive analysis of the biological functions of biomarkers, offering a more profound understanding of the disease mechanisms. The benefits of our sample-aware narrowPASEF strategy demonstrated the most substantial impact on low-abundance proteins. Overall, these results show promise for more valuable and robust biomarker discoveries in the future.

Proteomics↗