Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Multiple approaches to data-mining of proteomic data based on statistical and pattern classification methods.

The data-mining challenge presented is composed of two fundamental problems. Problem one is the separation of forty-one subjects into two classifications based on the data produced by the mass spectrometry of protein samples from each subject. Problem two is to find the specific differences between protein expression data of two sets of subjects. In each problem, one group of subjects has a disease, while the other group is nondiseased. Each problem was approached with the intent to introduce a new and potentially useful tool to analyze protein expression from mass spectrometry data. A variety of methodologies, both conventional and nonconventional were used in the analysis of these problems. The results presented show both overlap and discrepancies. What is important is the breadth of the techniques and the future direction this analysis will create.

Artificial Intelligence↗

Predotar: A tool for rapidly screening proteomes for N-terminal targeting sequences.

Probably more than 25% of the proteins encoded by the nuclear genomes of multicellular eukaryotes are targeted to membrane-bound compartments by N-terminal targeting signals. The major signals are those for the endoplasmic reticulum, the mitochondria, and in plants, plastids. The most abundant of these targeted proteins are well-known and well-studied, but a large proportion remain unknown, including most of those involved in regulation of organellar gene expression or regulation of biochemical pathways. The discovery and characterization of these proteins by biochemical means will be long and difficult. An alternative method is to identify candidate organellar proteins via their characteristic N-terminal targeting sequences. We have developed a neural network-based approach (Predotar--Prediction of Organelle Targeting sequences) for identifying genes encoding these proteins amongst eukaryotic genome sequences. The power of this approach for identifying and annotating novel gene families has been illustrated by the discovery of the pentatricopeptide repeat family.

Arabidopsis Proteins↗

Extractor for ESI quadrupole TOF tandem MS data enabled for high throughput batch processing.

BACKGROUND: Mass spectrometry based proteomics result in huge amounts of data that has to be processed in real time in order to efficiently feed identification algorithms and to easily integrate in automated environments. We present wiff2dta, a tool created to convert MS/MS data obtained using Applied Biosystem's QStar and QTrap 2000 and 4000 series. RESULTS: Comparing the performance of wiff2dta with the standard tools, we find wiff2dta being the fastest solution for extracting spectrum data from ABIs raw file format. wiff2dta is at least 10% faster than the standard tools. It is also capable of batch processing and can be easily integrated in high throughput environments. The program is freely available via http://www.protein-ms.de, http://sourceforge.net/projects/protms/ and is also available from Applied Biosystems. CONCLUSIONS: wiff2dta offers the possibility to run as stand-alone application or within a batch process as command-line tool integrated in automation and high-throughput environments. It is more efficient than the state-of-the-art tools provided.

Automation↗

Primary skin fibroblasts as human model system for proteome analysis.

Elucidation of cellular processes and their changes at the level of protein expression and post-translational modification patterns may allow identification of novel proteins and thereby mechanisms involved in the pathogenesis of multigenic diseases. The aim of this study was to test cultured, nontransformed primary fibroblasts derived from human skin biopsies as a suitable model system for proteome analysis. Therefore soluble protein fractions were separated on several overlapping ultrazoom gels covering the pH range from 3.5-9. Correlation analysis of gel-pairs revealed a highly reproducible protein expression pattern within (intra-assay) and between (inter-assay) independent experiments of a single fibroblast cell line (intra-cell line comparison). Spot intensity variations were less than a factor of two for more than 80% of identical spots. In addition, inter-cell line comparison exhibits no significant variations in spot intensities. To achieve further improvements in reproducibility we generated master gels for each pH range by combining averaged spot information derived from two different cell lines each analysed by two independent experiments using the raw master gel algorithm of the Z3 image analysis software. The resulting reference images of primary human fibroblasts provided a basis for investigating regulation by extracellular stimuli and drugs as well as their alterations in patients with different diseases.

Cells, Cultured↗

Changes in the serum proteome associated with the development of hepatocellular carcinoma in hepatitis C-related cirrhosis.

Early diagnosis of hepatocellular carcinoma (HCC) is the key to the delivery of effective therapies. The conventional serological diagnostic test, estimation of serum alpha-fetoprotein (AFP) lacks both sensitivity and specificity as a screening tool and improved tests are needed to complement ultrasound scanning, the major modality for surveillance of groups at high risk of HCC. We have analysed the serum proteome of 182 patients with hepatitis C-induced liver cirrhosis (77 with HCC) by surface-enhanced laser desorption/ionisation time-of-flight mass spectrometry (SELDI). The patients were split into a training set (84 non-HCC, 60 HCC) and a 'blind' test set (21 non-HCC, 17 HCC). Neural networks developed on the training set were able to classify the blind test set with 94% sensitivity (95% CI 73-99%) and 86% specificity (95% CI 65-95%). Two of the SELDI peaks (23/23.5 kDa) were elevated by an average of 50% in the serum of HCC patients (P<0.001) and were identified as kappa and lambda immunoglobulin light chains. This approach may permit identification of several individual proteins, which, in combination, may offer a novel way to diagnose HCC.

Amino Acid Sequence↗

Rapid prefractionation of complex protein lysates with centrifugal membrane adsorber units improves the resolving power of 2D-PAGE-based proteome analysis.

BACKGROUND: Two-dimensional gel electrophoresis (2D-PAGE) has proven over the years to be a reliable and efficient method for separation of hundreds of proteins based on charge and mass. Nevertheless, the complexity of even the simplest proteomes limits the resolving power of 2D-PAGE. This limitation can be partially alleviated by sample prefractionation using a variety of techniques. RESULTS: Here, we have used Vivapure Ion Exchange centrifugal adsorber units to rapidly prefractionate total fission yeast protein lysate based on protein charge. Three fractions were prepared by stepwise elution with increasing sodium chloride concentrations. Each of the fractions, as well as the total lysate, were analyzed by 2D-PAGE. This simple prefractionation procedure considerably increased the resolving power of 2D-PAGE. Whereas 308 spots could be detected by analysing total protein lysate, 910 spots were observed upon prefractionation. Thorough gel image analysis demonstrated that prefractionation visualizes an additional set of 458 unique fission yeast proteins not detected in whole cell lysate. CONCLUSIONS: Prefractionation with Vivapure Q spin columns proved to be a simple, fast, reproducible, and cost-effective means of increasing the resolving power of 2D-PAGE using standard laboratory equipment.

Centrifugation↗

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational↗

ECLIPSE: exploring the dark proteome of ESKAPE pathogens through the sequence similarity network of the Protein Universe Atlas.

MOTIVATION: The accelerating crisis of antimicrobial resistance among the critical so-called ESKAPE pathogens demands the urgent identification of novel molecular targets. However, a substantial fraction of ESKAPE proteomes remains functionally uncharacterized, with many genes annotated as encoding hypothetical proteins. These protein sequences often lack significant similarity to known protein families when conventional homology-based annotation methods are used and thus remain "dark". This limits our ability to explore their roles in pathogenicity, and it is thus crucial to bridge this substantial gap in pathogen biology by developing new strategies to illuminate these "dark" regions of the ESKAPE pan-proteome. RESULTS: We introduce ECLIPSE (ESKAPE Connectome Linkage and Inference for Proteome Sequence Exploration), a network-based computational framework that systematically identifies and prioritizes functionally dark protein families in ESKAPE pan-proteomes. ECLIPSE embeds target ESKAPE pathogen proteomes within the global sequence similarity network of the Protein Universe Atlas. It detects connected components composed entirely of unannotated proteins, called the "dark proteome." As a case study, we applied ECLIPSE to a pan-proteome of 3&#x2006;460&#x2006;657 protein sequences from 635 strains of Pseudomonas aeruginosa (PA). ECLIPSE identified 120&#x2006;985 proteins (4%) residing in completely dark connected components. Furthermore, we have performed a taxonomic diversity analysis using normalized Shannon indices to characterize each dark component by its enrichment in ESKAPE pathogens. The analysis utilized the evenness (E) value (see Methods 2.1), which distinguishes Pseudomonas-specific (target-specific) from ESKAPE-enriched dark components. We then developed the Dark Proteome Prioritization Score (DPPS), a composite multidimensional scoring framework (see Methods 2.5). It ranks these dark components by biological relevance across four orthogonal axes: (i) functional darkness, (ii) P. aeruginosa proportion in the Atlas, (iii) AMR-clade taxonomic restriction, and (iv) conservation across the 635 P. aeruginosa strains. This framework outputs a robust four-tier scoring system; the prioritized Tier I components were validated by weight sensitivity analysis and remained stable across 500 Monte Carlo weight perturbations. Structural characterization of one of the top-ranked ESKAPE-enriched dark components revealed that it belongs to the beta-barrel fold DUF1302 (PF06980) family, for which no experimentally solved three-dimensional structure exists in the PDB. The genomic context analysis indicates that it is co-localized with a LuxR-type transcriptional regulator. Collectively, ECLIPSE identifies evolutionarily conserved, structurally defined, and functionally dark proteins enriched across ESKAPE pathogens; these dark proteins can further be utilized as alternative antimicrobial targets for experimental characterization. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available for free at: Github: https://github.com/surabhilata/ECLIPSE.git, Zenodo: DOI: 10.5281/zenodo.21064323.

Proteome↗

Comparison of protein expression profiles between monolayer and spheroid cell culture of HT-29 cells revealed fragmentation of CK18 in three-dimensional cell culture.

The use of three-dimensional cell culture models, so-called multicellular tumor spheroids, is a special approach in experimental cancer research, because spheroids are similar to in vivo tumors in structural as well as functional sense. Cells grown in spheroids exhibit alterations of cell cycle regulation, induction of apoptosis and differentiation and can acquire multidrug resistance. In this study we investigated the protein expression in human colorectal cancer cells grown in monolayer and in spheroid cultures using proteomics. Evaluation by computer-assisted image analysis revealed overexpression of three cytokeratin 18 fragments that were generated in vivo. Cytokeratin 18 has previously been described as a target for caspase-mediated cleavage during apoptosis and our results indicate that apoptosis may take place in spheroids. Other proteins upregulated in spheroids include calreticulin precursor, a rho GDP dissociation inhibitor variant, several cytokeratins and peroxiredoxin 4. Some of these proteins have already been linked to chemoresistance and apoptotic phenomena.

Electrophoresis, Gel, Two-Dimensional↗

Large-scale prediction of disulphide bridges using kernel methods, two-dimensional recursive neural networks, and weighted graph matching.

The formation of disulphide bridges between cysteines plays an important role in protein folding, structure, function, and evolution. Here, we develop new methods for predicting disulphide bridges in proteins. We first build a large curated data set of proteins containing disulphide bridges to extract relevant statistics. We then use kernel methods to predict whether a given protein chain contains intrachain disulphide bridges or not, and recursive neural networks to predict the bonding probabilities of each pair of cysteines in the chain. These probabilities in turn lead to an accurate estimation of the total number of disulphide bridges and to a weighted graph matching problem that can be addressed efficiently to infer the global disulphide bridge connectivity pattern. This approach can be applied both in situations where the bonded state of each cysteine is known, or in ab initio mode where the state is unknown. Furthermore, it can easily cope with chains containing an arbitrary number of disulphide bridges, overcoming one of the major limitations of previous approaches. It can classify individual cysteine residues as bonded or nonbonded with 87% specificity and 89% sensitivity. The estimate for the total number of bridges in each chain is correct 71% of the times, and within one from the true value over 94% of the times. The prediction of the overall disulphide connectivity pattern is exact in about 51% of the chains. In addition to using profiles in the input to leverage evolutionary information, including true (but not predicted) secondary structure and solvent accessibility information yields small but noticeable improvements. Finally, once the system is trained, predictions can be computed rapidly on a proteomic or protein-engineering scale. The disulphide bridge prediction server (DIpro), software, and datasets are available through www.igb.uci.edu/servers/psss.html.

Amino Acid Sequence↗

From biological databases to platforms for biomedical discovery.

The use of high-throughput DNA sequencing and proteomic methods has led to an unprecedented increase in the amount of genomic and proteomic data. Application of computing technologies and development of computational tools to analyze and present these data has not kept pace with the accumulation of information. Here, we discuss the use of different database systems to store biological information and mention some of the key emerging computing technologies that are likely to have a key role in the future of bioinformatics.

Algorithms↗

The model organism as a system: integrating 'omics' data sets.

Various technologies can be used to produce genome-scale, or 'omics', data sets that provide systems-level measurements for virtually all types of cellular components in a model organism. These data yield unprecedented views of the cellular inner workings. However, this abundance of information also presents many hurdles, the main one being the extraction of discernable biological meaning from multiple omics data sets. Nevertheless, researchers are rising to the challenge by using omics data integration to address fundamental biological questions that would increase our understanding of systems as a whole.

Animals↗

Protein image alignment via piecewise affine transformations.

We present a new approach for aligning families of 2D gels. Instead of choosing one of the gels as reference and performing a pairwise alignment, we construct an ideal gel that is representative of the entire family and obtain a set of piecewise affine transformations that optimally align each gel of the family to the ideal gel. The coefficients defining the transformations as well as the ideal landmarks are obtained as the solution of a large-scale quadratic programming problem that can be solved efficiently by interior-point methods.

Animals↗

Inference of differential kinase interaction networks with KINference.

MOTIVATION: Differential kinase interaction networks (DKINs) are networks containing kinase-substrate links that are differentially active between two conditions. Existing methods are either able to predict condition-agnostic kinase-substrate links or condition-specific differential kinase activity, but do not provide differential kinase-substrate links. Moreover, existing methods for predicting kinase-substrate links usually rely on curated biochemical knowledge. Thus, there is a lack of data-driven DKIN inference methods that are also applicable when prior knowledge is scarce. RESULTS: To address this need, we present KINference. KINference combines computation of a baseline KIN representing the space of all possible kinase-substrate links with filters applied to nodes and edges to identify differentially active subnetworks that are relevant in the context of a specific phosphoproteomics dataset. For the node filters, we rely on functional relevance and differential phosphorylation scores; for the edge filters, we make use of prize-collecting Steiner trees and correlations between phosphorylation sites of kinases and their target proteins. Tests on two phosphoproteomics datasets (kinase inhibition in breast cancer cells, SARS-CoV-2 infection in Calu-3 cells) show that the proposed filters produce significant results in terms of overlap with known interactions between kinases and phosphorylation sites. Furthermore, a case study on the SARS-CoV-2 infection data, suggests a potential host pathway linked to virus replication, showcasing the process of hypothesis generation utilizing DKINs computed by KINference. AVAILABILITY AND IMPLEMENTATION: KINference is available as an R package at https://github.com/bionetslab/KINference and https://doi.org/10.5281/zenodo.15411150. Scripts to reproduce the results are available at https://github.com/bionetslab/KINference-Evaluation-Scripts and https://doi.org/10.5281/zenodo.15424599.

Humans↗

PLSKO: a robust knockoff generator to control false discovery rate in omics variable selection.

MOTIVATION: Integrating the knockoff framework with any variable-selection method delivers stringent false discovery rate (FDR) control without recourse to p-values, offering a powerful alternative for differential expression analysis of high-throughput omics datasets. However, existing knockoff generators rely on restrictive modelling assumptions or coarse approximations that often inflate the FDR when applied to real-world data. RESULTS: We introduce Partial Least Squares Knockoff (PLSKO), an efficient, assumption-free generator that remains robust across diverse omics platforms. Our extensive simulations show that PLSKO is the only method to maintain FDR control with sufficient power in complex non-linear settings. Our semi-simulation studies drawn from RNA-seq, proteomics, metabolomics, and microbiome experiments confirm PLSKO generates valid knockoff variables. In pre-eclampsia multi-omics case studies, we combine PLSKO with Aggregation Knockoff to address the randomness of knockoffs and improve power, and demonstrate the method's ability to recover biologically meaningful features. AVAILABILITY AND IMPLEMENTATION: Our proposed algorithm is available on Github (https://github.com/guannan-yang/PLSKO) and Zenodo (https://doi.org/10.5281/zenodo.16879594).

Algorithms↗

Differential expression of plasma proteins and pathway enrichments in pediatric diabetic ketoacidosis.

BACKGROUND: In children with type 1 diabetes (T1D), diabetic ketoacidosis (DKA) triggers a significant inflammatory response; however, the specific effector proteins and signaling pathways involved remain largely unexplored. This pediatric case-control study utilized plasma proteomics to explore protein alterations associated with severe DKA and to identify signaling pathways that associate with clinical variables. METHODS: We conducted a proteome analysis of plasma samples from 17 matched pairs of pediatric patients with T1D; one cohort with severe DKA and another with insulin-controlled diabetes. Proximity extension assays were used to quantify 3072 plasma proteins. Data analysis was performed using multivariate statistics, machine learning, and bioinformatics. RESULTS: This study identified 214 differentially expressed proteins (162 upregulated, 52 downregulated; adj P&#x2009;<&#x2009;0.05 and a fold change&#x2009;>&#x2009;2), reflecting cellular dysfunction and metabolic stress in severe DKA. We characterized protein expression across various organ systems and cell types, with notable alterations observed in white blood cells. Elevated inflammatory pathways suggest an enhanced inflammatory response, which may contribute to the complications of severe DKA. Additionally, upregulated pathways related to hormone signaling and nitrogen metabolism were identified, consistent with increased hormone release and associated metabolic processes, such as glycogenolysis and lipolysis. Changes in lipid and fatty acid metabolism were also observed, aligning with the lipolysis and ketosis characteristic of severe DKA. Finally, several signaling pathways were associated with clinical biochemical&#xa0;variables. CONCLUSIONS: Our findings highlight differentially expressed plasma proteins and enriched signaling pathways that were associated with clinical features, offering insights into the pathophysiology of severe DKA.

Humans↗

Fruits of human genome project and private venture, and their impact on life science.

A small knowledge base was created by organizing the Human Genome Project (HGP) and its related issues in "Science" magazines between 1996 and 2000. This base revealed the stunning achievement of HGP and a private venture and its impact on today's biology and life science. In the mid-1990, they encouraged the development of advanced high throughput automated DNA sequencers and the technologies that can analyse all genes at once in a systematic fashion. Using these technologies, they completed the genome sequence of human and various other organisms. These fruits opened the door to comparative genomics, functional genomics, the interdisprinary field between computer and biology, and proteomics. They have caused a shift in biological investigation from studying single genes or proteins to studying all genes or proteins at once, and causing revolutional changes in traditional biology, drug discovery and therapy. They have expanded the range of potential drug targets and have facilitated a shift in drug discovery programs toward rational target-based strategies. They have spawned pharmacogenomics that could give rise to a new generation of highly effective drugs that treat causes, not just symptoms. They should also cause a migration from the traditional medications that are safe and effective for every members of the population to personalized medicine and personalized therapy.

Biological Science Disciplines↗

Requirements of a brain selective estrogen: advances and remaining challenges for developing a NeuroSERM.

Our goal is to develop therapeutic agents that prevent age-associated neurodegenerative disease such as Alzheimer's. To achieve this goal, we are building on extensive knowledge regarding mechanisms of estrogen action in brain and the epidemiological human data indicating that estrogen/hormone therapy reduces the risk of developing Alzheimer's disease when administered at the time of the menopause and continued over several to many years. The mechanisms of estrogen action in neurons provides a systematic mechanistic rationale for determining why estrogen therapy is efficacious for prevention of Alzheimer's disease and why it is not efficacious for long-term treatment of the disease. Our preclinical research plan is a hybrid of both discovery and translational research to develop a brain selective estrogen receptor modulator (SERM). We have termed such molecules NeuroSERMs to denote their preferential selectivity for activating estrogen mechanisms in brain. Our strategy to develop NeuroSERMs is threefold: (1) determine the target of estrogen action in brain, specifically the estrogen receptor in hippocampal and cortical neurons required for the neurotrophic and neuroprotective actions of estrogen; (2) develop NeuroSERM candidate molecules using three in silico discovery and design strategies and (3) determine the neurotrophic and neuroprotective efficacy of candidate molecules using neuronal responses predictive of clinical efficacy. Using an academic translational research model, a team of scientists with expertise in molecular biology, computational chemistry, synthetic chemistry, proteomics, neurobiology and mitochondrial function have been assembled along with state of the art technologies required to develop candidate NeuroSERM molecules.

Aged↗