Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Hierarchical analysis of large-scale two-dimensional gel electrophoresis experiments.

Large-scale two-dimensional gel experiments have the potential to identify proteins that play an important role in elucidating cell mechanisms and in various stages of drug discovery. Such experiments, typically including hundreds or even thousands of related gels, are notoriously difficult to perform, and analysis of the gel images has until recently been virtually impossible. In this paper we describe a scalable computational model that permits the organization and analysis of a large gel collection. The model is implemented in Compugen's Z4000 system. Gels are organized in a hierarchical, multidimensional data structure that allow the user to view a large-scale experiment as a tree of numerous simpler experiments, and carry out the analysis one step at a time. Analyzed sets of gels form processing units that can be combined into higher level units in an iterative framework. The different conditions at the core of the experiment design, termed the dimensions of the experiment, are transformed from a multidimensional structure to a single hierarchy. The higher level comparison is performed with the aid of a synthetic "adaptor" gel image, called a Raw Master Gel (RMG). The RMG allows the inclusion of data from an entire set of gels to be presented as a gel image, thereby enabling the iterative process. Our model includes a flexible experimental design approach that allows the researcher to choose the condition to be analyzed a posteriori. It also enables data reuse, the performing of several different analysis designs on the same experimental data. The stability and reproducibility of a protein can be analyzed by tracking it up or down the hierarchical dimensions of the experiment.

Computational Biology↗

High-throughput proteomics using matrix-assisted laser desorption/ ionization mass spectrometry.

It has become evident that the mystery of life will not be deciphered just by decoding its blueprint, the genetic code. In the life and biomedical sciences, research efforts are now shifting from pure gene analysis to the analysis of all biomolecules involved in the machinery of life. One area of these postgenomic research fields is proteomics. Although proteomics, which basically encompasses the analysis of proteins, is not a new concept, it is far from being a research field that can rely on routine and large-scale analyses. At the time the term proteomics was coined, a gold-rush mentality was created, promising vast and quick riches (i.e., solutions to the immensely complex questions of life and disease). Predictably, the reality has been quite different. The complexity of proteomes and the wide variations in the abundances and chemical properties of their constituents has rendered the use of systematic analytical approaches only partially successful, and biologically meaningful results have been slow to arrive. However, to learn more about how cells and, hence, life works, it is essential to understand the proteins and their complex interactions in their native environment. This is why proteomics will be an important part of the biomedical sciences for the foreseeable future. Therefore, any advances in providing the tools that make protein analysis a more routine and large-scale business, ideally using automated and rapid analytical procedures, are highly sought after. This review will provide some basics, thoughts and ideas on the exploitation of matrix-assisted laser desorption/ ionization in biological mass spectrometry - one of the most commonly used analytical tools in proteomics - for high-throughput analyses.

Computational Biology↗

Identification of putative exported/secreted proteins in prokaryotic proteomes.

The increasing number of bacterial genomes being sequenced fuels an equal demand for methods to rapidly analyze the proteomes of these organisms. One group of proteins of pressing importance is the exported/secreted proteins, given their dominant immunogenicity and role in pathogenesis. With this in mind, a weight matrix algorithm and two artificial neural networks, one based on amino acid position within the N-terminus and the other on amino acid frequency, were developed for identification of such proteins. The neural networks and a hybrid method, combining the weight matrix algorithm and the amino acid frequency neural network, were tested independently against a standard data set of secreted and cytoplasmic proteins to determine their accuracy in predicting secreted prokaryotic proteins. The results of these analyses demonstrated that the amino acid position neural network provided the highest accuracy (Mathews correlation coefficient of 0.93) in predicting secreted proteins of Gram-negative bacteria, whereas the hybrid method was best (Mathews correlation coefficient of 0.97) for prediction of Gram-positive secreted proteins. These two methods were integrated into a single program (ExProt) designed to analyze whole proteomes. In addition to protein localization, ExProt also contains a neural network trained to identify the most probable signal peptidase I cleavage site of secreted proteins. When tested against the standard protein data set ExProt correctly predicted 73.5 and 84.5% of the cleavage sites in Gram-positive and Gram-negative secreted proteins, respectively. Comparative analysis of Gram-negative, Gram-positive, Mycobacterium tuberculosis, and Archaea proteomes with ExProt revealed that the fraction of putative exported/secreted proteins encoded by bacterial genomes ranged from 8% for Methanococcus jannaschii to 37% for Mycoplasma pneumoniae.

Algorithms↗

Proteome analysis of the rice etioplast: metabolic and regulatory networks and novel protein functions.

We report an extensive proteome analysis of rice etioplasts, which were highly purified from dark-grown leaves by a novel protocol using Nycodenz density gradient centrifugation. Comparative protein profiling of different cell compartments from leaf tissue demonstrated the purity of the etioplast preparation by the absence of diagnostic marker proteins of other cell compartments. Systematic analysis of the etioplast proteome identified 240 unique proteins that provide new insights into heterotrophic plant metabolism and control of gene expression. They include several new proteins that were not previously known to localize to plastids. The etioplast proteins were compared with proteomes from Arabidopsis chloroplasts and plastid from tobacco Bright Yellow 2 cells. Together with computational structure analyses of proteins without functional annotations, this comparative proteome analysis revealed novel etioplast-specific proteins. These include components of the plastid gene expression machinery such as two RNA helicases, an RNase II-like hydrolytic exonuclease, and a site 2 protease-like metalloprotease all of which were not known previously to localize to the plastid and are indicative for so far unknown regulatory mechanisms of plastid gene expression. All etioplast protein identifications and related data were integrated into a data base that is freely available upon request.

Amino Acid Sequence↗

Proteomic analysis.

The field of proteomics is becoming increasingly important as genome sequences are being completed and annotated. Recent advances in proteomics include experimental and mathematical proofs of the need to complement microarray analysis with protein analysis, improved sensitivity for mass spectrometric analysis of separated proteins, better informatic tools for gel analysis and protein spot annotation, first steps towards automated experimental procedures, and new technology for quantitation of protein changes.

Animals↗

A generalized higher-order correlation analysis framework for multi-omics network inference.

Multiple -omics (genomics, proteomics, etc.) profiles are commonly generated to gain insight into a disease or physiological system. Constructing multi-omics networks with respect to the trait(s) of interest provides an opportunity to understand relationships between molecular features but integration is challenging due to multiple data sets with high dimensionality. One approach is to use canonical correlation to integrate one or two omics types and a single trait of interest. However, these types of methods may be limited due to (1) not accounting for higher-order correlations existing among features, (2) computational inefficiency when extending to more than two omics data when using a penalty term-based sparsity method, and (3) lack of flexibility for focusing on specific correlations (e.g., omics-to-phenotype correlation versus omics-to-omics correlations). In this work, we have developed a novel multi-omics network analysis pipeline called Sparse Generalized Tensor Canonical Correlation Analysis Network Inference (SGTCCA-Net) that can effectively overcome these limitations. We also introduce an implementation to improve the summarization of networks for downstream analyses. Simulation and real-data experiments demonstrate the effectiveness of our novel method for inferring omics networks and features of interest.

Genomics↗

Proteomics in environmental pollution research: Advances, challenges, and future directions.

Environmental proteomics has emerged as a powerful approach for elucidating the molecular mechanisms underlying pollutant-induced biological effects. Although this field has developed rapidly, the systematic review of recent proteomics applications in environmental pollution research remains limited. This review explored the emerging roles of toxicoproteomics in biomarker discovery and mechanistic elucidation, as well as ecotoxicoproteomics in ecological risk assessment and bioremediation strategies. Here, we review the field, highlighting recent trends such as the integration of proteomics with genomics, transcriptomics, and metabolomics to provide a comprehensive view of biological responses to environmental stressors. We further discuss the growing application of artificial intelligence in improving proteomics data interpretation and accelerating biomarker discovery. In addition, recent technological advances in environmental proteomics are highlighted, including next-generation tissue microarray proteomics, nanoscale proteomics, single-cell proteomics, and spatial proteomics. Despite its potential, proteomics faces challenges, such as high operational costs, computational complexity in analysis, and technical limitations in low-abundance protein detection. We propose that the convergence of proteomics with artificial intelligence and multi-omics approaches offers promising solutions to these challenges, enhancing the practical application of proteomics in environmental monitoring and risk assessment.

Proteomics↗

The molecular similarity landscape of preclinical cancer models to patient tumors.

Selecting appropriate preclinical models is fundamental for translational oncology, yet a large-scale, multi-omic quantitative comparison of their similarity to primary human tumors is lacking. To address this, we integrated transcriptomic, proteomic, and genomic profiles from over 10,000 primary tumors from The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC), alongside 4,000 preclinical models. Using a robust computational framework, we revealed a clear hierarchy of transcriptomic and proteomic similarity to patient tumors: with patient-dervied xenografts (PDXs) having greater transcriptomic and proteomic similarity to patient tumors (>) compared with patient-derived organoids (PDOs), which are equal in hierarchy to that of PDX-dervied organoids (PDXOs) > cell lines. We also quantified high molecular conservation (Pearson correlation coefficient = 0.96) across paired in vitro to in vivo platform (organoids to PDX) transitions. Furthermore, genomic analysis demonstrated that whole-exome sequencing (WES) outperforms RNA-seq in detecting DNA variants, and it identified a clonal complexity hierarchy (cell lines > PDXOs > PDXs > PDOs) reflecting the effect of passaging history on intratumor heterogeneity. Ultimately, this study delivers a comprehensive quantitative benchmark, establishing a population-level hierarchy of molecular similarity between preclinical models and primary tumors and providing a data-driven reference for model selection. These findings offer a data-driven framework for selecting models that balance biological representativeness with experimental practicality.

Humans↗

Geometric algorithms for the analysis of 2D-electrophoresis gels.

In proteomics, two-dimensional gel electrophoresis (2-DE) is a separation technique for proteins. The resulting protein spots can be identified either by using picking robots and subsequent mass spectrometry or by visual cross inspection of a new gel image with an already analyzed master gel. Difficulties especially arise from inherent noise and irregular geometric distortions in 2-DE images. Aiming at the automated analysis of large series of 2-DE images, or at the even more difficult interlaboratory gel comparisons, the bottleneck is to solve the two most basic algorithmic problems with high quality: Identifying protein spots and computing a matching between two images. For the development of the analysis software CAROl at Freie Universität Berlin, we have reconsidered these two problems and obtained new solutions which rely on methods from computational geometry. Their novelties are: 1. Spot detection is also possible for complex regions formed by several "merged" (usually saturated) spots; 2. User-defined landmarks are not necessary for the matching. Furthermore, images for comparison are allowed to represent different parts of the entire protein pattern, which only partially "overlap." The implementation is done in a client server architecture to allow queries via the internet. We also discuss and point at related theoretical questions in computational geometry.

Algorithms↗

Biomarker discovery by proteomics: challenges not only for the analytical chemist.

This forum article outlines some of the major challenges in present day biomarker discovery research. Notably the dilemma of reaching sufficient concentration sensitivity versus the required analysis time per sample is highlighted using a model calculation. A number of possible developments and recent research findings are considered to show possible ways out of this dilemma. Finally, the challenge of processing large, megavariate datasets prior to statistical analysis is presented.

Animals↗

High-throughput proteomics for alcohol research.

This report summarizes the proceedings of a satellite symposium of the 2003 Research Society on Alcoholism meeting held on June 21, 2003, in Fort Lauderdale, FL. The goal of this symposium, sponsored by the NIAAA, was to identify new proteomic directions in alcohol research that will (1) enable studies that focus on characterizing protein function, biochemical pathways, and networks to understand alcohol-related illnesses; (2) identify protein-protein interactions, posttranslational modifications, and subcellular localizations; (3) identify molecular targets for medication development; (4) develop biomarkers for susceptibility, dependence, consumption, and relapse, as well as alcohol-induced pathologies; and (5) develop high-throughput drug screens to test the efficacy of therapeutics that control alcohol-induced diseases. The purpose of the symposium was also to promote the application of high-throughput proteomic approaches, including isolation of membrane-bound proteins, in situ proteomics, large-scale two-dimensional separations, protein microarray platforms, mass spectrometry, matrix-assisted laser desorption/ionization, matrix-assisted laser desorption/ionization time-of-flight, liquid chromatography-tandem mass spectrometry, and isotope-coded affinity tags. In addition, the development of protein network maps by using new bioinformatics approaches for database mining was also discussed.

Alcohol Drinking↗

Improved Ruthenium II tris (bathophenantroline disulfonate) staining and destaining protocol for a better signal-to-background ratio and improved baseline resolution.

In proteomics the ability to visualize proteins from electropherograms is essential. Here a new protocol for staining and destaining gels treated with Ruthenium II tris (bathophenantroline disulfonate) is presented. The method is compared with the silver-staining procedure of Swain and Ross, the Ruthenium II tris (bathophenantroline disulfonate) stain described by Rabilloud (Rabilloud T., Strub, S. M. Luche, S., Girardet, S. L. et al., Proteomics 2001, 1, 699-704) and the SYPRO Ruby gel stain. The method offers a better signal-to-background ratio with improved baseline resolution for both sodium dodecyl sulfate-polyacrylamide gels and two-dimensional gels.

2,2'-Dipyridyl↗

Review: prediction of in vivo fates of proteins in the era of genomics and proteomics.

Even after a nascent protein emerges from the ribosome, its fate is still controlled by its own amino acid sequence information. Namely, it may be co-/posttranslationally modified (e.g., phosphorylated, N-/O-glycosylated, and lipidated); it may be inserted into the membrane, translocated to an organelle, or secreted to the outside milieu; it may be processed for maturation or selective degradation; finally, its fragment may be presented on the cell surface as an antigen. Here, prediction methods of such protein fates from their amino acid sequences are reviewed. In many cases, artificial neural network techniques have been effectively used. The prediction of in vivo fates of proteins will be useful for characterizing newly identified candidate genes in a genome or for interpreting multiple spots in proteome analyses.

Animals↗

Scanning the available Dictyostelium discoideum proteome for O-linked GlcNAc glycosylation sites using neural networks.

Dictyostelium discoideum has been suggested as a eukaryotic model organism for glycobiology studies. Presently, the characteristics of acceptor sites for the N-acetylglucosaminyl-transferases in Dictyostelium discoideum, which link GlcNAc in an alpha linkage to hydroxyl residues, are largely unknown. This motivates the development of a species specific method for prediction of O-linked GlcNAc glycosylation sites in secreted and membrane proteins of D. discoideum. The method presented here employs a jury of artificial neural networks. These networks were trained to recognize the sequence context and protein surface accessibility in 39 experimentally determined O-alpha-GlcNAc sites found in D. discoideum glycoproteins expressed in vivo. Cross-validation of the data revealed a correlation in which 97% of the glycosylated and nonglycosylated sites were correctly identified. Based on the currently limited data set, an abundant periodicity of two (positions-3, -1, +1, +3, etc.) in Proline residues alternating with hydroxyl amino acids was observed upstream and downstream of the acceptor site. This was a consequence of the spacing of the glycosylated residues themselves which were peculiarly found to be situated only at even positions with respect to each other, indicating that these may be located within beta-strands. The method has been used for a rapid and ranked scan of the fraction of the Dictyostelium proteome available in public databases, remarkably 25-30% of which were predicted glycosylated. The scan revealed acceptor sites in several proteins known experimentally to be O-glycosylated at unmapped sites. The available proteome was classified into functional and cellular compartments to study any preferential patterns of glycosylation. A sequence based prediction server for GlcNAc O-glycosylations in D. discoideum proteins has been made available through the WWW at http://www.cbs.dtu.dk/services/DictyOGlyc/ and via E-mail to DictyOGlyc@cbs.dtu.dk.

Algorithms↗

MSight: an image analysis software for liquid chromatography-mass spectrometry.

Images obtained from high-throughput mass spectrometry (MS) contain information that remains hidden when looking at a single spectrum at a time. Image processing of liquid chromatography-MS datasets can be extremely useful for quality control, experimental monitoring and knowledge extraction. The importance of imaging in differential analysis of proteomic experiments has already been established through two-dimensional gels and can now be foreseen with MS images. We present MSight, a new software designed to construct and manipulate MS images, as well as to facilitate their analysis and comparison.

Chromatography, Liquid↗

Data mining techniques for cancer detection using serum proteomic profiling.

OBJECTIVE: Pathological changes in an organ or tissue may be reflected in proteomic patterns in serum. It is possible that unique serum proteomic patterns could be used to discriminate cancer samples from non-cancer ones. Due to the complexity of proteomic profiling, a higher order analysis such as data mining is needed to uncover the differences in complex proteomic patterns. The objectives of this paper are (1) to briefly review the application of data mining techniques in proteomics for cancer detection/diagnosis; (2) to explore a novel analytic method with different feature selection methods; (3) to compare the results obtained on different datasets and that reported by Petricoin et al. in terms of detection performance and selected proteomic patterns. METHODS AND MATERIAL: Three serum SELDI MS data sets were used in this research to identify serum proteomic patterns that distinguish the serum of ovarian cancer cases from non-cancer controls. A support vector machine-based method is applied in this study, in which statistical testing and genetic algorithm-based methods are used for feature selection respectively. Leave-one-out cross validation with receiver operating characteristic (ROC) curve is used for evaluation and comparison of cancer detection performance. RESULTS AND CONCLUSIONS: The results showed that (1) data mining techniques can be successfully applied to ovarian cancer detection with a reasonably high performance; (2) the classification using features selected by the genetic algorithm consistently outperformed those selected by statistical testing in terms of accuracy and robustness; (3) the discriminatory features (proteomic patterns) can be very different from one selection method to another. In other words, the pattern selection and its classification efficiency are highly classifier dependent. Therefore, when using data mining techniques, the discrimination of cancer from normal does not depend solely upon the identity and origination of cancer-related proteins.

Biomarkers, Tumor↗