Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Bioinformatics in proteomics.

Proteomics technologies are under continuous improvements and new technologies are introduced. Nowadays high throughput acquisition of proteome data is possible. The young and rapidly emerging field of bioinformatics in proteomics is introducing new algorithms to handle large and heterogeneous data sets and to improve the knowledge discovery process. For example new algorithms for image analysis of two dimensional gels have been developed within the last five years. Within mass spectrometry data analysis algorithms for peptide mass fingerprinting (PMF) and peptide fragmentation fingerprinting (PFF) have been developed. Local proteomics bioinformatics platforms emerge as data management systems and knowledge bases in Proteomics. We review recent developments in bioinformatics for proteomics with emphasis on expression proteomics.

Animals↗

Mining the malaria transcriptome.

Malaria remains the most devastating parasitic disease worldwide, and is responsible each year for >500 million infections and between one million and two million deaths of children under five years of age. Plasmodium falciparum is the most prevalent and deadly malaria parasite of humans, and a huge amount of data about it is now publicly available following completion of its genome sequence, the complete transcriptome of its asexual blood stages and proteomic analyses of its different life stages. Thus, new computational approaches are needed to analyze these data to yield biologically meaningful results that can be validated experimentally and, hopefully, lead to alternative control strategies. In this article, we highlight the importance of new computational approaches in mining the malaria transcriptome of the intraerythrocytic developmental cycle of P. falciparum.

Animals↗

PHProteomicDB: a module for two-dimensional gel electrophoresis database creation on personal web sites.

PHProteomicDB is a PHP-written module to help researchers in proteomics to share two-dimensional electrophoresis gel data using personal web sites. No technical or PHP knowledge is necessary except a few basics about web site management. PHProteomicDB has a user-friendly administration interface to enter and update data. It creates web pages on the fly displaying gel characteristics, gel pictures, and numbered gel spots with their related identifications pointing to their reference pages in protein databanks. The module is freely available at http://www.huvec.com/index.php3?rub=Download.

Animals↗

Human bronchoalveolar lavage: biofluid analysis with special emphasis on sample preparation.

Respiratory diseases are an important health problem throughout the world. Whether caused by industrial pollutants, infections, smoking, cancer or metabolic diseases, damage to the lungs and airways often lead to morbidity or death. Bronchoalveolar lavage (BAL) obtained by fiber-optic bronchoscopy is a biofluid mirroring the expression of normally secreted pulmonary proteins and the products of activated cells and destructive processes. The characterization of the proteome within this compartment provides an opportunity to establish temporal and prognostic indicators of airway disease. The objective of this study was to develop methods of analysis of BAL samples, which achieved the highest level of annotation of the expression map of this proteome. We have optimized the process of sample preparation after investigating a variety of techniques including dialysis, ultramembrane filtration, precipitation and gel filtration. We have further studied methods to remove albumin from BAL in order to unmask proteins hidden on two-dimensional gels. In a pilot application of the method, BAL protein profiles obtained from healthy nonsmokers and smokers at risk for developing chronic obstructive pulmonary disease showed distinct differences.

Albumins↗

Co-evolutionary analysis reveals insights into protein-protein interactions.

Protein-protein interactions play crucial roles in biological processes. Experimental methods have been developed to survey the proteome for interacting partners and some computational approaches have been developed to extend the impact of these experimental methods. Computational methods are routinely applied to newly discovered genes to infer protein function and plausible protein-protein interactions. Here, we develop and extend a quantitative method that identifies interacting proteins based upon the correlated behavior of the evolutionary histories of protein ligands and their receptors. We have studied six families of ligand-receptor pairs including: the syntaxin/Unc-18 family, the GPCR/G-alpha's, the TGF-beta/TGF-beta receptor system, the immunity/colicin domain collection from bacteria, the chemokine/chemokine receptors, and the VEGF/VEGF receptor family. For correlation scores above a defined threshold, we were able to find an average of 79% of all known binding partners. We then applied this method to find plausible binding partners for proteins with uncharacterized binding specificities in the syntaxin/Unc-18 protein and TGF-beta/TGF-beta receptor families. Analysis of the results shows that co-evolutionary analysis of interacting protein families can reduce the search space for identifying binding partners by not only finding binding partners for uncharacterized proteins but also recognizing potentially new binding partners for previously characterized proteins. We believe that correlated evolutionary histories provide a route to exploit the wealth of whole genome sequences and recent systematic proteomic results to extend the impact of these studies and focus experimental efforts to categorize physiologically or pathologically relevant protein-protein interactions.

Algorithms↗

MPAC: a computational framework for inferring pathway activities from multi-omic data.

Fully capturing cellular state requires examining genomic, epigenomic, transcriptomic, proteomic, and other assays for a biological sample and comprehensive computational modeling to reason with the complex and sometimes conflicting measurements. Modeling these so-called multi-omic data is especially beneficial in disease analysis, where observations across omic data types may reveal unexpected patient groupings and inform clinical outcomes and treatments. We present Multi-omic Pathway Analysis of Cells (MPAC), a computational framework that interprets multi-omic data through prior knowledge from biological pathways. MPAC leverages network relationships encoded in pathways through a factor graph to infer consensus activity levels for proteins and associated pathway entities from multi-omic data, runs permutation testing to eliminate spurious activity predictions, and groups biological samples by pathway activities to allow identifying and prioritizing proteins with potential clinical relevance, e.g., associated with patient prognosis. Using DNA copy number alteration and RNA-seq data from head and neck squamous cell carcinoma patients from The Cancer Genome Atlas as an example, we demonstrate that MPAC predicts a patient subgroup related to immune responses not identified by analysis with either input omic data type alone. Key proteins identified via this subgroup have pathway activities related to clinical outcome as well as immune cell compositions. Our MPAC R package, available at https://bioconductor.org/packages/MPAC, enables similar multi-omic analyses on new datasets.

Journal Article↗

Initial proteome analysis of model microorganism Haemophilus influenzae strain Rd KW20.

The proteome of Haemophilus influenzae strain Rd KW20 was analyzed by liquid chromatography (LC) coupled with ion trap tandem mass spectrometry (MS/MS). This approach does not require a gel electrophoresis step and provides a rapidly developed snapshot of the proteome. In order to gain insight into the central metabolism of H. influenzae, cells were grown microaerobically and anaerobically in a rich medium and soluble and membrane proteins of strain Rd KW20 were proteolyzed with trypsin and directly examined by LC-MS/MS. Several different experimental and computational approaches were utilized to optimize the proteome coverage and to ensure statistically valid protein identification. Approximately 25% of all predicted proteins (open reading frames) of H. influenzae strain Rd KW20 were identified with high confidence, as their component peptides were unambiguously assigned to tandem mass spectra. Approximately 80% of the predicted ribosomal proteins were identified with high confidence, compared to the 33% of the predicted ribosomal proteins detected by previous two-dimensional gel electrophoresis studies. The results obtained in this study are generally consistent with those obtained from computational genome analysis, two-dimensional gel electrophoresis, and whole-genome transposon mutagenesis studies. At least 15 genes originally annotated as conserved hypothetical were found to encode expressed proteins. Two more proteins, previously annotated as predicted coding regions, were detected with high confidence; these proteins also have close homologs in related bacteria. The direct proteomics approach to studying protein expression in vivo reported here is a powerful method that is applicable to proteome analysis of any (micro)organism.

Aerobiosis↗

A glycoproteome database of normal human liver tissue.

PURPOSE: To extensively investigate the glycoproteins of normal human liver tissue, constructing the glycoprotein profile and database of the normal human liver tissue. METHODS: The total proteins were extracted from the normal human liver tissue and then subjected to two-dimensional electrophoresis (2-DE). Finally, 2-DE gels were stained according to the methods of multiplexed proteomics (MP) technology. Glycoprotein spots were excised from 2-DE gel and then characterized by matrix assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF-MS). RESULTS: The PDQuest software detected 1,011 glycoprotein spots and 1,923 total protein spots in the 2-DE gels of sample from the normal human liver tissue. Furthermore, 116 species of glycoproteins were successfully identified via peptide mass profiling using MALDI-TOF-MS/MS and annotated to our databases. In addition, we also applied bioinformatics softwares to predict N- or O-glycosylation sites of identified glycoproteins. CONCLUSION: This study demonstrates the feasibility of a novel technological platform to contruct glycoprotein databases. These results lay the foundation for future physiological and pathological studies of the human liver.

Databases, Protein↗

Bioinformatics in proteomics: application, terminology, and pitfalls.

Bioinformatics applies data mining, i.e., modern computer-based statistics, to biomedical data. It leverages on machine learning approaches, such as artificial neural networks, decision trees and clustering algorithms, and is ideally suited for handling huge data amounts. In this article, we review the analysis of mass spectrometry data in proteomics, starting with common pre-processing steps and using single decision trees and decision tree ensembles for classification. Special emphasis is put on the pitfall of overfitting, i.e., of generating too complex single decision trees. Finally, we discuss the pros and cons of the two different decision tree usages.

Computational Biology↗

Functional proteomics using microchannel plate detectors.

We describe the development of a novel detection system used for the functional imaging of proteins separated on electrophoretic gels. A microchannel plate detector is used here for real-time imaging of low levels of tritiated protein separated by two-dimensional (2-D) electrophoresis. The system employs radioisotope-free, low noise microchannel plates originally developed for photon counting in X-ray astronomy. Using the detector configuration described here, proteins were resolved on mini gels by either one or two-dimensional electrophoresis, transferred onto polyvinylidene difluoride membranes and directly imaged. Tritiated diisopropylfluorophosphate (DFP) was used as a selective label for the serine hydrolase class of enzymes and their distribution in the central nervous system was examined. This survey revealed approximately 24 protein spots by 2-D electrophoresis. We also investigated the relative sensitivity of these proteins towards DFP and found the peptidase, acylpeptide hydrolase to be the most sensitive brain protein towards this reagent. Using a number of different tritiated standards, it was found that the system can image as little as 0.1 Bq/mm(2) of tritium corresponding to 320 attomol of DFP labelled protein/mm(2). Moreover, the system has a wide dynamic range (>10(6)) allowing samples of high and low activity to be quantified on the same gel.

Animals↗

Proteomics reveals protein profile changes in doxorubicin--treated MCF-7 human breast cancer cells.

MCF-7 cells are extensively used as a cell model to investigate human breast tumors and the cellular mechanism of antitumor drugs such as doxorubicin (DOX), an anthracycline antitumor drug widely used in clinical chemotherapy. To understand the effects of DOX on the protein expression, we perform a comprehensive proteomics to survey global changes in proteins after DOX treatment in MCF-7 cells. Exposure of MCF-7 cells to 0.1 microM DOX for 2 days induced a differentiation-like phenotype with prominent perinuclear autocatalytic vacuoles, abundant filamentous material, and irregular microvilli at the cell surface. In this study, we also present a proteome reference map of MCF-7 cells with 21 identified protein spots via analysis of N-terminal sequencing, mass spectrometry, immunoblot and/or computer matching with protein database. Based on the proteome map, we found that DOX causes a markedly decrease in the levels of three isoforms of heat shock protein 27 (HSP27) whereas the levels of other stress associated proteins including HSP60, calreticulin, and protein disulfide isomerase were not significantly altered in DOX-treated MCF-7 cells. Taken together, we suggest that that action of DOX on breast tumor cells may be partly related to dysregulation of HSP27 expression. Modulation of HSP27 levels may be a clinically useful potential target for design of antitumor drugs and controlling breast tumor growth.

Antineoplastic Agents↗

Will genomics revolutionise pharmaceutical R&D?

The successful identification of drug targets requires an understanding of the high-level functional interactions between the key components of cells, organs and systems, and how these interactions change in disease states. This information does not reside in the genome, or in the individual proteins that genes code for, it is to be found at a higher level. Genomics will succeed in revolutionising pharmaceutical research and development only if these interactions are also understood by determining the logic of healthy and diseased states. The rapid growth in biological databases, models of cells, tissues and organs, and in computing power has made it possible to explore functionality all the way from the level of genes to whole organs and systems. Combined with genomic and proteomic data, in silico simulation technology is set to transform all stages of drug discovery and development. The major obstacle to achieving this will be obtaining the relevant experimental data at levels higher than genomics and proteomics.

Cell Physiological Phenomena↗

Quantitative and reproducible two-dimensional gel analysis using Phoretix 2D Full.

Quantitative two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) is used to determine changes in individual protein levels in complex protein mixtures. To provide reliable data, the software used for 2-D gel image analysis must provide a linear response over a wide dynamic range of data output. Here, we show that Phoretix 2D Full analysis of 2-D gels stained with colloidal Coomassie Brilliant Blue G-250 can provide a linear measure of changes in protein quantity. We show using a complex mixture of Arabidopsis thaliana proteins, that this is true for essentially all focused proteins, in a data output range greater than three orders of magnitude. An analysis of the factors that affect errors in the results demonstrated that reproducibility of the data is significantly improved by user seeding, whereas it is reduced by use of the background subtraction algorithms.

Algorithms↗

Using cellular automata images and pseudo amino acid composition to predict protein subcellular location.

The avalanche of newly found protein sequences in the post-genomic era has motivated and challenged us to develop an automated method that can rapidly and accurately predict the localization of an uncharacterized protein in cells because the knowledge thus obtained can greatly speed up the process in finding its biological functions. However, it is very difficult to establish such a desired predictor by acquiring the key statistical information buried in a pile of extremely complicated and highly variable sequences. In this paper, based on the concept of the pseudo amino acid composition (Chou, K. C. PROTEINS: Structure, Function, and Genetics, 2001, 43: 246-255), the approach of cellular automata image is introduced to cope with this problem. Many important features, which are originally hidden in the long amino acid sequences, can be clearly displayed through their cellular automata images. One of the remarkable merits by doing so is that many image recognition tools can be straightforwardly applied to the target aimed here. High success rates were observed through the self-consistency, jackknife, and independent dataset tests, respectively.

Algorithms↗

Spectroscopic imaging of protein crystals in crystallization drops.

Automatic imaging and scoring of crystallization drops is an essential step in high-throughput crystallography. Presently, white-light images of crystallization drops are acquired robotically and the images are analyzed and scored using pattern recognition algorithms. However, the scoring part remains unreliable as crystals and microcrystals are not always recognized by existing feature-extraction and recognition algorithms. We propose a fundamental shift in crystal monitoring through spectroscopic imaging of crystallization drops. This method converts the problem of automatic crystal detection from one of pattern recognition into one of intensity (concentration) analysis. The latter can be more robust and reliable.

Crystallization↗

Automatic classification and pattern discovery in high-throughput protein crystallization trials.

Conceptually, protein crystallization can be divided into two phases search and optimization. Robotic protein crystallization screening can speed up the search phase, and has a potential to increase process quality. Automated image classification helps to increase throughput and consistently generate objective results. Although the classification accuracy can always be improved, our image analysis system can classify images from 1,536-well plates with high classification accuracy (85%) and ROC score (0.87), as evaluated on 127 human-classified protein screens containing 5,600 crystal images and 189,472 non-crystal images. Data mining can integrate results from high-throughput screens with information about crystallizing conditions, intrinsic protein properties, and results from crystallization optimization. We apply association mining, a data mining approach that identifies frequently occurring patterns among variables and their values. This approach segregates proteins into groups based on how they react in a broad range of conditions, and clusters cocktails to reflect their potential to achieve crystallization. These results may lead to crystallization screen optimization, and reveal associations between protein properties and crystallization conditions. We also postulate that past experience may lead us to the identification of initial conditions favorable to crystallization for novel proteins.

Algorithms↗