Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Computational analysis of crystallization trials.

A system for the automatic categorization of the results of crystallization experiments generated by robotic screening is presented. Images from robotically generated crystallization screens are taken at preset time intervals and analyzed by the computer program Crystal Experiment Evaluation Program (CEEP). This program attempts to automatically categorize the individual crystal experiments into a number of simple classes ranging from clear drop to mountable crystal. The algorithm first selects features from the images via edge detection and texture analysis. Classification is achieved via a self-organizing neural net generated from a set of hand-classified images used as a training set. New images are then classified according to this neural net. It is demonstrated that incorporation of time-series information may enhance the accuracy of classification. Preliminary results from the screening of the proteome of Thermotoga maritima are presented showing the utility of the system.

Algorithms↗

Behavioural manipulation in a grasshopper harbouring hairworm: a proteomics approach.

The parasitic Nematomorph hairworm, Spinochordodes tellinii (Camerano) develops inside the terrestrial grasshopper, Meconema thalassinum (De Geer) (Orthoptera: Tettigoniidae), changing the insect's responses to water. The resulting aberrant behaviour makes infected insects more likely to jump into an aquatic environment where the adult parasite reproduces. We used proteomics tools (i.e. two-dimensional gel electrophoresis (2-DE), computer assisted comparative analysis of host and parasite protein spots and MALDI-TOF mass spectrometry) to identify these proteins and to explore the mechanisms underlying this subtle behavioural modification. We characterized simultaneously the host (brain) and the parasite proteomes at three stages of the manipulative process, i.e. before, during and after manipulation. For the host, there was a differential proteomic expression in relation to different effects such as the circadian cycle, the parasitic status, the manipulative period itself, and worm emergence. For the parasite, a differential proteomics expression allowed characterization of the parasitic and the free-living stages, the manipulative period and the emergence of the worm from the host. The findings suggest that the adult worm alters the normal functions of the grasshopper's central nervous system (CNS) by producing certain 'effective' molecules. In addition, in the brain of manipulated insects, there was found to be a differential expression of proteins specifically linked to neurotransmitter activities. The evidence obtained also suggested that the parasite produces molecules from the family Wnt acting directly on the development of the CNS. These proteins show important similarities with those known in other insects, suggesting a case of molecular mimicry. Finally, we found many proteins in the host's CNS as well as in the parasite for which the function(s) are still unknown in the published literature (www) protein databases. These results support the hypothesis that host behavioural changes are mediated by a mix of direct and indirect chemical manipulation.

Animals↗

Interpretation of shotgun proteomic data: the protein inference problem.

The shotgun proteomic strategy based on digesting proteins into peptides and sequencing them using tandem mass spectrometry and automated database searching has become the method of choice for identifying proteins in most large scale studies. However, the peptide-centric nature of shotgun proteomics complicates the analysis and biological interpretation of the data especially in the case of higher eukaryote organisms. The same peptide sequence can be present in multiple different proteins or protein isoforms. Such shared peptides therefore can lead to ambiguities in determining the identities of sample proteins. In this article we illustrate the difficulties of interpreting shotgun proteomic data and discuss the need for common nomenclature and transparent informatic approaches. We also discuss related issues such as the state of protein sequence databases and their role in shotgun proteomic analysis, interpretation of relative peptide quantification data in the presence of multiple protein isoforms, the integration of proteomic and transcriptional data, and the development of a computational infrastructure for the integration of multiple diverse datasets.

Amino Acid Sequence↗

Quality control and peak finding for proteomics data collected from nipple aspirate fluid by surface-enhanced laser desorption and ionization.

BACKGROUND: Recently, researchers have been using mass spectroscopy to study cancer. For use of proteomics spectra in a clinical setting, stringent quality-control procedures will be needed. METHODS: We pooled samples of nipple aspirate fluid from healthy breasts and breasts with cancer to prepare a control sample. Aliquots of the control sample were used on two spots on each of three IMAC ProteinChip arrays (Ciphergen Biosystems, Inc.) on 4 successive days to generate 24 SELDI spectra. In 36 subsequent experiments, the control sample was applied to two spots of each ProteinChip array, and the resulting spectra were analyzed to determine how closely they agreed with the original 24 spectra. RESULTS: We describe novel algorithms that (a) locate peaks in unprocessed proteomics spectra and (b) iteratively combine peak detection with baseline correction. These algorithms detected approximately 200 peaks per spectrum, 68 of which are detected in all 24 original spectra. The peaks were highly correlated across samples. Moreover, we could explain 80% of the variance, using only six principal components. Using a criterion that rejects a chip if the Mahalanobis distance from both control spectra to the center of the six-dimensional principal component space exceeds the 95% confidence limit threshold, we rejected 5 of the 36 chips. CONCLUSIONS: Mahalanobis distance in principal component space provides a method for assessing the reproducibility of proteomics spectra that is robust, effective, easily computed, and statistically sound.

Algorithms↗

MAVisto: a tool for the exploration of network motifs.

UNLABELLED: MAVisto is a tool for the exploration of motifs in biological networks. It provides a flexible motif search algorithm and different views for the analysis and visualization of network motifs. These views help to explore interesting motifs: the frequency of motif occurrences can be compared with randomized networks, a list of motifs along with information about structure and number of occurrences depending on the reuse of network elements shows potentially interesting motifs, a motif fingerprint reveals the overall distribution of motifs of a given size and the distribution of a particular motif in the network can be visualized using an advanced layout algorithm. AVAILABILITY: MAVisto is platform independent and available free of charge as a Java webstart application at http://mavisto.ipk-gatersleben.de/ CONTACT: schwoebb@ipk-gatersleben.de SUPPLEMENTARY INFORMATION: Can be found at http://mavisto.ipk-gatersleben.de/

Computer Simulation↗

Proteomic profiling of urinary proteins in renal cancer by surface enhanced laser desorption ionization and neural-network analysis: identification of key issues affecting potential clinical utility.

Recent advances in proteomic profiling technologies, such as surface enhanced laser desorption ionization mass spectrometry, have allowed preliminary profiling and identification of tumor markers in biological fluids in several cancer types and establishment of clinically useful diagnostic computational models. There are currently no routinely used circulating tumor markers for renal cancer, which is often detected incidentally and is frequently advanced at the time of presentation with over half of patients having local or distant tumor spread. We have investigated the clinical utility of surface enhanced laser desorption ionization profiling of urine samples in conjunction with neural-network analysis to either detect renal cancer or to identify proteins of potential use as markers, using samples from a total of 218 individuals, and examined critical technical factors affecting the potential utility of this approach. Samples from patients before undergoing nephrectomy for clear cell renal cell carcinoma (RCC; n = 48), normal volunteers (n = 38), and outpatients attending with benign diseases of the urogenital tract (n = 20) were used to successfully train neural-network models based on either presence/absence of peaks or peak intensity values, resulting in sensitivity and specificity values of 98.3-100%. Using an initial "blind" group of samples from 12 patients with RCC, 11 healthy controls, and 9 patients with benign diseases to test the models, sensitivities and specificities of 81.8-83.3% were achieved. The robustness of the approach was subsequently evaluated with a group of 80 samples analyzed "blind" 10 months later, (36 patients with RCC, 31 healthy volunteers, and 13 patients with benign urological conditions). However, sensitivities and specificities declined markedly, ranging from 41.0% to 76.6%. Possible contributing factors including sample stability, changing laser performance, and chip variability were examined, which may be important for the long-term robustness of such approaches, and this study highlights the need for rigorous evaluation of such factors in future studies.

Adult↗

The modal distribution of protein isoelectric points reflects amino acid properties rather than sequence evolution.

Two-dimensional gel electrophoresis, a routine application in proteomics, separates proteins according to their molecular mass (M(r)) and isoelectric point (pI). As the genomic sequences for more and more organisms are determined, the M(r) and pI of all their proteins can be estimated computationally. The examination of several of these theoretical proteome plots has revealed a multimodal pI distribution, however, no conclusive explanation for this unusual distribution has so far been presented. We examined the pI distribution of 115 fully sequenced genomes and observed that the modal distribution does not reflect phylogeny or sequence evolution, but rather the chemical properties of amino acids. We provide a statistical explanation of why the observed distributions of pI values are multimodal.

Algorithms↗

Proteomic approaches to studying drug targets and resistance in Plasmodium.

Ever increasing drug resistance by Plasmodium falciparum, the most virulent of human malaria parasites, is creating new challenges in malaria chemotherapy. The entire genome sequences of P. falciparum and the rodent malaria parasite, P. yoelii yoelii are now available. Extensive genome sequence data from other Plasmodium species including another important human malaria parasite, P. vivax are also available. Powerful research techniques coupled to genomic resources are needed to help identify new drug and vaccine targets against malaria. Applied to Plasmodium, proteomics combines high-resolution protein or peptide separation with mass spectrometry and computer software to rapidly identify large numbers of proteins expressed from various stages of parasite development. Proteomic methods can be applied to study sub-cellular localization, cell function, organelle composition, changes in protein expression patterns in response to drug exposure, drug-protein binding and validation of data from genomic annotation and transcript expression studies. Recent high-throughput proteomic approaches have provided a wealth of protein expression data on P. falciparum, while smaller-scale studies examining specific drug-related hypotheses are also appearing. Of particular interest is the study of mechanisms of action and resistance of drugs such as the quinolines, whose targets currently may not be predictable from genomic data. Coupling the Plasmodium sequence data with bioinformatics, proteomics and RNA transcript expression profiling opens unprecedented opportunities for exploring new malaria control strategies. This review will focus on pharmacological research in malaria and other intracellular parasites using proteomic techniques, emphasizing resources and strategies available for Plasmodium.

Animals↗

HiRes--a tool for comprehensive assessment and interpretation of metabolomic data.

UNLABELLED: The increasing role of metabolomics in system biology is driving the development of tools for comprehensive analysis of high-resolution NMR spectral datasets. This task is quite challenging since unlike the datasets resulting from other 'omics', a substantial preprocessing of the data is needed to allow successful identification of spectral patterns associated with relevant biological variability. HiRes is a unique stand-alone software tool that combines standard NMR spectral processing functionalities with techniques for multi-spectral dataset analysis, such as principal component analysis and non-negative matrix factorization. In addition, HiRes contains extensive abilities for data cleansing, such as baseline correction, solvent peak suppression, removal of frequency shifts owing to experimental conditions as well as auxiliary information management. Integration of these components together with multivariate analytical procedures makes HiRes very capable of addressing the challenges for assessment and interpretation of large metabolomic datasets, greatly simplifying this otherwise lengthy and difficult process and assuring optimal information retrieval. AVAILABILITY: HiRes is freely available for research purposes at http://hatch.cpmc.columbia.edu/highresmrs.html

Algorithms↗

Cell cycle-dependent protein dynamics in budding yeast resolved by deconvolution of bulk proteomics.

The cell division cycle is characterised by oscillatory dynamics in regulatory mechanisms and biosynthesis, coordinated with genome replication and segregation. To understand these dynamics, quantitative cell cycle-dependent protein concentration data are essential. Unfortunately, accurately resolving cell cycle-dependent protein dynamics is challenging because single-cell proteomics is currently infeasible and bulk proteomics requires - inherently imperfect - cell synchronisation. Here, we developed a computational method to deconvolve cell cycle-dependent protein concentration dynamics and applied it to new budding yeast bulk proteome data. Key to this method was a yeast population model, parameterised with experimental cell cycle progression and volume growth data, for quantifying the desynchronisation in sampled populations. We performed deconvolution on 3272 proteins, using cross-validation to determine regularisation parameters, and identified 539 proteins with cell cycle-dependent dynamics. Many of these dynamics were consistent with known yeast biology and dynamic proteins were enriched for several metabolic process, extending previous observations and supporting the emerging picture of metabolic activity as varying substantially over cell cycle phases. We consider the generated cell cycle-resolved budding yeast proteome data a key resource.

Journal Article↗

Microwave-enhanced ink staining for fast and sensitive protein quantification in proteomic studies.

A novel microwave-enhanced ink staining method was developed for rapid and sensitive estimation of protein content in sample buffers containing chaotropes, dyes, detergents, and reducing agents. Dye-based Blue-Black ink was used to quantitatively visualize proteins spotted on a nitrocellulose membrane. The total staining time was greatly reduced to 3 min by brief exposure to microwave radiation. The stained membrane was washed with distilled water, baked in a microwave oven for complete desiccation, transparentized with mineral oil, and documented by a desktop scanner or densitometer. Only 1 microL of protein sample (protein solubilized in SDS-PAGE sample buffer or IEF rehydration buffer) was used for protein spotting. The novel solid-phase protein assay gives a 500-fold dynamic range from 19.5 to 10000 ng/microL and can be scaled up for high-throughput protein quantification analysis. The fast, sensitive and low-cost microwave-enhanced ink staining procedure is ideal for protein quantification in proteomic analysis.

Animals↗

Proteomic discovery analysis of quantitatively assessed emphysema in the general population. The MESA Lung Study.

BACKGROUND: Pulmonary emphysema occurs frequently in older adults, often without airflow limitation. Its presence predicts symptoms, respiratory hospitalizations and deaths, and all-cause mortality. Proteomics may provide further insights into emphysema pathogenesis and inform therapeutic targets. OBJECTIVE: We performed a proteomic discovery analysis of percent emphysema on computed tomography (CT) in a population-based, multiethnic sample from the Multi-Ethnic Study of Atherosclerosis (MESA) Lung Study. Replication was performed in two chronic obstructive pulmonary disease (COPD)-based studies, the SubPopulations and InteRmediate Outcome Measures in COPD Study (SPIROMICS) and the Genetic Epidemiology of COPD (COPDGene) Study. METHODS: MESA recruited participants from the general population in 2000-02. The MESA Lung Study performed full-lung CT scans in 2010-12. Percent emphysema was defined as the percentage of lung voxels&#x2009;<&#x2009;-950 Hounsfield units. Over 7,200 plasma aptamers were measured via SomaScan. Cross-sectional linear and least absolute shrinkage and selection operator (LASSO) regression models were adjusted for demographics, anthropometrics, smoking, renal function, and scanner parameters. Statistical significance was defined as a false discovery rate p-value&#x2009;<&#x2009;0.05. Gene Ontology (GO)/Reactome enrichment analyses were performed. LASSO-selected proteins' predictive performance was evaluated. RESULTS: Among 2,504 participants in the MESA Lung Study, mean age was 69.4&#xa0;years, 1,291 had ever smoked, and median percent emphysema-like lung was 1.4%. In total, 1,234 aptamers were significantly associated with percent emphysema in the MESA Lung Study, and 35 replicated in the SPIROMICS and COPDGene Studies. Novel associations included protein family with sequence similarity (FAM) 177A1, syntenin-2, ubiquitin carboxyl-terminal hydrolase 25, and uncharacterized protein C20orf173. Previously identified emphysema-associated proteins included soluble advanced glycosylation end product-specific receptor (sRAGE), protein S100-A12, high mobility group protein B1, and roundabout homolog 2. Enrichment analyses identified 40 GO biological processes, including chemokine production and regulation and cell-cell adhesion and regulation, and two Reactome pathways, including RAGE signaling. In tenfold cross-validation, novel proteins were largely retained by LASSO (R2&#x2009;=&#x2009;5.4%), improved overall model performance (R2&#x2009;=&#x2009;24.8%), and uniquely explained greater variance in percent emphysema. CONCLUSIONS: This analysis in a general population sample identified novel and previously characterized proteins whose functional roles were validated by GO/Reactome enriched pathways, offering new insights into emphysema pathophysiology and therapeutics.

Humans↗

Structural proteomics: methods in deriving protein structural information and issues in data management.

Structural proteomics is an emerging paradigm that is gaining importance in the post-genomic era as a valuable discipline to process the protein target information being deciphered. The field plays a crucial role in assigning function to sequenced proteins, defining pathways in which the targets are involved, and understanding structure-function relationships of the protein targets. A key component of this research sector is accessing the three-dimensional structures of protein targets by both experimental and theoretical methods. This then leads to the question of how to store, retrieve, and manipulate vast amounts of sequence (1-D) and structural (3-D) information in a relational format so that extensive data analysis can be achieved. We at SBI have addressed both of these fundamental requirements of structural proteomics. We have developed an extensive collection of three-dimensional protein structures from sequence data and have implemented a relational architecture for data management. In this article we will discuss our approaches to structural proteomics and the tools that life science researchers can use in their discovery efforts.

Computational Biology↗

Experimental standards for high-throughput proteomics.

Proteome analysis, utilizing high-throughput proteomics approaches, involves studying proteins that a whole organism (or specific tissue or cellular compartment) expresses under certain conditions. Intrinsic difficulties of these studies, as well as the enormous volumes of data they typically produce, make the proteome analysis and interpretation very difficult. As with any high-throughput approach, proteomics experiments should be carefully designed, analyzed, and verified. In addition to computational standards,experimental standards--simple and complex mixtures of known proteins--for high-throughput proteomics have to be developed and utilized. This article discusses such experimental standards and their implementations.

Animals↗

Assessing protein co-evolution in the context of the tree of life assists in the prediction of the interactome.

The identification of the whole set of protein interactions taking place in an organism is one of the main tasks in genomics, proteomics and systems biology. One of the computational techniques used by many investigators for studying and predicting protein interactions is the comparison of evolutionary histories (phylogenetic trees), under the hypothesis that interacting proteins would be subject to a similar evolutionary pressure resulting in a similar topology of the corresponding trees. Here, we present a new approach to predict protein interactions from phylogenetic trees, which incorporates information on the overall evolutionary histories of the species (i.e. the canonical "tree of life") in order to correct by the expected background similarity due to the underlying speciation events. We test the new approach in the largest set of annotated interacting proteins for Escherichia coli. This assessment of co-evolution in the context of the tree of life leads to a highly significant improvement (P(N) by sign test approximately 10E-6) in predicting interaction partners with respect to the previous technique, which does not incorporate information on the overall speciation tree. For half of the proteins we found a real interactor among the 6.4% top scores, compared with the 16.5% by the previous method. We applied the new method to the whole E.coli proteome and propose functions for some hypothetical proteins based on their predicted interactors. The new approach allows us also to detect non-canonical evolutionary events, in particular horizontal gene transfers. We also show that taking into account these non-canonical evolutionary events when assessing the similarity between evolutionary trees improves the performance of the method predicting interactions.

Algorithms↗