Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Computational methods for alternative splicing prediction.

The fact that a large majority of mammalian genes are subject to alternative splicing indicates that this phenomenon represents a major mechanism for increasing proteome complexity. Here, we provide an overview of current methods for the computational prediction of alternative splicing based on the alignment of genome and transcript sequences. Specific features and limitations of different approaches and software are discussed, particularly those affecting prediction accuracy and assembly of alternative transcripts.

Algorithms↗

Protein stains for proteomic applications: which, when, why?

This review recollects literature data on sensitivity and dynamic range for the most commonly used colorimetric and fluorescent dyes for general protein staining, and summarizes procedures for the most common PTM-specific detection methods. It also compiles some important points to be considered in imaging and evaluation. In addition to theoretical considerations, examples are provided to illustrate differential staining of specific proteins with different detection methods. This includes a large body of original data on the comparative evaluation of several pre- and post-electrophoresis stains used in parallel on a single specimen, horse serum run in 2-DE (IPG-DALT). A number of proteins/protein spots are found to be over- or under-revealed with some of the staining procedures.

Colorimetry↗

Megavariate data analysis of mass spectrometric proteomics data using latent variable projection method.

There are many data mining techniques for processing and general learning of multivariate data. However, we believe the wavelet transformation and latent variable projection method are particularly useful for spectroscopic and chromatographic data. Projection based methods are designed to handle hugely multivariate nature of such data effectively. For the actual analysis of the data we have used latent variable projection methods such as principal component analysis (PCA) and partial least squares projection to latent structures based discriminant analysis (PLS-DA) to analyze the raw data presented to the participants of the First Duke Proteomics Data Mining Conference. PCA was used to solve problem #1 (clustering problem) and the PLS-DA was used to solve problem #2 (classification problem). The idea of internal and external cross-validation was used to validate the model obtained from the classification analysis. The simple two-component PLS-DA model obtained from the analysis performed well. The model has completely separated the two groups from all the data. The same model applied on two-thirds of the data showed good performance by external validation with independent test set of remaining 13 specimens obtained by setting aside the spectra of every third specimen (accuracy of 85%).

Artificial Intelligence↗

Transcriptomic and proteomic characterization of the Fur modulon in the metal-reducing bacterium Shewanella oneidensis.

The availability of the complete genome sequence for Shewanella oneidensis MR-1 has permitted a comprehensive characterization of the ferric uptake regulator (Fur) modulon in this dissimilatory metal-reducing bacterium. We have employed targeted gene mutagenesis, DNA microarrays, proteomic analysis using liquid chromatography-mass spectrometry, and computational motif discovery tools to define the S. oneidensis Fur regulon. Using this integrated approach, we identified nine probable operons (containing 24 genes) and 15 individual open reading frames (ORFs), either with unknown functions or encoding products annotated as transport or binding proteins, that are predicted to be direct targets of Fur-mediated repression. This study suggested, for the first time, possible roles for four operons and eight ORFs with unknown functions in iron metabolism or iron transport-related functions. Proteomic analysis clearly identified a number of transporters, binding proteins, and receptors related to iron uptake that were up-regulated in response to a fur deletion and verified the expression of nine genes originally annotated as pseudogenes. Comparison of the transcriptome and proteome data revealed strong correlation for genes shown to be undergoing large changes at the transcript level. A number of genes encoding components of the electron transport system were also differentially expressed in a fur deletion mutant. The gene omcA (SO1779), which encodes a decaheme cytochrome c, exhibited significant decreases in both mRNA and protein abundance in the fur mutant and possessed a strong candidate Fur-binding site in its upstream region, thus suggesting that omcA may be a direct target of Fur activation.

Bacterial Proteins↗

Model building and model checking for biochemical processes.

A central claim of computational systems biology is that, by drawing on mathematical approaches developed in the context of dynamic systems, kinetic analysis, computational theory and logic, it is possible to create powerful simulation, analysis, and reasoning tools for working biologists to decipher existing data, devise new experiments, and ultimately to understand functional properties of genomes, proteomes, cells, organs, and organisms. In this article, a novel computational tool is described that achieves many of the goals of this new discipline. The novelty of this system involves an automaton-based semantics of the temporal evolution of complex biochemical reactions starting from the representation given as a set of differential equations. The related tools also provide ability to qualitatively reason about the systems using a propositional temporal logic that can express an ordered sequence of events succinctly and unambiguously. The implementation of mathematical and computational models in the Simpathica and XSSYS systems is described briefly. Several example applications of these systems to cellular and biochemical processes are presented: the two most prominent are Leibler et al.'s repressilator (an artificial synthesized oscillatory network), and Curto- Voit-Sorribas-Cascante's purine metabolism reaction model.

Biochemical Phenomena↗

Quality control metrics for LC-MS feature detection tools demonstrated on Saccharomyces cerevisiae proteomic profiles.

Quantitative proteomic profiling using liquid chromatography-mass spectrometry is emerging as an important tool for biomarker discovery, prompting development of algorithms for high-throughput peptide feature detection in complex samples. However, neither annotated standard data sets nor quality control metrics currently exist for assessing the validity of feature detection algorithms. We propose a quality control metric, Mass Deviance, for assessing the accuracy of feature detection tools. Because the Mass Deviance metric is derived from the natural distribution of peptide masses, it is machine- and proteome-independent and enables assessment of feature detection tools in the absence of completely annotated data sets. We validate the use of Mass Deviance with a second, independent metric that is based on isotopic distributions, demonstrating that we can use Mass Deviance to identify aberrant features with high accuracy. We then demonstrate the use of independent metrics in tandem as a robust way to evaluate the performance of peptide feature detection algorithms. This work is done on complex LC-MS profiles of Saccharomyces cerevisiae which present a significant challenge to peptide feature detection algorithms.

Algorithms↗

The evolution of controlled multitasked gene networks: the role of introns and other noncoding RNAs in the development of complex organisms.

Eukaryotic phenotypic diversity arises from multitasking of a core proteome of limited size. Multitasking is routine in computers, as well as in other sophisticated information systems, and requires multiple inputs and outputs to control and integrate network activity. Higher eukaryotes have a mosaic gene structure with a dual output, mRNA (protein-coding) sequences and introns, which are released from the pre-mRNA by posttranscriptional processing. Introns have been enormously successful as a class of sequences and comprise up to 95% of the primary transcripts of protein-coding genes in mammals. In addition, many other transcripts (perhaps more than half) do not encode proteins at all, but appear both to be developmentally regulated and to have genetic function. We suggest that these RNAs (eRNAs) have evolved to function as endogenous network control molecules which enable direct gene-gene communication and multitasking of eukaryotic genomes. Analysis of a range of complex genetic phenomena in which RNA is involved or implicated, including co-suppression, transgene silencing, RNA interference, imprinting, methylation, and transvection, suggests that a higher-order regulatory system based on RNA signals operates in the higher eukaryotes and involves chromatin remodeling as well as other RNA-DNA, RNA-RNA, and RNA-protein interactions. The evolution of densely connected gene networks would be expected to result in a relatively stable core proteome due to the multiple reuse of components, implying that cellular differentiation and phenotypic variation in the higher eukaryotes results primarily from variation in the control architecture. Thus, network integration and multitasking using trans-acting RNA molecules produced in parallel with protein-coding sequences may underpin both the evolution of developmentally sophisticated multicellular organisms and the rapid expansion of phenotypic complexity into uncontested environments such as those initiated in the Cambrian radiation and those seen after major extinction events.

Animals↗

Using models of the myocyte for functional interpretation of cardiac proteomic data.

There has been significant progress towards the development of highly integrative computational models of the cardiac myocyte over the past decade. Models now incorporate descriptions of voltage-gated ionic currents and membrane transporters, mechanisms of calcium-induced calcium release and intracellular calcium cycling, mitochondrial ATP production and its coupling to energy-requiring membrane transport processes and mechanisms of force generation. There is an extensive literature documenting both the reconstructive and predictive abilities of these models and there is no question that an interplay between quantitative modelling and experimental investigation has become a central component of modern cardiovascular research. As data regarding the cardiovascular proteome in both health and disease emerge, integrative models of the myocyte are becoming useful tools for interpreting the functional significance of changes in protein expression and post-translational modifications (PTMs). Data of particular importance include information on: (a) changes of expressed protein level, (b) changes of protein PTMs, (c) protein localization, and (d) protein-protein interactions, as it is often possible to incorporate and interpret the functional significance of such findings using computational models. We provide two examples of how models may be used in this fashion. In the first example, we show how information on altered expression of the sarcoplasmic reticulum Ca2+-ATPase, when interpreted through the use of a computational model, has provided key insights into fundamental mechanisms regulating cardiac action potential duration. In the second example, we show how information on the effects of phosphorylation of L-type Ca2+ channels, when interpreted through the use of a model, provides insights on how this post-translational modification alters the properties of excitation-contraction coupling and risk for arrhythmia.

Animals↗

Exploring the proteome of Plasmodium.

With the entire genomic sequence of several species of Plasmodium soon to be available, researchers are now focusing on methods to study gene and protein expression at the whole organism level. Traditional methods of characterising and identifying large numbers of proteins from a complex protein mixture have relied predominantly on two-dimensional gel electrophoresis combined with N-terminal sequencing or mass spectrometry of individually prepared proteins. New proteomics methods are now available that are based on resolving small peptides derived from complex protein mixtures by high-resolution liquid chromatography and directly identifying them by tandem mass spectrometry (LC/LC/MS/MS) and sophisticated computer search algorithms against whole genome sequence databases. These newer proteomic methods have the potential to accelerate the reproducible identification of large numbers of proteins from various life cycle stages of Plasmodium and may help to better understand parasite biology and lead to the identification of new targets of vaccines and drugs.

Animals↗

Challenges to be faced in the reconstruction of metabolic networks from public databases.

In the post-genomic era, the biochemical information for individual compounds, enzymes, reactions to be found within named organisms has become readily available. The well-known KEGG and BioCyc databases provide a comprehensive catalogue for this information and have thereby substantially aided the scientific community. Using these databases, the complement of enzymes present in a given organism can be determined and, in principle, used to reconstruct the metabolic network. However, such reconstructed networks contain numerous properties contradicting biological expectation. The metabolic networks for a number of organisms are reconstructed from KEGG and BioCyc databases, and features of these networks are related to properties of their originating database.

Algorithms↗

The biology of the post-genomic era: the proteomics.

The complete identification of coding sequences in a number of species has led to announce the beginning of the post-genomic era, new tools have become available to study complex phenomena in biological systems. Rapid advances in genomic sequencing and bioinformatics have established the field of genomics to investigate thousands genes' activity through mRNA display. However, recent studies have demonstrated a lack of correlation between the transcriptional profiles and the actual protein levels in cells, so investigation of the expressed part of the genome is also required to link genomic data to biological function. It is possible that evolutional development occured by increasing complexity of regulation processes at the level of RNA and protein molecules instead of simple increase in gene number, so investigation of proteins and protein complexes became important fields of our post-genomic era. High-resolution two-dimensional gels combined with sensitive mass spectrometry can reveal virtually all proteins present in cells opening new insights into functions of cells, tissues and whole organisms.

Animals↗

Bioinformatics strategies for proteomic profiling.

Clinical proteomics is an emerging field that involves the analysis of protein expression profiles of clinical samples for de novo discovery of disease-associated biomarkers and for gaining insight into the biology of disease processes. Mass spectrometry represents an important set of technologies for protein expression measurement. Among them, surface-enhanced laser desorption/ionization time-of-flight mass spectrometry (SELDI TOF-MS), because of its high throughput and on-chip sample processing capability, has become a popular tool for clinical proteomics. Bioinformatics plays a critical role in the analysis of SELDI data, and therefore, it is important to understand the issues associated with the analysis of clinical proteomic data. In this review, we discuss such issues and the bioinformatics strategies used for proteomic profiling.

Computational Biology↗

Systematic identification of immunoreceptor tyrosine-based inhibitory motifs in the human proteome.

Immunoreceptor tyrosine-based inhibitory motifs (ITIMs) are short sequences of the consensus (ILV)-x-x-Y-x-(LV) in the cytoplasmic tail of immune receptors. The phosphorylation of tyrosines in ITIMs is known to be an important signalling mechanism regulating the activation of immune cells. The shortness of the motif makes it difficult to predict ITIMs in large protein databases. Simple pattern searches find ITIMs in approximately 30% of the protein sequences in the RefSeq database. The majority are false positive predictions. We propose a new database search strategy for ITIM-bearing transmembrane receptors based on the use of sequence context, i.e. the predictions of signal peptides, transmembrane helices (TMs) and protein domains. Our new protocol allowed us to narrow down the number of potential human ITIM receptors to 109 proteins (0.7% of RefPep). Of these, 36 have been described as ITIM receptors in the literature before. Many ITIMs are conserved between orthologous human and mouse proteins which represent novel ITIM receptor candidates. Publicly available DNA array expression data revealed that ITIM receptors are not exclusively expressed in blood cells. We hypothesise that ITIM signalling is not restricted to immune cells, but also functions in diverse solid organs of mouse and man.

Computational Biology↗

Neural network prediction of peptide separation in strong anion exchange chromatography.

MOTIVATION: The still emerging combination of technologies that enable description and characterization of all expressed proteins in a biological system is known as proteomics. Although many separation and analysis technologies have been employed in proteomics, it remains a challenge to predict peptide behavior during separation processes. New informatics tools are needed to model the experimental analysis method that will allow scientists to predict peptide separation and assist with required data mining steps, such as protein identification. RESULTS: We developed a software package to predict the separation of peptides in strong anion exchange (SAX) chromatography using artificial neural network based pattern classification techniques. A multi-layer perceptron is used as a pattern classifier and it is designed with feature vectors extracted from the peptides so that the classification error is minimized. A genetic algorithm is employed to train the neural network. The developed system was tested using 14 protein digests, and the sensitivity analysis was carried out to investigate the significance of each feature. AVAILABILITY: The software and testing results can be downloaded from ftp://ftp.bbc.purdue.edu.

Algorithms↗

An enhanced Java graph applet interface for visualizing interactomes.

UNLABELLED: We have developed several new navigation features for a Java graph applet previously released for visualizing protein-protein interactions. This graph viewer can be used to navigate any molecular interactome dataset. We have successfully implemented this tool for exploring protein networks stored in the Bioverse interaction database. AVAILABILITY: http://bioverse.compbio.washington.edu/viewer CONTACT: ram@compbio.washington.edu.

Animals↗

Preprocessing of two-dimensional gel electrophoresis images.

Proteomics produces a huge amount of two-dimensional gel electrophoresis images. Their analysis can yield a lot of information concerning proteins responsible for different diseases or new unidentified proteins. However, an automatic analysis of such images requires an efficient tool for reducing noise in images. This allows proper detection of the spots' borders, which is important in protein quantification (as the spots' areas are used to determine the amounts of protein present in an analyzed mixture). Also in the feature-based matching methods the detected features (spots) can be described by additional attributes, such as area or shape. In our study, a comparison of different methods of noise reduction is performed in order to find out a method best suited for reducing noise in gel images. Among the compared methods there are the classical methods of linear filtering, e.g., the mean and Gaussian filtering, the nonlinear method, i.e., median filtering, and also the methods better suited for processing of nonstationary signals, such as spatially adaptive linear filtering and filtering in the wavelet domain. The best results are obtained by filtering of gel images in the wavelet domain, using the BayesThresh method of threshold value determination.

Electrophoresis, Gel, Two-Dimensional↗

Managing core resources for genomics and proteomics.

Recent years have seen an explosive growth in biological data, which is often not published in a conventional sense but rather deposited in a database. This trend and the need for computational analyses of the data make databases essential tools for biological research. Data from a variety of sources, covering a wide range of biological information, are stored in different, often quite specialized, databases. The provision of such databases as useful resources for the scientific community is a demanding task since the data not only have to be stored in a consistent way, but also have to be easily accessible and highly integrated with other databases. Furthermore, it is necessary to provide users with effective tools to search the databases and to analyze the data. At the European Bioinformatics Institute (EBI), we develop and maintain a number of biological databases and provide a variety of bioinformatics tools to facilitate database and similarity searches and data analysis. In this review, we will provide examples of the core resources maintained at the EBI and summarize important issues of database management of such resources.

Computational Biology↗