Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

BIOCHAM: an environment for modeling biological systems and formalizing experimental knowledge.

UNLABELLED: BIOCHAM (the BIOCHemical Abstract Machine) is a software environment for modeling biochemical systems. It is based on two aspects: (1) the analysis and simulation of boolean, kinetic and stochastic models and (2) the formalization of biological properties in temporal logic. BIOCHAM provides tools and languages for describing protein networks with a simple and straightforward syntax, and for integrating biological properties into the model. It then becomes possible to analyze, query, verify and maintain the model with respect to those properties. For kinetic models, BIOCHAM can search for appropriate parameter values in order to reproduce a specific behavior observed in experiments and formalized in temporal logic. Coupled with other methods such as bifurcation diagrams, this search assists the modeler/biologist in the modeling process. AVAILABILITY: BIOCHAM (v. 2.5) is a free software available for download, with example models, at http://contraintes.inria.fr/BIOCHAM/.

Algorithms↗

Dietary protein-related changes in hepatic transcription correspond to modifications in hepatic protein expression in growing pigs.

In a previous investigation we showed by expression profiling based on transcription analysis using differential display RT-PCR (DDRT-PCR) and real-time RT-PCR that a soy protein diet (SPI) significantly changes the hepatic transcription pattern compared with a casein diet (CAS). The present study was conducted to determine whether the transcriptional modulation is translated into protein expression. The hepatic mRNA abundance of four genes (EP24.16, LC3, NPAP60L, RFC2) that showed diet-related expression in previous DDRT-PCR experiments was analyzed by real-time RT-PCR. Two pigs that showed the most prominent SPI-related changes of transcription and two casein-fed pigs were selected and their hepatic protein pattern was studied comparatively by two-dimensional gel electrophoresis and peptide mass fingerprinting. The two-dimensional protein gel electrophoresis revealed a predominant SPI-associated upregulation of protein expression that corresponded to the results of the mRNA study. Of 380 diet-related protein spots displayed, 215 appeared exclusively or enlarged in the two SPI pigs; 10 of 39 diet-related expressed protein spots extracted could be identified by peptide mass fingerprinting and database search. Compared with the transcriptomics approach, the proteomics approach led in part to the identification of the same diet-associated expressed molecules (plasminogen, trypsin, phospholipase A2, glutathione-S-transferase alpha, retinal binding protein) or at least molecules belonging to the same metabolic pathways (protein and amino acid metabolism, oxidative stress response, lipid metabolism). The present results at the proteome level confirm SPI-related increased oxidative stress response and significant effects on protein biosynthesis already observed at the transcriptome level.

Animals↗

Collecting and harvesting biological data: the GPCRDB and NucleaRDB information systems.

The amount of genomic and proteomic data that is entered each day into databases and the experimental literature is outstripping the ability of experimental scientists to keep pace. While generic databases derived from automated curation efforts are useful, most biological scientists tend to focus on a class or family of molecules and their biological impact. Consequently, there is a need for molecular class-specific or other specialized databases. Such databases collect and organize data around a single topic or class of molecules. If curated well, such systems are extremely useful as they allow experimental scientists to obtain a large portion of the available data most relevant to their needs from a single source. We are involved in the development of two such databases with substantial pharmacological relevance. These are the GPCRDB and NucleaRDB information systems, which collect and disseminate data related to G protein-coupled receptors and intra-nuclear hormone receptors, respectively. The GPCRDB was a pilot project aimed at building a generic molecular class-specific database capable of dealing with highly heterogeneous data. A first version of the GPCRDB project has been completed and it is routinely used by thousands of scientists. The NucleaRDB was started recently as an application of the concept for the generalization of this technology. The GPCRDB is available via the WWW at http://www.gpcr.org/7tm/ and the NucleaRDB at http://www.receptors.org/NR/.

Binding, Competitive↗

The Online Bioinformatics Resources Collection at the University of Pittsburgh Health Sciences Library System--a one-stop gateway to online bioinformatics databases and software tools.

To bridge the gap between the rising information needs of biological and medical researchers and the rapidly growing number of online bioinformatics resources, we have created the Online Bioinformatics Resources Collection (OBRC) at the Health Sciences Library System (HSLS) at the University of Pittsburgh. The OBRC, containing 1542 major online bioinformatics databases and software tools, was constructed using the HSLS content management system built on the Zope Web application server. To enhance the output of search results, we further implemented the Vivísimo Clustering Engine, which automatically organizes the search results into categories created dynamically based on the textual information of the retrieved records. As the largest online collection of its kind and the only one with advanced search results clustering, OBRC is aimed at becoming a one-stop guided information gateway to the major bioinformatics databases and software tools on the Web. OBRC is available at the University of Pittsburgh's HSLS Web site (http://www.hsls.pitt.edu/guides/genetics/obrc).

Computational Biology↗

Optimal classification of protein sequences and selection of representative sets from multiple alignments: application to homologous families and lessons for structural genomics.

Hierarchical classification is probably the most popular approach to group related proteins. However, there are a number of problems associated with its use for this purpose. One is that the resulting tree showing a nested sequence of groups may not be the most suitable representation of the data. Another is that visual inspection is the most common method to decide the most appropriate number of subsets from a tree. In fact, classification of proteins in general is bedevilled with the need for subjective thresholds to define group membership (e.g., 'significant' sequence identity for homologous families). Such arbitrariness is not only intellectually unsatisfying but also has important practical consequences. For instance, it hinders meaningful identification of protein targets for structural genomics. I describe an alternative approach to cluster related proteins without the need for an a priori threshold: one, through its use of dynamic programming, which is guaranteed to produce globally optimal solutions at all levels of partition granularity. Grouping proteins according to weights assigned to their aligned sequences makes it possible to delineate dynamically a 'core-periphery' structure within families. The 'core' of a protein family comprises the most typical sequences while the 'periphery' consists of the atypical ones. Further, a new sequence weighting scheme that combines the information in all the multiply aligned positions of an alignment in a novel way is put forward. Instead of averaging over all positions, this procedure takes into account directly the distribution of sequence variability along an alignment. The relationships between sequence weights and sequence identity are investigated for 168 families taken from HOMSTRAD, a database of protein structure alignments for homologous families. An exact solution is presented for the problem of how to select the most representative pair of sequences for a protein family. Extension of this approach by a greedy algorithm allows automatic identification of a minimal set of aligned sequences. The results of this analysis are available on the Web at http://mathbio.nimr.mrc.ac.uk/~amay.

Algorithms↗

Weighted-support vector machines for predicting membrane protein types based on pseudo-amino acid composition.

Membrane proteins are generally classified into the following five types: (1) type I membrane proteins, (2) type II membrane proteins, (3) multipass transmembrane proteins, (4) lipid chain-anchored membrane proteins and (5) GPI-anchored membrane proteins. Prediction of membrane protein types has become one of the growing hot topics in bioinformatics. Currently, we are facing two critical challenges in this area: first, how to take into account the extremely complicated sequence-order effects, and second, how to deal with the highly uneven sizes of the subsets in a training dataset. In this paper, stimulated by the concept of using the pseudo-amino acid composition to incorporate the sequence-order effects, the spectral analysis technique is introduced to represent the statistical sample of a protein. Based on such a framework, the weighted support vector machine (SVM) algorithm is applied. The new approach has remarkable power in dealing with the bias caused by the situation when one subset in the training dataset contains many more samples than the other. The new method is particularly useful when our focus is aimed at proteins belonging to small subsets. The results obtained by the self-consistency test, jackknife test and independent dataset test are encouraging, indicating that the current approach may serve as a powerful complementary tool to other existing methods for predicting the types of membrane proteins.

Algorithms↗

Kernel-based data fusion and its application to protein function prediction in yeast.

Kernel methods provide a principled framework in which to represent many types of data, including vectors, strings, trees and graphs. As such, these methods are useful for drawing inferences about biological phenomena. We describe a method for combining multiple kernel representations in an optimal fashion, by formulating the problem as a convex optimization problem that can be solved using semidefinite programming techniques. The method is applied to the problem of predicting yeast protein functional classifications using a support vector machine (SVM) trained on five types of data. For this problem, the new method performs better than a previously-described Markov random field method, and better than the SVM trained on any single type of data.

Algorithms↗

Semantic similarity measures as tools for exploring the gene ontology.

Many bioinformatics resources hold data in the form of sequences. Often this sequence data is associated with a large amount of annotation. In many cases this data has been hard to model, and has been represented as scientific natural language, which is not readily computationally amenable. The development of the Gene Ontology provides us with a more accessible representation of some of this data. However it is not clear how this data can best be searched, or queried. Recently we have adapted information content based measures for use with the Gene Ontology (GO). In this paper we present detailed investigation of the properties of these measures, and examine various properties of GO, which may have implications for its future design.

Classification↗

Proteome analysis. I. Gene products are where the biological action is.

Two-dimensional electrophoresis has rapidly become the method of choice for resolving complex mixtures of proteins. Since the technique was pioneered in 1975, 2-D gel methods have undergone a series of enhancements to optimize resolution and reproducibility. Recent improvements in the sensitivity of mass spectrometry have allowed the direct identification of polypeptides from 2-D gels by a procedure termed "mass profiling". In combination, these two techniques have made possible the characterization of the complete collection of gene products, or proteome, of an organism. Proteomes are increasingly being documented as interactive informational databases available on the World Wide Web (WWW). This availability of organismic global protein patterns will no doubt be an invaluable resource aiding the discovery of diagnostic and therapeutic disease markers.

Amino Acid Sequence↗

[Bioinformatics].

Explore the source record for details and available documents.

Computational Biology↗

Autumn 2005 Workshop of the Human Proteome Organisation Proteomics Standards Initiative (HUPO-PSI) Geneva, September, 4-6, 2005.

The autumn workshop of the Proteomics Standards Initiative of the Human Proteomics Organisation met to further advance the development of the existing standards in the fields of molecular interactions and mass spectrometry. In addition, new areas were addressed, in particular developing standards for the description and exchange of data from gel electrophoresis experiments. The General Proteomics Standards group is now working closely with the FuGE (Functional Genomics Experiment) efforts to define a general standard in which to encode data that will enable a systems biology approach to data analysis. Common to all these efforts is the field of protein modifications, and work has been initiated to establish an ontology in this field that can be used by both workers in the field of proteomics and the wider scientific community.

Databases, Protein↗

Characterization of the mouse brain proteome using global proteomic analysis complemented with cysteinyl-peptide enrichment.

We report a global proteomic approach for analyzing brain tissue and for the first time a comprehensive characterization of the whole mouse brain proteome. Preparation of the whole brain sample incorporated a highly efficient cysteinyl-peptide enrichment (CPE) technique to complement a global enzymatic digestion method. Both the global and the cysteinyl-enriched peptide samples were analyzed by SCX fractionation coupled with reversed phase LC-MS/MS analysis. A total of 48,328 different peptides were confidently identified (>98% confidence level), covering 7792 nonredundant proteins ( approximately 34% of the predicted mouse proteome). A total of 1564 and 1859 proteins were identified exclusively from the cysteinyl-peptide and the global peptide samples, respectively, corresponding to 25% and 31% improvements in proteome coverage compared to analysis of only the global peptide or cysteinyl-peptide samples. The identified proteins provide a broad representation of the mouse proteome with little bias evident due to protein pI, molecular weight, and/or cellular localization. Approximately 26% of the identified proteins with gene ontology (GO) annotations were membrane proteins, with 1447 proteins predicted to have transmembrane domains, and many of the membrane proteins were found to be involved in transport and cell signaling. The MS/MS spectrum count information for the identified proteins was used to provide a measure of relative protein abundances. The mouse brain peptide/protein database generated from this study represents the most comprehensive proteome coverage for the mammalian brain to date, and the basis for future quantitative brain proteomic studies using mouse models. The proteomic approach presented here may have broad applications for rapid proteomic analyses of various mouse models of human brain diseases.

Animals↗

Mass spectrometry of the human pituitary proteome: identification of selected proteins.

The field of proteomics involves the combined application of advanced separation techniques, mass spectrometry, and bioinformatics tools to characterize proteins in complex biological mixtures. Here we report the identification of nine proteins from the human pituitary proteome, using the proteomics approach. The pituitary proteins were separated by two-dimensional electrophoresis, and were visualized by silver staining. The proteins of interest were subjected to in-gel digestion with trypsin, and the masses of the resulting peptides were determined by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry. This tryptic mass map was used to identify the proteins through a search of a protein-sequence database. The identified proteins include important hormones, and enzymes with various catalytic activities. These proteins will be used to construct a two-dimensional reference database of the human pituitary. This database will be employed to study changes in the pituitary proteome that are associated with the formation of pituitary tumors.

Databases, Factual↗

Proteomic approaches to characterize protein modifications: new tools to study the effects of environmental exposures.

Proteomics is the study of proteomes, which are the collections of proteins expressed in cells. Whereas genomes are essentially invariant in different cells in an organism, proteomes vary from cell to cell, with time and as a function of environmental stimuli and stress. The integration of new mass spectrometry (MS) methods, data analysis algorithms, and information from databases of protein and gene sequences has enabled the characterization of proteomes. Many environmental agents directly or indirectly generate reactive electrophiles that covalently modify proteins. Although considerable evidence supports a key role for protein adducts in adverse effects of chemicals, limitations in analytical technology have slowed progress in this area. New applications of liquid chromatography-tandem mass spectrometry (LC-MS-MS) now offer the potential to identify protein targets of reactive electrophiles and to map adducts at the level of amino acid sequence. Use of the data-analysis tools Sequest and SALSA (Scoring Algorithm for Spectral Analysis) together with LC-MS-MS analyses of protein digests enables the identification of modified forms of proteins in a sample. These approaches can map adducts to specific amino acids in protein targets and are being adapted to searches for protein adducts in complex proteomes. These tools will facilitate the identification of new biomarkers of chemical exposure and studies of mechanisms by which protein modifications contribute to the adverse effects of environmental exposures.

Algorithms↗