Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Novel approach for peptide quantitation and sequencing based on 15N and 13C metabolic labeling.

Here we describe a method for protein identification and quantification using stable isotopes via in vivo metabolic labeling of the hyperthermophilic crenarchaeon Sulfolobus solfataricus. Stable isotope labeling for quantitative proteomics is becoming increasingly popular; however, its usefulness in protein identification has not been fully exploited. We use both 15N and 13C labeling to create three different versions of the same peptide, corresponding to the unlabeled, 15N and 13C labeled versions. The peptide then appears as three different peaks in a TOF-MS scan and three corresponding sets of MS/MS spectra are obtained. With this information, the elemental carbon and nitrogen compositions for each peptide and each fragment can be calculated. When this is used as a constraint in database searching and/or de novo sequencing, the confidence of a match is increased (for an example intact peptide from 34 choices to 1). This makes the method a useful proteomic tool for both sequenced and unsequenced organisms. Furthermore, it allows for accurate protein quantitation (standard deviations over >4 peptides per protein were within 10%) of three phenotypes in one MS experiment. Abundances for each peptide are calculated by determining the relative areas of each of the three peaks in the TOF-MS spectrum.

Amino Acid Sequence↗

Identification of differentially expressed proteins of gamma-ray irradiated rat intestinal epithelial IEC-6 cells by two-dimensional gel electrophoresis and matrix-assisted laser desorption/ionisation-time of flight mass spectrometry.

To identify proteins involved in the processes of cellular and molecular response to radiation damage repair in intestinal epithelial IEC-6 cells, we comparatively analyzed the proteome of irradiated IEC-6 cells with that of normal cells. A series of methods were used, including two-dimensional gel electrophoresis (Z-DE), PDQuest software analysis of 2-DE gels, peptide mass fingerprinting based on matrix-assisted laser desorption/ionisation-time of flight-mass spectrometry (MALDI-TOF-MS), and Swiss-Prot database searching, to separate and identify differentially expressed proteins. Western blotting and reverse transcriptase polymerase chain reaction (RT-PCR) were used to validate the differentially expressed proteins. Image analysis revealed that averages of 608 +/- 39 and 595 +/- 31 protein spots were detected in normal and irradiated IEC-6 cells, respectively. Sixteen differential protein spots were isolated from gels, and measured with MALDI-TOF-MS. A total of 14 spots yielded good spectra, and 11 spots matched with known proteins after database searching. These proteins were mainly involved in anti-oxidation, metabolism, and protein post-translational processes. Western blotting confirmed that stress-70 protein was down-regulated by gamma-irradiation. Up-regulation of ERP29 was confirmed by RT-PCR, indicating that it is involved in ionizing radiation. The clues provided by the comparative proteome strategy utilized here will shed light on molecular mechanisms of radiation damage repair in intestinal epithelial cells.

Animals↗

Design of a data model for developing laboratory information management and analysis systems for protein production.

Data management has emerged as one of the central issues in the high-throughput processes of taking a protein target sequence through to a protein sample. To simplify this task, and following extensive consultation with the international structural genomics community, we describe here a model of the data related to protein production. The model is suitable for both large and small facilities for use in tracking samples, experiments, and results through the many procedures involved. The model is described in Unified Modeling Language (UML). In addition, we present relational database schemas derived from the UML. These relational schemas are already in use in a number of data management projects.

Algorithms↗

Bioinformatics.

Explore the source record for details and available documents.

Cells↗

Data mining for protein-protein interactions in invertebrate model organisms.

Well-annotated genome databases are available for many invertebrate species, notably the fruitfly, Drosophila melanogaster, and the nematode, Caenorhabditis elegans. An adequate interpretation of this information at the biological level requires the exploration of the interactions between the gene products. Knowledge of protein interactions and the components of cell signalling pathways in the fly and worm are particularly valuable as hypotheses can be rapidly tested using the powerful genetic toolkits available. Invertebrates offer additional experimental advantages when attempting to characterise protein-protein interactions (PPIs). Their relatively small genome size compared to mammals helps to reduce missed interactions due to redundancy, and their function can be addressed using forward (mutants) and reverse (RNA interference) genetics. However, the researcher looking for evidence of PPIs for a protein of interest is faced with the challenge of extracting interaction data from sources that are highly varied, such as the results of microarray experiments in the unstructured text of research papers. This challenge is greatly reduced by a range of public databases of curated information, as well as publicly available, enhanced search engines, which can provide either direct experimental evidence for a PPI, or valuable clues for generating new hypotheses.

Animals↗

Towards building the silicon cell: a modular approach.

Systems Biology aims to understand quantitatively how properties of biological systems can be understood as functions of the characteristics of, and interactions between their macromolecular components. Whereas, traditional biochemistry focused on isolation and characterization of cellular components, the challenge for Systems Biology lies in integration of this knowledge and the knowledge about molecular interactions. Computer models play an important role in this integration. We here discuss an approach with which we aim to link kinetic models on small parts of metabolism together, so as to form detailed kinetic models of larger chunks of metabolism, and ultimately of the entire living cell. Specifically, we will discuss techniques that can be used to model a sub-network in isolation of a larger network of which it is a part, while still maintaining the dynamics of the larger complete network. We will start by outlining the JWS online system, the silicon cell project, and the type of models we propose. JWS online is a model repository, which can be used for the storage, simulation and analysis of kinetic models. We advocate to integrate a top-down approach, where measurements on the complete system are used to derive fluxes in a detailed structural model, with a bottom-up approach, consisting of the integration of molecular mechanism-based detailed kinetic models into the structural model.

Animals↗

The complement of enzymatic sets in different species.

We present here a comprehensive analysis of the complement of enzymes in a large variety of species. As enzymes are a relatively conserved group there are several classification systems available that are common to all species and link a protein sequence to an enzymatic function. Enzymes are therefore an ideal functional group to study the relationship between sequence expansion, functional divergence and phenotypic changes. By using information retrieved from the well annotated SWISS-PROT database together with sequence information from a variety of fully sequenced genomes and information from the EC functional scheme we have aimed here to estimate the fraction of enzymes in genomes, to determine the extent of their functional redundancy in different domains of life and to identify functional innovations and lineage specific expansions in the metazoa lineage. We found that prokaryote and eukaryote species differ both in the fraction of enzymes in their genomes and in the pattern of expansion of their enzymatic sets. We observe an increase in functional redundancy accompanying an increase in species complexity. A quantitative assessment was performed in order to determine the degree of functional redundancy in different species. Finally, we report a massive expansion in the number of mammalian enzymes involved in signalling and degradation.

Animals↗

Parallel tandem: a program for parallel processing of tandem mass spectra using PVM or MPI and X!Tandem.

A method for the rapid correlation of tandem mass spectra to a list of protein sequences in a database has been developed. The combination of the fast and accurate computational search algorithm, X!Tandem, and a Linux cluster parallel computing environment with PVM or MPI, significantly reduces the time required to perform the correlation of tandem mass spectra to protein sequences in a database. A file of tandem mass spectra is divided into a specified number of files, each containing an equal number of the spectra from the larger file. These files are then searched in parallel against a protein sequence database. The results of each parallel output file are collated into one file for viewing through a web interface. Thousands of spectra can be searched in an accurate, practical, and time effective manner. The source code for running Parallel Tandem utilizing either PVM or MPI on Linux operating system is available from http://www.thegpm.org. This source code is made available under Artistic License from the authors.

Algorithms↗

Comparison of probability and likelihood models for peptide identification from tandem mass spectrometry data.

We evaluate statistical models used in two-hypothesis tests for identifying peptides from tandem mass spectrometry data. The null hypothesis H(0), that a peptide matches a spectrum by chance, requires information on the probability of by-chance matches between peptide fragments and peaks in the spectrum. Likewise, the alternate hypothesis H(A), that the spectrum is due to a particular peptide, requires probabilities that the peptide fragments would indeed be observed if it was the causative agent. We compare models for these probabilities by determining the identification rates produced by the models using an independent data set. The initial models use different probabilities depending on fragment ion type, but uniform probabilities for each ion type across all of the labile bonds along the backbone. More sophisticated models for probabilities under both H(A) and H(0) are introduced that do not assume uniform probabilities for each ion type. In addition, the performance of these models using a standard likelihood model is compared to an information theory approach derived from the likelihood model. Also, a simple but effective model for incorporating peak intensities is described. Finally, a support-vector machine is used to discriminate between correct and incorrect identifications based on multiple characteristics of the scoring functions. The results are shown to reduce the misidentification rate significantly when compared to a benchmark cross-correlation based approach.

Databases, Protein↗

Expressed peptide tags: an additional layer of data for genome annotation.

While genome sequencing is becoming ever more routine, genome annotation remains a challenging process. Identification of the coding sequences within the genomic milieu presents a tremendous challenge, especially for eukaryotes with their complex gene architectures. Here, we present a method to assist the annotation process through the use of proteomic data and bioinformatics. Mass spectra of digested protein preparations of the organism of interest were acquired and searched against a protein database created by a six-frame translation of the genome. The identified peptides were mapped back to the genome, compared to the current annotation, and then categorized as supporting or extending the current genome annotation. We named the classified peptides Expressed Peptide Tags (EPTs). The well-annotated bacterium Rhodopseudomonas palustris was used as a control for the method and showed a high degree of correlation between EPT mapping and the current annotation, with 86% of the EPTs confirming existing gene calls and less than 1% of the EPTs expanding on the current annotation. The eukaryotic plant pathogens Phytophthora ramorum and Phytophthora sojae, whose genomes have been recently sequenced and are much less well-annotated, were also subjected to this method. A series of algorithmic steps were taken to increase the confidence of EPT identification for these organisms, including generation of smaller subdatabases to be searched against, and definition of EPT criteria that accommodates the more complex eukaryotic gene architecture. As expected, the analysis of the Phytophthora species showed less correlation between EPT mapping and their current annotation. While approximately 76% of Phytophthora EPTs supported the current annotation, a portion of them (7.7% and 12.9% for P. ramorum and P. sojae, respectively) suggested modification to current gene calls or identified novel genes that were missed by the current genome annotation of these organisms.

Amino Acid Sequence↗

Completeness in structural genomics.

Structural genomics has the goal of obtaining useful, three-dimensional models of all proteins by a combination of experimental structure determination and comparative model building. We evaluate different strategies for optimizing information return on effort. The strategy that maximizes structural coverage requires about seven times fewer structure determinations compared with the strategy in which targets are selected at random. With a choice of reasonable model quality and the goal of 90% coverage, we extrapolate the estimate of the total effort of structural genomics. It would take approximately 16,000 carefully selected structure determinations to construct useful atomic models for the vast majority of all proteins. In practice, unless there is global coordination of target selection, the total effort will likely increase by a factor of three. The task can be accomplished within a decade provided that selection of targets is highly coordinated and significant funding is available.

Algorithms↗

Unraveling protein interaction networks with near-optimal efficiency.

The functional characterization of genes and their gene products is the main challenge of the genomic era. Examining interaction information for every gene product is a direct way to assemble the jigsaw puzzle of proteins into a functional map. Here we demonstrate a method in which the information gained from pull-down experiments, in which single proteins act as baits to detect interactions with other proteins, is maximized by using a network-based strategy to select the baits. Because of the scale-free distribution of protein interaction networks, we were able to obtain fast coverage by focusing on highly connected nodes (hubs) first. Unfortunately, locating hubs requires prior global information about the network one is trying to unravel. Here, we present an optimized 'pay-as-you-go' strategy that identifies highly connected nodes using only local information that is collected as successive pull-down experiments are performed. Using this strategy, we estimate that 90% of the human interactome can be covered by 10,000 pull-down experiments, with 50% of the interactions confirmed by reciprocal pull-down experiments.

Algorithms↗

Chromosomal toxin-antitoxin systems in Pseudomonas putida are rather selfish than beneficial.

Chromosomal toxin-antitoxin (TA) systems are widespread genetic elements among bacteria, yet, despite extensive studies in the last decade, their biological importance remains ambivalent. The ability of TA-encoded toxins to affect stress tolerance when overexpressed supports the hypothesis of TA systems being associated with stress adaptation. However, the deletion of TA genes has usually no effects on stress tolerance, supporting the selfish elements hypothesis. Here, we aimed to evaluate the cost and benefits of chromosomal TA systems to Pseudomonas putida. We show that multiple TA systems do not confer fitness benefits to this bacterium as deletion of 13 TA loci does not influence stress tolerance, persistence or biofilm formation. Our results instead show that TA loci are costly and decrease the competitive fitness of P. putida. Still, the cost of multiple TA systems is low and detectable in certain conditions only. Construction of antitoxin deletion strains showed that only five TA systems code for toxic proteins, while other TA loci have evolved towards reduced toxicity and encode non-toxic or moderately potent proteins. Analysis of P. putida TA systems' homologs among fully sequenced Pseudomonads suggests that the TA loci have been subjected to purifying selection and that TA systems spread among bacteria by horizontal gene transfer.

Anti-Bacterial Agents↗

Support vector machine-based method for subcellular localization of human proteins using amino acid compositions, their order, and similarity search.

Here we report a systematic approach for predicting subcellular localization (cytoplasm, mitochondrial, nuclear, and plasma membrane) of human proteins. First, support vector machine (SVM)-based modules for predicting subcellular localization using traditional amino acid and dipeptide (i + 1) composition achieved overall accuracy of 76.6 and 77.8%, respectively. PSI-BLAST, when carried out using a similarity-based search against a nonredundant data base of experimentally annotated proteins, yielded 73.3% accuracy. To gain further insight, a hybrid module (hybrid1) was developed based on amino acid composition, dipeptide composition, and similarity information and attained better accuracy of 84.9%. In addition, SVM modules based on a different higher order dipeptide i.e. i + 2, i + 3, and i + 4 were also constructed for the prediction of subcellular localization of human proteins, and overall accuracy of 79.7, 77.5, and 77.1% was accomplished, respectively. Furthermore, another SVM module hybrid2 was developed using traditional dipeptide (i + 1) and higher order dipeptide (i + 2, i + 3, and i + 4) compositions, which gave an overall accuracy of 81.3%. We also developed SVM module hybrid3 based on amino acid composition, traditional and higher order dipeptide compositions, and PSI-BLAST output and achieved an overall accuracy of 84.4%. A Web server HSLPred (www.imtech.res.in/raghava/hslpred/ or bioinformatics.uams.edu/raghava/hslpred/) has been designed to predict subcellular localization of human proteins using the above approaches.

Algorithms↗

Automated identification of putative methyltransferases from genomic open reading frames.

We have analyzed existing methodologies and created novel methodologies for the automatic assignment of S-adenosylmethionine (AdoMet)-dependent methyltransferase functionality to genomic open reading frames based on predicted protein sequences. A large class of the AdoMet-dependent methyltransferases shares a common binding motif for the AdoMet cofactor in the form of a seven-strand twisted beta-sheet; this structural similarity is mirrored in a degenerate sequence similarity that we refer to as methyltransferase signature motifs. These motifs are the basis of our assignments. We find that simple pattern matching based on the motif sequence is of limited utility and that a new method of "sensitized matrices for scoring methyltransferases" (SM2) produced with modified versions of the MEME and MAST tools gives greatly improved results for the Saccharomyces cerevisiae yeast genome. From our analysis, we conclude that this class of methyltransferases makes up approximately 0.6-1.6% of the genes in the yeast, human, mouse, Drosophila melanogaster, Caenorhabditis elegans, Arabidopsis thaliana, and Escherichia coli genomes. We provide lists of unidentified genes that we consider to have a high probability of being methyltransferases for future biochemical analyses.

Amino Acid Motifs↗