Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Two-dimensional electrophoresis database of fluorescence-labeled proteins of colon cancer cells.

We constructed a novel database of the proteome of DLD-1 colon cancer cells by two-dimensional polyacrylamide gel electrophoresis (2D-PAGE) of fluorescence-labeled proteins followed by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF) analysis. The database consists of 258 functionally categorized proteins corresponding to 314 protein spots. The majority of the proteins are oxidoreductases, cytoskeletal proteins and nucleic acid binding proteins. Phosphatase treatment showed that 28% of the protein spots on the gel are phosphorylated, and mass spectrometric analysis identified 21 of them. Proteins of DLD-1 cells and of laser-microdissected colon cancer tissues showed similar distribution on 2D gels, suggesting the utility of our database for clinical proteomics.

Amino Acid Sequence↗

Local database and the search program for proteomic analysis of sperm proteins in the ascidian Ciona intestinalis.

Separation of proteins by two-dimensional electrophoresis and following mass spectrometry (MS) is now a conventional technique for proteomic analysis. For proteomic analysis of a certain tissue with a limited information of primary structures of proteins, we have developed an analytical system for peptide mass fingerprinting in gene products in the testis of the ascidian Ciona intestinalis. Ciona sperm proteins were separated by two-dimensional gel electrophoresis and the tryptic fragments were subjected to MALDI-TOF/MS. The mass pattern was searched against on-line databases but resulted in less identification of these proteins. We have constructed a MS database from Ciona testis ESTs and the genome draft sequence, along with a newly devised, perl-based search program PerMS for peptide mass fingerprinting. This system could identify more than 80% of Ciona sperm proteins, suggesting that it could be widely applied for proteomic analysis for a limited tissue with less genomic information.

Animals↗

Toward computer-based cleavage site prediction of cysteine endopeptidases.

Identification of relevant substrates is essential for elucidation of in vivo functions of peptidases. The recent availability of the complete genome sequences of many eukaryotic organisms holds the promise of identifying specific peptidase substrates by systematic proteome analyses in combination with computer-based screening of genome databases. Currently available proteomics and bioinformatics tools are not sufficient for reliable endopeptidase substrate predictions. To address these shortcomings the bioinformatics tool 'PEPS' (Prediction of Endopeptidase Substrates) has been developed and is presented here. PEPS uses individual rule-based endopeptidase cleavage site scoring matrices (CSSM). The efficiency of PEPS in predicting putative caspase 3, cathepsin B and cathepsin L cleavage sites is demonstrated in comparison to established algorithms. Mortalin, a member of the heat shock protein family HSP70, was identified by PEPS as a putative cathepsin L substrate. Comparative proteome analyses of cathepsin L-deficient and wild-type mouse fibroblasts showed that mortalin is enriched in the absence of cathepsin L. These results indicate that CSSM/PEPS can correctly predict relevant peptidase substrates.

Animals↗

PEDRo: a database for storing, searching and disseminating experimental proteomics data.

BACKGROUND: Proteomics is rapidly evolving into a high-throughput technology, in which substantial and systematic studies are conducted on samples from a wide range of physiological, developmental, or pathological conditions. Reference maps from 2D gels are widely circulated. However, there is, as yet, no formally accepted standard representation to support the sharing of proteomics data, and little systematic dissemination of comprehensive proteomic data sets. RESULTS: This paper describes the design, implementation and use of a Proteome Experimental Data Repository (PEDRo), which makes comprehensive proteomics data sets available for browsing, searching and downloading. It is also serves to extend the debate on the level of detail at which proteomics data should be captured, the sorts of facilities that should be provided by proteome data management systems, and the techniques by which such facilities can be made available. CONCLUSIONS: The PEDRo database provides access to a collection of comprehensive descriptions of experimental data sets in proteomics. Not only are these data sets interesting in and of themselves, they also provide a useful early validation of the PEDRo data model, which has served as a starting point for the ongoing standardisation activity through the Proteome Standards Initiative of the Human Proteome Organisation.

Animals↗

Proteomic analysis of the mouse mammary gland is a powerful tool to identify novel proteins that are differentially expressed during mammary development.

After lactation, the mouse mammary gland undergoes apoptosis and tissue remodelling as the gland reverts to its prepregnant state. This complex change was investigated using 2-DE. An integrated database was produced from lactation and involution proteomes. Forty-four molecular cluster indexes (MCIs) that showed altered expression from lactation to involution were selected for MS analysis. Of these, 32 gave protein annotations, 18 of which were unequivocal proteins. Selected proteins were then studied across all of development, including pregnancy, using data integrated from another proteome database. Two proteins, the RNA polymerase B transcription factor 3 (BTF3) and the minichromosome maintenance protein 3 (MCM3), although initially selected on the basis of the lactation/involution criteria, had expression profiles that indicated an additional role in mammary development and were further analysed. BTF3, a transcription factor previously not described in the mammary gland, was up-regulated strongly in pregnancy, indicating an involvement in alveolar growth. MCM3's expression was greatest in pregnancy and late involution, decreasing through lactation. Immunohistochemistry localised MCM3 to the mammary epithelium, where a greater proportion of cells stained than for the proliferation marker Ki67. MCM3 expression during lactation may identify cells that are licensed to repopulate the gland during cell loss in lactation and following involution.

Animals↗

Glycome project: concept, strategy and preliminary application to Caenorhabditis elegans.

Glycans play a central role as potential mediators between complex cell societies, because all living organisms consist of cells covered with diverse carbohydrate chains reflecting various cell types and states. However, we have no idea how diverse these carbohydrate chains actually are. The main purpose of this article is to persuade life scientists to realize the fundamental importance of taking some action by becoming involved in "glycomics". "Glycome" is a term meaning the whole set of glycans produced by individual organisms, as the third bioinformative macromolecules to be elucidated next to the genome and proteome. Here a basic strategy is presented. The essence of the project includes the following: (a) glycopeptides, but not glycans released from their core proteins, are targeted for linkage to genome databases; (b) Caenorhabditis elegans is used as the first model organism for this project, since its genome project has already been completed; (c) four essential attributes are adopted to characterize each glycopeptide: (i) cosmid identification number (ID), (ii) molecular weight (M(r)), (iii) retention (Rs) of pyridylaminated (PA) oligosaccharides in 2-D mapping, and (iv) dissociation constants (Kd's) of PA-oligosaccharides for a set of lectins. Thus, the obtained ID, M(r), R and Kd's construct the glycome database, which will be open as the previous genome and proteome databases. For the project to proceed the "glyco-catch" method is proposed, where a group of target glycopeptides are captured by means of lectin-affinity chromatography after protease digestion. Already glycopeptides from asialofetuin and ovalbumin were successfully captured by galectin-agarose and Con A-agarose, respectively. Further, to examine the practical validity of the method, we extracted membrane proteins from C. elegans with 1% Triton X-100, and isolated specific glycopeptides by use of the same galectin column. One of the glycopeptides was successfully identified in the C. elegans genome database. Finally, for determination of Kd between glycopeptides and lectins, a recently reinforced frontal affinity chromatography (FAC) is proposed as an alternative to define glycan structures in place of determining every covalent structure.

Animals↗

Protein identification by tandem mass spectrometry and sequence database searching.

The shotgun proteomics strategy, based on digesting proteins into peptides and sequencing them using tandem mass spectrometry (MS/MS), has become widely adopted. The identification of peptides from acquired MS/MS spectra is most often performed using the database search approach. We provide a detailed description of the peptide identification process and review the most commonly used database search programs. The appropriate choice of the search parameters and the sequence database are important for successful application of this method, and we provide general guidelines for carrying out efficient analysis of MS/MS data. We also discuss various reasons why database search tools fail to assign the correct sequence to many MS/MS spectra, and draw attention to the problem of false-positive identifications that can significantly diminish the value of published data. To assist in the evaluation of peptide assignments to MS/MS spectra, we review the scoring schemes implemented in most frequently used database search tools. We also describe statistical approaches and computational tools for validating peptide assignments to MS/MS spectra, including the concept of expectation values, reversed database searching, and the empirical Bayesian analysis of PeptideProphet. Finally, the process of inferring the identities of the sample proteins given the list of peptide identifications is outlined, and the limitations of shotgun proteomics with regard to discrimination between protein isoforms are discussed.

Amino Acid Sequence↗

3D-GENOMICS: a database to compare structural and functional annotations of proteins between sequenced genomes.

The 3D-GENOMICS database (http://www.sbg.bio. ic.ac.uk/3dgenomics/) provides structural annotations for proteins from sequenced genomes. In August 2003 the database included data for 93 proteomes. The annotations stored in the database include homologous sequences from various sequence databases, domains from SCOP and Pfam, patterns from Prosite and other predicted sequence features such as transmembrane regions and coiled coils. In addition to annotations at the sequence level, several precomputed cross- proteome comparative analyses are available based on SCOP domain superfamily composition. Annotations are available to the user via a web interface to the database. Multiple points of entry are available so that a user is able to: (i) directly access annotations for a single protein sequence via keywords or accession codes, (ii) examine a sequence of interest chosen from a summary of annotations for a particular proteome, or (iii) access precomputed frequency-based cross-proteome comparative analyses.

Amino Acid Sequence↗

Progress in in silico functional genomics: the malaria Metabolic Pathways database.

Genomic, transcriptomic and proteomic data can be turned into biologically meaningful information if they are synthesized into processes. Such amalgamation has been done for the most virulent malaria parasite Plasmodium falciparum in the Metabolic Pathways database. The dialectics of construction of metabolic pathways and other biological processes using this database is presented here. Additional features, such as links to other databases and the incorporation of transcriptomic clocks, are elucidated. Comparison of Metabolic Pathways to other similar databases is analyzed.

Animals↗

[Advances in plant proteomics--I. Key techniques of proteome].

With the completion of genome sequences of the model plants, such as rice (Oryza sativa L.) and Arabidopsis thaliana, it has come into plant functional genomics era, which becomes a hard base of the appearance and development of plant proteomics. This review focuses on the background of proteomics, concept of proteomics and key techniques of proteomics. The key techniques of proteomics include separation, such as 2-DE (Two-Dimensional Electrophoresis), RP-HPLC (Reverse Phase High Performance Liquid Chromatography) and SELDI (Surface Enhanced Laser Desorption/ Ionization) protein chip, mass spectrometry, such as MALDI-TOF-MS (Matrix Assisted Laser Desorption/Ionization-Time Of Flight-Mass Spectrometry) and ESI-MS/MS (Electrospray Ionization Mass Spectrometry/Mass Spectrometry), databases related to proteomics, quantitative proteome, TAP (Tandem Affinity Purification) and yeast two-hybrid system. Challenges and prospects of proteomic techniques are discussed.

Computational Biology↗

Overview of the HUPO Plasma Proteome Project: results from the pilot phase with 35 collaborating laboratories and multiple analytical groups, generating a core dataset of 3020 proteins and a publicly-available database.

HUPO initiated the Plasma Proteome Project (PPP) in 2002. Its pilot phase has (1) evaluated advantages and limitations of many depletion, fractionation, and MS technology platforms; (2) compared PPP reference specimens of human serum and EDTA, heparin, and citrate-anti-coagulated plasma; and (3) created a publicly-available knowledge base (www.bioinformatics.med.umich.edu/hupo/ppp; www.ebi.ac.uk/pride). Thirty-five participating laboratories in 13 countries submitted datasets. Working groups addressed (a) specimen stability and protein concentrations; (b) protein identifications from 18 MS/MS datasets; (c) independent analyses from raw MS-MS spectra; (d) search engine performance, subproteome analyses, and biological insights; (e) antibody arrays; and (f) direct MS/SELDI analyses. MS-MS datasets had 15 710 different International Protein Index (IPI) protein IDs; our integration algorithm applied to multiple matches of peptide sequences yielded 9504 IPI proteins identified with one or more peptides and 3020 proteins identified with two or more peptides (the Core Dataset). These proteins have been characterized with Gene Ontology, InterPro, Novartis Atlas, OMIM, and immunoassay-based concentration determinations. The database permits examination of many other subsets, such as 1274 proteins identified with three or more peptides. Reverse protein to DNA matching identified proteins for 118 previously unidentified ORFs. We recommend use of plasma instead of serum, with EDTA (or citrate) for anticoagulation. To improve resolution, sensitivity and reproducibility of peptide identifications and protein matches, we recommend combinations of depletion, fractionation, and MS/MS technologies, with explicit criteria for evaluation of spectra, use of search algorithms, and integration of homologous protein matches. This Special Issue of PROTEOMICS presents papers integral to the collaborative analysis plus many reports of supplementary work on various aspects of the PPP workplan. These PPP results on complexity, dynamic range, incomplete sampling, false-positive matches, and integration of diverse datasets for plasma and serum proteins lay a foundation for development and validation of circulating protein biomarkers in health and disease.

Algorithms↗

Data mining crystallization databases: knowledge-based approaches to optimize protein crystal screens.

Protein crystallization is a major bottleneck in protein X-ray crystallography, the workhorse of most structural proteomics projects. Because the principles that govern protein crystallization are too poorly understood to allow them to be used in a strongly predictive sense, the most common crystallization strategy entails screening a wide variety of solution conditions to identify the small subset that will support crystal nucleation and growth. We tested the hypothesis that more efficient crystallization strategies could be formulated by extracting useful patterns and correlations from the large data sets of crystallization trials created in structural proteomics projects. A database of crystallization conditions was constructed for 755 different proteins purified and crystallized under uniform conditions. Forty-five percent of the proteins formed crystals. Data mining identified the conditions that crystallize the most proteins, revealed that many conditions are highly correlated in their behavior, and showed that the crystallization success rate is markedly dependent on the organism from which proteins derive. Of the proteins that crystallized in a 48-condition experiment, 60% could be crystallized in as few as 6 conditions and 94% in 24 conditions. Consideration of the full range of information coming from crystal screening trials allows one to design screens that are maximally productive while consuming minimal resources, and also suggests further useful conditions for extending existing screens.

Archaeal Proteins↗

Proteome analysis in the study of lymphoma cells.

This review provides an overview on recent studies in the field of proteome analysis of lymphoma cells, and highlights the potentials of such studies for a better knowledge of drug effects at the molecular level. After giving general information on the field of proteome analysis of lymphoma cells, some characteristics of the strategies used during this analysis are pointed out, such as cell extraction strategies and affinity captures. Therefore, the issue of proteome analysis of lymphoma cells content will be covered with respect to those protein extracts that can be prepared in saline solutions, such as cytoplasm proteins, or that are associated with the cell membranes. The question of which kinds of information have been retrieved from lymphoma-cell proteomics is discussed on the basis of several examples-lymphoma cell-mapping studies and constitution of protein databases, and comparative proteome analysis studies of the modifications that result from a drug treatment.

Animals↗

The human proteome organization (HUPO) and environmental health.

The Human Proteome Organization, or HUPO, was formed to promote research and large-scale analysis of the human proteome. By consolidating national proteome organizations into an international body, HUPO will coordinate international initiatives, biological resources, protocols, standards and data for studying the human proteome. HUPO has identified five key areas to advance study of the human proteome, specifically in bioinformatics, new technologies, the plasma proteome, cell models, and a public antibody initiative. Consideration of three major issue areas may help develop HUPO's strategy for human proteome study. First is the need to distinguish the value of high throughput platforms from discovery platforms in proteomics. Second is the importance for international planning on integrating both transcriptome and proteome data and databases. Last is that effects of the environment from chemical, physical, and biological exposures alter the expression and structure of the proteome, which become manifest in long-term adverse health effects and disease. Environmental health research stands to greatly benefit from the shared resources, data, and vision of the HUPO organization as a valuable resource in exploiting knowledge of the human proteome toward improving public health.

Computational Biology↗

The Integr8 project--a resource for genomic and proteomic data.

Integr8 (http://www.ebi.ac.uk/integr8/) is providing an integration layer for the exploitation of genomic and proteomic data by drawing on databases maintained at major bioinformatics centres in Europe. Main aims are to store the relationships of biological entities to each other and to entries in other databases, to provide a framework that allows for new kinds of data to be integrated, and to offer an entity-centric view of complete genomes and proteomes. Basic tools for data integration comprise the Proteome Analysis database, the International Protein Index (IPI), the Universal Protein sequence archive (UniParc) and the Genome Reviews. Entry points for the Integr8 portal depend on the users entity of interest: from browsing the taxonomy or with a predetermined species of interest, the species page can be used, and a simple search page leads to different applications when looking for certain protein sequences or genes. Customisable statistics data are available from the BioMart application, and pre-prepared data can be downloaded from the FTP site.

Computational Biology↗

Mapping the proteome of Drosophila melanogaster: analysis of embryos and adult heads by LC-IMS-MS methods.

Multidimensional separations combined with mass spectrometry are used to study the proteins that are present in two states of Drosophila melanogaster: the whole embryo and the adult head. The approach includes the incorporation of a gas-phase separation dimension in which ions are dispersed according to differences in their mobilities and is described as a means of providing a detailed analytical map of the proteins that are present. Overall, we find evidence for 1133 unique proteins. In total, 780 are identified in the head, and 660 are identified in the embryo. Only 307 proteins are in common to both developmental stages, indicating that there are significant differences in these proteomes. A comparison of the proteome to a database of mRNAs that are found from analysis by cDNA approaches (i.e., transcriptome) also shows little overlap. All of this information is discussed in terms of the relationship between the predicted genome, and measured transcriptomes and proteomes. Additionally, the merits and weaknesses of current technologies are assessed in some detail.

Animals↗

Construction of a Francisella tularensis two-dimensional electrophoresis protein database.

We have started the construction of a two-dimensional database of the proteome of Francisella tularensis, a bacterium that is responsible for the highly pathogenic disease tularemia. The genome of this intracellular pathogen is not completely sequenced yet and, currently, information about only 66 proteins is available from NCBI database. We have analyzed the F. tularensis live vaccine strain by two-dimensional gel electrophoresis with immobilized pH 3-10 gradient in the first dimension and 9-16% gradient or tricine SDS-PAGE in the second dimension. In both cases about 2000 spots were detected. Furthermore, we compared the protein pattern of the nonvirulent F. tularensis live vaccine strain with protein profiles of two wild type clinical isolates and more than 50 differentially expressed proteins were counted. The separated proteins are going to be identified by peptide mass fingerprinting. However, due to the lack of complete genome sequence data only eight proteins were unambiguously identified. Among them, acid phosphatase and the most basic isoform of a hypothetical 23 kDa protein are characteristic only for virulent strains.

Bacterial Proteins↗