Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Perturbation and interpretation of nitrogen isotope distribution patterns in proteomics.

This study provides a discussion on the applications and limitations of (15)NH(4)(+) metabolic labeling in proteomic studies. The hyperthemophilic crenarchaeon Sulfolobus solfataricus was used as a model organism throughout this study. The distribution of nitrogen was studied in four different experiments in which this distribution was manipulated in a unique way. The experiments included full adaptation to media with relative isotope abundances (RIA) of 0.36%, 50%, and >98% (15)NH(4)(+). The incorporation efficiency was calculated on the basis of a comparison between theoretical and experimental spectra. In the case of full adaptation, incorporation efficiencies reflected the RIA (0.36%, 47.5% and 99% respectively). Labeling efficiencies were calculated on the basis of peak areas in TOF-MS spectra. It is shown that in the case of full adaptation, labeling efficiencies are 100%. In addition, we demonstrate that (15)NH(4)(+) labeling can be used in protein turnover studies, even when labeling is incomplete. In this case, incorporation efficiencies of 88-93% (lower than the RIA) were measured, providing evidence for amino acid recycling. Labeling efficiencies were always between 63% and 94% providing evidence for protein degradation. Finally, it was shown that isotope distributions can be useful in peptide identification.

Ammonia↗

Bioinformatics in the post-genome era.

Recent years saw a dramatic increase in genomic and proteomic data in public archives. Now with the complete genome sequences of human and other species in hand, detailed analyses of the genome sequences will undoubtedly improve our understanding of biological systems and at the same time require sophisticated bioinformatic tools. Here we review what computational challenges are ahead and what are the new exciting developments in this exciting field.

Animals↗

Proteomics of the chloroplast: systematic identification and targeting analysis of lumenal and peripheral thylakoid proteins.

The soluble and peripheral proteins in the thylakoids of pea were systematically analyzed by using two-dimensional electrophoresis, mass spectrometry, and N-terminal Edman sequencing, followed by database searching. After correcting to eliminate possible isoforms and post-translational modifications, we estimated that there are at least 200 to 230 different lumenal and peripheral proteins. Sixty-one proteins were identified; for 33 of these proteins, a clear function or functional domain could be identified, whereas for 10 proteins, no function could be assigned. For 18 proteins, no expressed sequence tag or full-length gene could be identified in the databases, despite experimental determination of a significant amount of amino acid sequence. Nine previously unidentified proteins with lumenal transit peptides are presented along with their full-length genes; seven of these proteins possess the twin arginine motif that is characteristic for substrates of the TAT pathway. Logoplots were used to provide a detailed analysis of the lumenal targeting signals, and all nuclear-encoded proteins identified on the two-dimensional gels were used to test predictions for chloroplast localization and transit peptides made by the software programs ChloroP, PSORT, and SignalP. A combination of these three programs was found to provide a useful tool for evaluating chloroplast localization and transit peptides and also could reveal possible alternative processing sites and dual targeting. The potential of proteomics for plant biology and homology-based searching with mass spectrometry data is discussed.

Amino Acid Sequence↗

In silico prediction of the peroxisomal proteome in fungi, plants and animals.

In an attempt to improve our abilities to predict peroxisomal proteins, we have combined machine-learning techniques for analyzing peroxisomal targeting signals (PTS1) with domain-based cross-species comparisons between eight eukaryotic genomes. Our results indicate that this combined approach has a significantly higher specificity than earlier attempts to predict peroxisomal localization, without a loss in sensitivity. This allowed us to predict 430 peroxisomal proteins that almost completely lack a localization annotation. These proteins can be grouped into 29 families covering most of the known steps in all known peroxisomal pathways. In general, plants have the highest number of predicted peroxisomal proteins, and fungi the smallest number.

Amino Acid Sequence↗

Escherichia coli--a model system that benefits from and contributes to the evolution of proteomics.

The large body of knowledge about Escherichia coli makes it a useful model organism for the expression of heterologous proteins. Proteomic studies have helped to elucidate the complex cellular responses of E. coli and facilitated its use in a variety of biotechnology applications. Knowledge of basic cellular processes provides the means for better control of heterologous protein expression. Beyond such important applications, E. coli is an ideal organism for testing new analytical technologies because of the extensive knowledge base available about the organism. For example, improved technology for characterization of unknown proteins using mass spectrometry has made two-dimensional electrophoresis (2DE) studies more useful and more rewarding, and much of the initial testing of novel protocols is based on well-studied samples derived from E. coli. These techniques have facilitated the construction of more accurate 2DE maps. In this review, we present work that led to the 2DE databases, including a new map based on tandem time-of-flight (TOF) mass spectrometry (MS); describe cellular responses relevant to biotechnology applications; and discuss some emerging proteomic techniques.

Electrophoresis, Gel, Two-Dimensional↗

Ovarian cancer detection by logical analysis of proteomic data.

A new type of efficient and accurate proteomic ovarian cancer diagnosis systems is proposed. The system is developed using the combinatorics and optimization-based methodology of logical analysis of data (LAD) to the Ovarian Dataset 8-7-02 (http://clinicalproteomics.steem.com), which updates the one used by Petricoin et al. in The Lancet 2002, 359, 572-577. This mass spectroscopy-generated dataset contains expression profiles of 15 154 peptides defined by their mass/charge ratios (m/z) in serum of 162 ovarian cancer and 91 control cases. Several fully reproducible models using only 7-9 of the 15 154 peptides were constructed, and shown in multiple cross-validation tests (k-folding and leave-one-out) to provide sensitivities and specificities of up to 100%. A special diagnostic system for stage I ovarian cancer patients is shown to have similarly high accuracy. Other results: (i) expressions of peptides with relatively low m/z values in the dataset are shown to be better at distinguishing ovarian cancer cases from controls than those with higher m/z values; (ii) two large groups of patients with a high degree of similarities among their formal (mathematical) profiles are detected; (iii) several peptides with a blocking or promoting effect on ovarian cancer are identified.

Algorithms↗

Website review: interPro (the integrated resource of protein domains and functional sites).

The family and motif databases, PROSITE, PRINTS, Pfam and ProDom, have been integrated into a powerful resource for protein secondary annotation. As of June 2000, InterPro had processed 384 572 proteins in SWISS-PROT and TrEMBL. Because the contributing databases have different clustering principles and scoring sensitivities, the combined assignments compliment each other for grouping protein families and delineating domains. The graphic displays of all matches above the scoring thresholds enables judgements to be made on the concordances or differences between the assignments. The website links can be used to analyse novel sequences and for queries across the proteomes of 32 organisms, including the partial human set, by domain and/or protein family. An analysis of selected HtrA/DegQ proteases demonstrates the utility of this website for detailed comparative genomics. Further information on the project can be found at the European Bioinformatics Institute at http://www.ebi.ac.uk/interpro/

Amino Acid Motifs↗

PlasmoDB: the Plasmodium genome resource. A database integrating experimental and computational data.

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates the recently completed P. falciparum genome sequence and annotation, as well as draft sequence and annotation emerging from other Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for intra- and inter-species comparisons. Sequence information is integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects and proteomics studies. The relational schema used to build PlasmoDB, GUS (Genomics Unified Schema) employs a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically-based, queries of the database. A stand-alone version of the database is also available on CD-ROM (P. falciparum GenePlot), facilitating access to the data in situations where internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to facilitate utilization of the vast quantities of genomic-scale data produced by the global malaria research community. The software used to develop PlasmoDB has been used to create a second Apicomplexan parasite genome database, ToxoDB (http://ToxoDB.org).

Animals↗

Technology development at the interface of proteome research and genomics: mapping nonpolymorphic proteins on the physical map of mouse chromosomes.

Data obtained from protein spots by peptide mass fingerprinting are used to identify the corresponding genes in sequence databases. The relevant cDNAs are obtained as clones from the Integrated Molecular Analysis of Genome Expression (I.M.A.G.E.) consortium. Mapping of I.M.A.G.E. clones is performed in two steps: first, cDNA clones are hybridized against a 10-hit genomic mouse bacterial artificial chromosome (BAC) library. Second, interspersed repetitive sequence polymerase chain reaction (IRS-PCR) using a single primer directed against the mouse B1 repeat element is performed on BACs. As each cDNA detects several BACs, and each individual BAC has a 50% chance to recover an IRS-PCR fragment, the majority of cDNAs produce at least a single IRS-PCR fragment. Individual IRS fragments are hybridized against high-density spotted filter grids containing the three-dimensional permutated pools of yeast artificial chromosome (YAC) library resources that are currently being used to construct a physical map of the mouse genome. IRS fragments that hybridize to YAC clones already placed into contigs immediately provide highly precise map positions. This technology therefore is able to draw links between proteins detected by 2-D gel electrophoresis and the corresponding gene loci in the mouse genome.

Animals↗

Global and targeted quantitative proteomics for biomarker discovery.

The extraordinary developments made in proteomic technologies in the past decade have enabled investigators to consider designing studies to search for diagnostic and therapeutic biomarkers by scanning complex proteome samples using unbiased methods. The major technology driving these studies is mass spectrometry (MS). The basic premises of most biomarker discovery studies is to use the high data-gathering capabilities of MS to compare biological samples obtained from healthy and disease-afflicted patients and identify proteins that are differentially abundant between the two specimen. To meet the need to compare the abundance of proteins in different samples, a number of quantitative approaches have been developed. In this article, many of these will be described with an emphasis on their advantageous and disadvantageous for the discovery of clinically useful biomarkers.

Biomarkers↗

Protein modeling and structure prediction with a reduced representation.

Protein modeling could be done on various levels of structural details, from simplified lattice or continuous representations, through high resolution reduced models, employing the united atom representation, to all-atom models of the molecular mechanics. Here I describe a new high resolution reduced model, its force field and applications in the structural proteomics. The model uses a lattice representation with 800 possible orientations of the virtual alpha carbon-alpha carbon bonds. The sampling scheme of the conformational space employs the Replica Exchange Monte Carlo method. Knowledge-based potentials of the force field include: generic protein-like conformational biases, statistical potentials for the short-range conformational propensities, a model of the main chain hydrogen bonds and context-dependent statistical potentials describing the side group interactions. The model is more accurate than the previously designed lattice models and in many applications it is complementary and competitive in respect to the all-atom techniques. The test applications include: the ab initio structure prediction, multitemplate comparative modeling and structure prediction based on sparse experimental data. Especially, the new approach to comparative modeling could be a valuable tool of the structural proteomics. It is shown that the new approach goes beyond the range of applicability of the traditional methods of the protein comparative modeling.

Amino Acid Sequence↗

Zinc finger proteins and other transcription regulators as response proteins in benzo[a]pyrene exposed cells.

Proteomic analysis, which combines two-dimensional electrophoresis (2-DE) and mass spectrometry (MS), is an important approach to screen proteins responsive to specific stimuli. Benzo[a]pyrene (B[a]P), a prototype of polycyclic hydrocarbons (PAHs), is a potent procarcinogen generated from the combustion of fossil fuel and cigarette smoke. To further probe the molecular mechanism of mutagenesis and carcinogenesis, and to find potential molecular markers involved in cellular responses to B[a]P exposure, we performed proteomic analysis of whole cellular proteins in human amnion epithelial cells after B[a]P-treatment. Image visualization and statistical analysis indicated that more than 40 proteins showed significant changes following B[a]P-treatment (P < 0.05). Among them, 20 proteins existed only in the control groups, while six were only present in B[a]P-treated cells. In addition, the expression of 10 proteins increased whereas 11 decreased after B[a]P-treatment. These proteins were subjected to in-gel tryptic digestion followed by matrix-assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF-MS) analysis. Using peptide mass fingerprinting (PMF) to search the nrNCBI database, we identified 22 proteins. Most of these proteins have unknown functions and have not been previously connected to a response to B[a]P exposure. To further annotate the characteristics of these proteins, GOblet analysis was carried out and results indicated that they were involved in multiple biological processes including regulation of transcription, cell proliferation, cell aging and other processes. However, expression changes were noted in a number of transcription regulators, including eight zinc finger proteins as well as SNF2L1 (SWI/SNF related, matrix associated, actin dependent regulator of chromatin, subfamily a, member 1), which is closely linked to the chromatin remodeling process. These data may provide new clues to further understand the implication of these proteins in cellular responses to carcinogen exposure as well as the molecular mechanisms of B[a]P-induced mutagenesis and carcinogenesis.

Amino Acid Sequence↗

Predicting transmembrane beta-barrels in proteomes.

Very few methods address the problem of predicting beta-barrel membrane proteins directly from sequence. One reason is that only very few high-resolution structures for transmembrane beta-barrel (TMB) proteins have been determined thus far. Here we introduced the design, statistics and results of a novel profile-based hidden Markov model for the prediction and discrimination of TMBs. The method carefully attempts to avoid over-fitting the sparse experimental data. While our model training and scoring procedures were very similar to a recently published work, the architecture and structure-based labelling were significantly different. In particular, we introduced a new definition of beta- hairpin motifs, explicit state modelling of transmembrane strands, and a log-odds whole-protein discrimination score. The resulting method reached an overall four-state (up-, down-strand, periplasmic-, outer-loop) accuracy as high as 86%. Furthermore, accurately discriminated TMB from non-TMB proteins (45% coverage at 100% accuracy). This high precision enabled the application to 72 entirely sequenced Gram-negative bacteria. We found over 164 previously uncharacterized TMB proteins at high confidence. Database searches did not implicate any of these proteins with membranes. We challenge that the vast majority of our 164 predictions will eventually be verified experimentally. All proteome predictions and the PROFtmb prediction method are available at http://www.rostlab.org/ services/PROFtmb/.

Markov Chains↗

Proteomic identification of oxidatively modified retinal proteins in a chronic pressure-induced rat model of glaucoma.

PURPOSE: Based on the evidence of an amplified production of reactive oxygen species (ROS) during glaucomatous neurodegeneration, proteomic analysis was performed to determine oxidative modification of retinal proteins after experimental elevation of intraocular pressure (IOP). METHODS: IOP elevation was induced in rats by hypertonic saline injections into episcleral veins. Protein expression was determined by two-dimensional polyacrylamide gel electrophoresis (2D-PAGE) of retinal protein lysates obtained from eyes matched for the cumulative IOP exposure and axon loss. To determine protein oxidation levels, protein carbonyls were detected through 2D-oxyblot analysis of 2,4-dinitrophenylhydrazine (DNPH)-treated samples using an anti-DNP antibody. For identification of oxidized proteins, peptide masses were analyzed by matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF/MS) and liquid chromatography-tandem mass spectrometry (LC/MS/MS). In addition to use of different engines in a bioinformatic database search and performance of peptide sequencing and 2D-Western blot analysis for confirmation of the identified proteins, immunohistochemistry was used for further validation of the proteomic findings. RESULTS: Comparison of 2D-oxyblots with Coomassie Blue-stained 2D-gels revealed that approximately 60 protein spots obtained with retinal protein lysates from ocular hypertensive eyes (of >400 spots) exhibited protein carbonyl immunoreactivity, which reflects oxidatively modified proteins. There was a significant increase in anti-carbonyl reactivity in individual protein spots obtained with retinal protein lysates from ocular hypertensive eyes compared with the control (P < 0.01). The identified proteins through peptide mass fingerprinting and peptide sequencing included glyceraldehyde-3-phosphate dehydrogenase, a glycolytic enzyme; HSP72, a stress protein; and glutamine synthetase, an excitotoxicity-related protein. Immunolabeling of retina sections with specific antibodies demonstrated cellular localization of these proteins as well as retinal distribution of the increased protein carbonyl immunoreactivity in ocular hypertensive eyes. CONCLUSIONS: The findings of this in vivo study provide novel evidence for oxidative modification of many retinal proteins in ocular hypertensive eyes and identify three specific targets of retinal protein oxidation in these eyes, thereby supporting the association of oxidative damage with neurodegeneration in glaucoma. By using a proteomic approach, this study also exemplifies that proteomics provide a very promising way to elucidate pathogenic mechanisms in glaucoma at the protein level.

Animals↗

Application of proteomics and protein analysis for biomarker and target finding for immunotherapy.

Regulatory T-cells play a central role in the maintenance of the immunological balance and are powerful inhibitors of T-cell activation both in vivo and in vitro. The enhancement of suppressor-cell function might be a target for immunotherapeutic approaches for the treatment of immune-mediated diseases like multiple sclerosis and Crohn's disease.The method of choice to elucidate the still unclear effector functions of regulatory T-cells is the differential proteome analyses performed with human and murine T-cell populations. To this end, whole-protein extracts of conventional and regulatory T-cells are separated by high-resolution two-dimensional gel electrophoresis according to Klose. The proteomes are analyzed by a 2DE gel image analysis software, ProteomWeaver. The protein spots that are found differentially expressed are picked from the gels and prepared for matrix-assisted laser desorption/ionization (MALDI) mass spectrometrical analysis automatically. The high-resolution 2DE-PAGE and the automated spot handling and protein identification allows one to rapidly find new potential candidate proteins that are of functional relevance for regulatory T-cells, to be used as targets for drug development or as biomarkers for research and diagnostic purposes.

Amino Acid Sequence↗

Optimal docking area: a new method for predicting protein-protein interaction sites.

Understanding energetics and mechanism of protein-protein association remains one of the biggest theoretical problems in structural biology. It is assumed that desolvation must play an essential role during the association process, and indeed protein-protein interfaces in obligate complexes have been found to be highly hydrophobic. However, the identification of protein interaction sites from surface analysis of proteins involved in non-obligate protein-protein complexes is more challenging. Here we present Optimal Docking Area (ODA), a new fast and accurate method of analyzing a protein surface in search of areas with favorable energy change when buried upon protein-protein association. The method identifies continuous surface patches with optimal docking desolvation energy based on atomic solvation parameters adjusted for protein-protein docking. The procedure has been validated on the unbound structures of a total of 66 non-homologous proteins involved in non-obligate protein-protein hetero-complexes of known structure. Optimal docking areas with significant low-docking surface energy were found in around half of the proteins. The 'ODA hot spots' detected in X-ray unbound structures were correctly located in the known protein-protein binding sites in 80% of the cases. The role of these low-surface-energy areas during complex formation is discussed. Burial of these regions during protein-protein association may favor the complexed configurations with near-native interfaces but otherwise arbitrary orientations, thus driving the formation of an encounter complex. The patch prediction procedure is freely accessible at http://www.molsoft.com/oda and can be easily scaled up for predictions in structural proteomics.

Binding Sites↗

Prediction of protease types in a hybridization space.

Regulating most physiological processes by controlling the activation, synthesis, and turnover of proteins, proteases play pivotal regulatory roles in conception, birth, digestion, growth, maturation, ageing, and death of all organisms. Different types of proteases have different functions and biological processes. Therefore, it is important for both basic research and drug discovery to consider the following two problems. (1) Given the sequence of a protein, can we identify whether it is a protease or non-protease? (2) If it is, what protease type does it belong to? Although the two problems can be solved by various experimental means, it is both time-consuming and costly to do so. The avalanche of protein sequences generated in the post-genetic era has challenged us to develop an automated method for making a fast and reliable identification. By hybridizing the functional domain composition and pseudo-amino acid composition, we have introduced a new method called "FunD-PseAA predictor" that is operated in a hybridization space. To avoid redundancy and bias, demonstrations were performed on a dataset where none of the proteins has >or=25% sequence identity to any other. The overall success rate thus obtained by the jackknife cross-validation test in identifying protease and non-protease was 92.95%, and that in identifying the protease type was 94.75% among the following six types: (1) aspartic, (2) cysteine, (3) glutamic, (4) metallo, (5) serine, and (6) threonine. Demonstration was also made on an independent dataset, and the corresponding overall success rates were 98.36% and 97.11%, respectively, suggesting the FunD-PseAA predictor is very powerful and may become a useful tool in bioinformatics and proteomics.

Algorithms↗

A glycoproteome database of normal human liver tissue.

PURPOSE: To extensively investigate the glycoproteins of normal human liver tissue, constructing the glycoprotein profile and database of the normal human liver tissue. METHODS: The total proteins were extracted from the normal human liver tissue and then subjected to two-dimensional electrophoresis (2-DE). Finally, 2-DE gels were stained according to the methods of multiplexed proteomics (MP) technology. Glycoprotein spots were excised from 2-DE gel and then characterized by matrix assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF-MS). RESULTS: The PDQuest software detected 1,011 glycoprotein spots and 1,923 total protein spots in the 2-DE gels of sample from the normal human liver tissue. Furthermore, 116 species of glycoproteins were successfully identified via peptide mass profiling using MALDI-TOF-MS/MS and annotated to our databases. In addition, we also applied bioinformatics softwares to predict N- or O-glycosylation sites of identified glycoproteins. CONCLUSION: This study demonstrates the feasibility of a novel technological platform to contruct glycoprotein databases. These results lay the foundation for future physiological and pathological studies of the human liver.

Databases, Protein↗