Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Report on the first congress of the Spanish Proteomics Society.

This report describes the first congress of the Spanish Proteomics Society (SEProt) and the Foundation Meeting of the European Proteomics Association (EuPA), which were held in Cordoba, Spain, February 13-17, 2005. The EuPA meeting joined together thirty European representatives from 17 national European proteomics organisations to discuss the opportunity of creating a supranational association aimed to promote all aspects of proteomics as a scientific discipline throughout Europe. The main subjects of the SEProt congress, which was attended by over 350 researchers mainly from Spain and other European countries, were the current status and applications of proteomics in medical and biological sciences.

Congresses as Topic↗

Scanning the available Dictyostelium discoideum proteome for O-linked GlcNAc glycosylation sites using neural networks.

Dictyostelium discoideum has been suggested as a eukaryotic model organism for glycobiology studies. Presently, the characteristics of acceptor sites for the N-acetylglucosaminyl-transferases in Dictyostelium discoideum, which link GlcNAc in an alpha linkage to hydroxyl residues, are largely unknown. This motivates the development of a species specific method for prediction of O-linked GlcNAc glycosylation sites in secreted and membrane proteins of D. discoideum. The method presented here employs a jury of artificial neural networks. These networks were trained to recognize the sequence context and protein surface accessibility in 39 experimentally determined O-alpha-GlcNAc sites found in D. discoideum glycoproteins expressed in vivo. Cross-validation of the data revealed a correlation in which 97% of the glycosylated and nonglycosylated sites were correctly identified. Based on the currently limited data set, an abundant periodicity of two (positions-3, -1, +1, +3, etc.) in Proline residues alternating with hydroxyl amino acids was observed upstream and downstream of the acceptor site. This was a consequence of the spacing of the glycosylated residues themselves which were peculiarly found to be situated only at even positions with respect to each other, indicating that these may be located within beta-strands. The method has been used for a rapid and ranked scan of the fraction of the Dictyostelium proteome available in public databases, remarkably 25-30% of which were predicted glycosylated. The scan revealed acceptor sites in several proteins known experimentally to be O-glycosylated at unmapped sites. The available proteome was classified into functional and cellular compartments to study any preferential patterns of glycosylation. A sequence based prediction server for GlcNAc O-glycosylations in D. discoideum proteins has been made available through the WWW at http://www.cbs.dtu.dk/services/DictyOGlyc/ and via E-mail to DictyOGlyc@cbs.dtu.dk.

Algorithms↗

Population proteomics: the concept, attributes, and potential for cancer biomarker research.

This review outlines the concept of population proteomics and its implication in the discovery and validation of cancer-specific protein modulations. Population proteomics is an applied subdiscipline of proteomics engaging in the investigation of human proteins across and within populations to define and better understand protein diversity. Population proteomics focuses on interrogation of specific proteins from large number of individuals, utilizing top-down, targeted affinity mass spectrometry approaches to probe protein modifications. Deglycosylation, sequence truncations, side-chain residue modifications, and other modifications have been reported for myriad of proteins, yet little is know about their incidence rate in the general population. Such information can be gathered via population proteomics and would greatly aid the biomarker discovery efforts. Discovery of novel protein modifications is also expected from such large scale population proteomics, expanding the protein knowledge database. In regard to cancer protein biomarkers, their validation via population proteomics-based approaches is advantageous as mass spectrometry detection is used both in the discovery and validation process, which is essential for the detection of those structurally modified protein biomarkers.

Biomarkers, Tumor↗

ModifiComb, a new proteomic tool for mapping substoichiometric post-translational modifications, finding novel types of modifications, and fingerprinting complex protein mixtures.

A major challenge in proteomics is to fully identify and characterize the post-translational modification (PTM) patterns present at any given time in cells, tissues, and organisms. Here we present a fast and reliable method ("ModifiComb") for mapping hundreds types of PTMs at a time, including novel and unexpected PTMs. The high mass accuracy of Fourier transform mass spectrometry provides in many cases unique elemental composition of the PTM through the difference DeltaM between the molecular masses of the modified and unmodified peptides, whereas the retention time difference DeltaRT between their elution in reversed-phase liquid chromatography provides an additional dimension for PTM identification. Abundant sequence information obtained with complementary fragmentation techniques using ion-neutral collisions and electron capture often locates the modification to a single residue. The (DeltaM, DeltaRT) maps are representative of the proteome and its overall modification state and may be used for database-independent organism identification, comparative proteomic studies, and biomarker discovery. Examples of newly found modifications include +12.000 Da (+C atom) incorporation into proline residues of peptides from proline-rich proteins found in human saliva. This modification is hypothesized to increase the known activity of the peptide.

Adult↗

High throughput profile-profile based fold recognition for the entire human proteome.

BACKGROUND: In order to maintain the most comprehensive structural annotation databases we must carry out regular updates for each proteome using the latest profile-profile fold recognition methods. The ability to carry out these updates on demand is necessary to keep pace with the regular updates of sequence and structure databases. Providing the highest quality structural models requires the most intensive profile-profile fold recognition methods running with the very latest available sequence databases and fold libraries. However, running these methods on such a regular basis for every sequenced proteome requires large amounts of processing power. In this paper we describe and benchmark the JYDE (Job Yield Distribution Environment) system, which is a meta-scheduler designed to work above cluster schedulers, such as Sun Grid Engine (SGE) or Condor. We demonstrate the ability of JYDE to distribute the load of genomic-scale fold recognition across multiple independent Grid domains. We use the most recent profile-profile version of our mGenTHREADER software in order to annotate the latest version of the Human proteome against the latest sequence and structure databases in as short a time as possible. RESULTS: We show that our JYDE system is able to scale to large numbers of intensive fold recognition jobs running across several independent computer clusters. Using our JYDE system we have been able to annotate 99.9% of the protein sequences within the Human proteome in less than 24 hours, by harnessing over 500 CPUs from 3 independent Grid domains. CONCLUSION: This study clearly demonstrates the feasibility of carrying out on demand high quality structural annotations for the proteomes of major eukaryotic organisms. Specifically, we have shown that it is now possible to provide complete regular updates of profile-profile based fold recognition models for entire eukaryotic proteomes, through the use of Grid middleware such as JYDE.

Algorithms↗

Proteome analysis of rice uppermost internodes at the milky stage.

Uppermost internodes, which connect the part between the ear and lower stem, form an important pathway transporting mineral nutrition from roots and photosynthates from leaves (especially the flag leaf) to the ear. The milky stage is the first stage of seed ripening. The uppermost internodes of rice at the milky stage are critical for seed quality and yield. Total soluble proteins of the uppermost internodes of rice (Oryza sativa L. ssp. indica) at the milky stage were analyzed using proteomic methods. Using 2-DE, 762 reproducible protein spots were detected. Among them, 132 abundant proteins were analyzed using MALDI-TOF-MS. Searching in the National Center for Biotechnology Information database, we could identify 98 proteins, which represent 80 gene products. These proteins belong to 11 functional groups with energy production-associated proteins in the first place. The large accumulation of proteins involved in metabolism, signaling, and stress resistance indicated that the uppermost internodes of rice have a high physiological and stress-resistant activity. In addition, our results will also enrich the database of the rice proteome.

Electrophoresis, Gel, Two-Dimensional↗

Integration of gel-based proteome data with pProRep.

UNLABELLED: pProRep is a web application integrating electrophoretic and mass spectral data from proteome analyses into a relational database. The graphical web-interface allows users to upload, analyse and share experimental proteome data. It offers researchers the possibility to query all previously analysed datasets and can visualize selected features, such as the presence of a certain set of ions in a peptide mass spectrum, on the level of the two-dimensional gel. AVAILABILITY: The pProRep package and instructions for its use can be downloaded from http://www.ptools.ua.ac.be/pProRep. The application requires a web server that runs PHP 5 (http://www.php.net) and MySQL. Some (non-essential) extensions need additional freely available libraries: details are described in the installation instructions.

Computational Biology↗

Proteomic analysis of progressive factors in uterine cervical cancer.

Human papillomavirus (HPV) infections play a crucial role in the progress of cervical cancer. The high-risk HPV types are frequently associated with the development of malignant lesions. Some of the latest studies have demonstrated that the high-risk HPV 16 and 18 are predominantly detected in the more aggressive cancers. In the present study, we aimed to establish the proteomic profiles and characterization of the tumor related proteins by using two-dimensional gel electrophoresis (2-DE) and matrix-assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF MS). For proteomic analysis, patients infected by HPV 16 or 18 were included in this study. We compared nuclear protein and cytoplasmic protein, separately by using the subcellular fraction. Differential protein spots between cervical cancer with high-risk HPV, HPV 16 or HPV 18, and HaCaT cell lines were characterized by 2-DE. Those proteins analyzed by peptide mass fingerprinting based on MALDI-TOF MS and database searching were the products of oncogenes or proto-oncogenes, and the others were involved in the regulation of cell cycle, for general genomic stability, telomerase activation, and cell immortalization. However, there was no difference in protein characterization for cervical cancer between HPV 16 and HPV 18 infection. Nonetheless, these data are valuable for the mass identification of differentially expressed proteins involved in human uterine cervical cancer. Moreover, the data has enormous value for establishing the human uterine cervical cancer proteome database that can be used in screening a molecular marker for the further study of human uterine cervical cancer, and also for studying any correlation among the cancers induced by HPV.

Adult↗

Bioinformatics Resources for In Silico Proteome Analysis.

In the growing field of proteomics, tools for the in silico analysis of proteins and even of whole proteomes are of crucial importance to make best use of the accumulating amount of data. To utilise this data for healthcare and drug development, first the characteristics of proteomes of entire species-mainly the human-have to be understood, before secondly differentiation between individuals can be surveyed. Specialised databases about nucleic acid sequences, protein sequences, protein tertiary structure, genome analysis, and proteome analysis represent useful resources for analysis, characterisation, and classification of protein sequences. Different from most proteomics tools focusing on similarity searches, structure analysis and prediction, detection of specific regions, alignments, data mining, 2D PAGE analysis, or protein modelling, respectively, comprehensive databases like the proteome analysis database benefit from the information stored in different databases and make use of different protein analysis tools to provide computational analysis of whole proteomes.

Journal Article↗

An object model and database for functional genomics.

MOTIVATION: Large-scale functional genomics analysis is now feasible and presents significant challenges in data analysis, storage and querying. Data standards are required to enable the development of public data repositories and to improve data sharing. There is an established data format for microarrays (microarray gene expression markup language, MAGE-ML) and a draft standard for proteomics (PEDRo). We believe that all types of functional genomics experiments should be annotated in a consistent manner, and we hope to open up new ways of comparing multiple datasets used in functional genomics. RESULTS: We have created a functional genomics experiment object model (FGE-OM), developed from the microarray model, MAGE-OM and two models for proteomics, PEDRo and our own model (Gla-PSI-Glasgow Proposal for the Proteomics Standards Initiative). FGE-OM comprises three namespaces representing (i) the parts of the model common to all functional genomics experiments; (ii) microarray-specific components; and (iii) proteomics-specific components. We believe that FGE-OM should initiate discussion about the contents and structure of the next version of MAGE and the future of proteomics standards. A prototype database called RNA And Protein Abundance Database (RAPAD), based on FGE-OM, has been implemented and populated with data from microbial pathogenesis. AVAILABILITY: FGE-OM and the RAPAD schema are available from http://www.gusdb.org/fge.html, along with a set of more detailed diagrams. RAPAD can be accessed by registration at the site.

Abstracting and Indexing↗

General framework for developing and evaluating database scoring algorithms using the TANDEM search engine.

MOTIVATION: Tandem mass spectrometry (MS/MS) identifies protein sequences using database search engines, at the core of which is a score that measures the similarity between peptide MS/MS spectra and a protein sequence database. The TANDEM application was developed as a freely available database search engine for the proteomics research community. To extend TANDEM as a platform for further research on developing improved database scoring methods, we modified the software to allow users to redefine the scoring function and replace the native TANDEM scoring function while leaving the remaining core application intact. Redefinition is performed at run time so multiple scoring functions are available to be selected and applied from a single search engine binary. We introduce the implementation of the pluggable scoring algorithm and also provide implementations of two TANDEM compatible scoring functions, one previously described scoring function compatible with PeptideProphet and one very simple scoring function that quantitative researchers may use to begin their development. This extension builds on the open-source TANDEM project and will facilitate research into and dissemination of novel algorithms for matching MS/MS spectra to peptide sequences. The pluggable scoring schema is also compatible with related search applications P3 and Hunter, which are part of the X! suite of database matching algorithms. The pluggable scores and the X! suite of applications are all written in C++. AVAILABILITY: Source code for the scoring functions is available from http://proteomics.fhcrc.org

Algorithms↗

Proteomic analyses using an accurate mass and time tag strategy.

An accurate mass and time (AMT) tag approach for proteomic analyses has been developed over the past several years to facilitate comprehensive high-throughput proteomic measurements. An AMT tag database for an organism, tissue, or cell line is established by initially performing standard shotgun proteomic analysis and, most importantly, by validating peptide identifications using the mass measurement accuracy of Fourier transform ion cyclotron resonance (FTICR) mass spectrometry (MS) and liquid chromatography (LC) elution time constraint. Creation of an AMT tag database largely obviates the need for subsequent MS/MS analyses, and thus facilitates high-throughput analyses. The strength of this technology resides in the ability to achieve highly efficient and reproducible one-dimensional reversed-phased LC separations in conjunction with highly accurate mass measurements using FTICR MS. Recent improvements allow for the analysis of as little as picrogram amounts of proteome samples by minimizing sample handling and maximizing peptide recovery. The nanoproteomics platform has also demonstrated the ability to detect >10(6) differences in protein abundances and identify more abundant proteins from subpicogram amounts of samples. The AMT tag approach is poised to become a new standard technique for the in-depth and high-throughput analysis of complex organisms and clinical samples, with the potential to extend the analysis to a single mammalian cell.

Amino Acid Sequence↗

Proteomics of Halophilic archaea.

Halophilic archaea is a member of the Halobacteriacea family, the only family in the Halobacteriales order. Most Halophilic archaea require 1.5M NaCl both to grow and retain the structural integrity of the cells. The proteins of these organisms have thus been adapted to be active and stable in the hypersaline condition. Consequently, the unique properties of these biocatalysts have resulted in several novel applications in industrial processes. Halophilic archaea are also to be useful for bioremediation of hypersaline environment. Proteome data have expended enormously with the significant advance recently achieved in two-dimensional gel electrophoresis (2-DE) and mass spectrometry (MS). The whole genome sequencing of Halobacterium species NRC-1 was completed and this would also provide tremendous help to analyze the protein mass data from the similar strain Halobacterium salinarum. Proteomics coupled with genomic databases now has become a basic tool to understand or identify the function of genes and proteins. In addition, the bioinformatics approach will facilitate to predict the function of novel proteins of Halophilic archaea. This review will discuss current proteome study of Halophilic archaea and introduce the efficient procedures for screening, predicting, and confirming the function of novel halophilic enzymes.

Amino Acid Sequence↗

Structural proteomics: methods in deriving protein structural information and issues in data management.

Structural proteomics is an emerging paradigm that is gaining importance in the post-genomic era as a valuable discipline to process the protein target information being deciphered. The field plays a crucial role in assigning function to sequenced proteins, defining pathways in which the targets are involved, and understanding structure-function relationships of the protein targets. A key component of this research sector is accessing the three-dimensional structures of protein targets by both experimental and theoretical methods. This then leads to the question of how to store, retrieve, and manipulate vast amounts of sequence (1-D) and structural (3-D) information in a relational format so that extensive data analysis can be achieved. We at SBI have addressed both of these fundamental requirements of structural proteomics. We have developed an extensive collection of three-dimensional protein structures from sequence data and have implemented a relational architecture for data management. In this article we will discuss our approaches to structural proteomics and the tools that life science researchers can use in their discovery efforts.

Computational Biology↗

A reference map of a human pituitary adenoma proteome.

In order to compare the proteomes from different cell types of pituitary adenomas for our long-term goal to clarify the molecular mechanisms that participate in the formation of pituitary adenoma, and to detect any tumor-related marker for an "early-stage" diagnosis, the two-dimensional gel electrophoresis (2-DE) reference map of a pituitary adenoma tissue proteome is described here. A vertical, two-dimensional (2-D) polyacrylamide gel electrophoresis system and PDQuest image analysis software have been used to provide a high level of between-gel reproducibility and to accurately array each protein expressed in a pituitary adenoma tissue. Mass spectrometry (matrix-assisted laser desorption/ionization-time of flight MALDI-TOF and liquid chromatography-electrospray ionization-quadrupole-ion trap LC-ESI-Q-IT) and protein databases were used to characterize each protein in the 2-D gel. The results demonstrate that a good reproducibility of the 2-D gel pattern was attained. The position deviation of matched spots among four 2-D gels was 1.95 +/- 0.45 mm in the isoelectric focusing direction, and 1.70 +/- 0.53 mm in the sodium dodecyl sulfate-polyacrylamide gel electrophoresis direction. A total of ca. 1000 protein spots were separated by 2-DE, and 135 protein spots that represent 111 proteins were characterized with mass spectrometry (96 spots for MALDI-TOF, 39 spots for LC-ESI-Q-IT). The characterized proteins include pituitary hormones, cellular signals, enzymes, cellular-defense proteins, cell-structure proteins, transport proteins, etc. Those proteins were located in the cytoplasmic, cellular membrane, mitochondrial, endoplasmic reticulum, nuclear, ribonucleosome, extracellular fractions, or were secreted in plasma, etc. Those identified proteins contribute to a functional profile of the pituitary adenoma proteome. These data will be used to expand the proteome database of the human pituitary, which can be accessed in the website http://www.utmem.edu /proteomics.

Adenoma↗

Analysis of gene ontology features in microarray data using the Proteome BioKnowledge Library.

Microarray technology has resulted in an explosion of complex, valuable data. Integrating data analysis tools with a comprehensive underlying database would allow efficient identification of common properties among differentially regulated genes. In this study we sought to compare the utility of various databases in microarray analysis. The Proteome BioKnowledge Library (BKL), a manually curated, proteome-wide compilation of the scientific literature, was used to generate a list of Gene Ontology (GO) Biological Process (BP) terms enriched among proteins involved in cardiovascular disease. Analysis of DNA microarray data generated in a study of rat vascular smooth muscle cell responses revealed significant enrichment in a number of GO BPs that were also enriched among cardiovascular disease-related proteins. Using annotation from LocusLink and chip annotation from the Gene Expression Omnibus yielded fewer enriched cardiovascular disease-associated GO BP terms. Data sets of orthologous genes from mouse and human were generated using the BKL Retriever. Analysis of these sets focusing on BKL Disease annotation, revealed a significant association of these genes with cardiovascular disease. These results and the extensive presence of experimental evidence for BKL GO and Disease features, underscore the benefits of using this database for microarray analysis.

Animals↗

Clinical proteomics: present and future prospects.

Advances in proteomics technology offer great promise in the understanding and treatment of the molecular basis of disease. The past decade of proteomics research, the study of dynamic protein expression, post-translational modifications, cellular and sub-cellular protein distribution, and protein-protein interactions, has culminated in the identification of many disease-related biomarkers and potential new drug targets. While proteomics remains the tool of choice for discovery research, new innovations in proteomic technology now offer the potential for proteomic profiling to become standard practice in the clinical laboratory. Indeed, protein profiles can serve as powerful diagnostic markers, and can predict treatment outcome in many diseases, in particular cancer. A number of technical obstacles remain before routine proteomic analysis can be achieved in the clinic; however the standardisation of methodologies and dissemination of proteomic data into publicly available databases is starting to overcome these hurdles. At present the most promising application for proteomics is in the screening of specific subsets of protein biomarkers for certain diseases, rather than large scale full protein profiling. Armed with these technologies the impending era of individualised patient-tailored therapy is imminent. This review summarises the advances in proteomics that has propelled us to this exciting age of clinical proteomics, and highlights the future work that is required for this to become a reality.

Journal Article↗