Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Profiling core proteomes of human cell lines by one-dimensional PAGE and liquid chromatography-tandem mass spectrometry.

Protein expression profiles vary considerably between human cell lines and tissues, which is in part a reflection of their specialized roles within an organism. It is of considerable practical use to establish which proteins constitute the primary components of the respective proteomes. When compiled into databases, such information can facilitate the assessment of selectivity and specificity of a wide range of proteomic experiments. Here we describe the major constituents of proteomes of six human immortalized cell lines. By employing a combination of one-dimensional SDS-PAGE and nanocapillary liquid chromatography-tandem mass spectrometry (LC-MS/MS), we identified up to 1785 non-redundant cytoplasmic and nuclear proteins from a single cell line using 50 and 30 microg of total protein from the corresponding fractions. Up to 38 proteins could be identified from a single band in one liquid chromatography-MS/MS experiment. When combined with systematic gridding of gel lanes into 48 slices, a dynamic range for protein identification of approximately 1:2000 can be envisaged for this approach. Identified proteins range from 4-553 kDa in size, cover the pI range between 3.4 and 12.8, and include 255 proteins with predicted transmembrane domains. Repeated analysis of peptides derived from the same gel band showed that the reproducibility of nanocapillary liquid chromatography-MS/MS of such complex mixtures is about 60-70% suggesting that a particular analytical experiment would need to be repeated about three times to arrive at a representative estimate of the set of highly abundant proteins in a given proteome. Given its technical simplicity, sensitivity, and wealth of generated information, we have adopted this experimental approach to characterize every cell line and tissue that is the subject of experimentation in our laboratory. The combined dataset for the six cell lines consists of 2341 non-redundant human proteins and thus constitutes one of the largest collections of human proteomic data published to date.

Cell Line↗

Integrative proteomics: structure, function, and interaction report on the 3rd joint meeting of the British Society for Proteome Research and the European Bioinformatics Institute, July 2006.

This report summarizes the highlights of the recent British Society for Proteome Research (BSPR) meeting jointly organized with the European Bioinformatics Institute (EBI) which was held at the Wellcome Trust Genome Campus, Hinxton, Cambridge, UK in July 2006. This was the third annual scientific meeting organized by the BSPR and EBI and the theme of this years meeting was Integrative Proteomics: Structure, function and interaction. A wealth of local and overseas speakers were invited to discuss both their own work and specific challenges present in modern day proteomic based experiments.

Computational Biology↗

Integrated approach for manual evaluation of peptides identified by searching protein sequence databases with tandem mass spectra.

Quantitative proteomics relies on accurate protein identification, which often is carried out by automated searching of a sequence database with tandem mass spectra of peptides. When these spectra contain limited information, automated searches may lead to incorrect peptide identifications. It is therefore necessary to validate the identifications by careful manual inspection of the mass spectra. Not only is this task time-consuming, but the reliability of the validation varies with the experience of the analyst. Here, we report a systematic approach to evaluating peptide identifications made by automated search algorithms. The method is based on the principle that the candidate peptide sequence should adequately explain the observed fragment ions. Also, the mass errors of neighboring fragments should be similar. To evaluate our method, we studied tandem mass spectra obtained from tryptic digests of E. coli and HeLa cells. Candidate peptides were identified with the automated search engine Mascot and subjected to the manual validation method. The method found correct peptide identifications that were given low Mascot scores (e.g., 20-25) and incorrect peptide identifications that were given high Mascot scores (e.g., 40-50). The method comprehensively detected false results from searches designed to produce incorrect identifications. Comparison of the tandem mass spectra of synthetic candidate peptides to the spectra obtained from the complex peptide mixtures confirmed the accuracy of the evaluation method. Thus, the evaluation approach described here could help boost the accuracy of protein identification, increase number of peptides identified, and provide a step toward developing a more accurate next-generation algorithm for protein identification.

Algorithms↗

RADARS, a bioinformatics solution that automates proteome mass spectral analysis, optimises protein identification, and archives data in a relational database.

RADARS, a rapid, automated, data archiving and retrieval software system for high-throughput proteomic mass spectral data processing and storage, is described. The majority of mass spectrometer data files are compatible with RADARS, for consistent processing. The system automatically takes unprocessed data files, identifies proteins via in silico database searching, then stores the processed data and search results in a relational database suitable for customized reporting. The system is robust, used in 24/7 operation, accessible to multiple users of an intranet through a web browser, may be monitored by Virtual Private Network, and is secure. RADARS is scalable for use on one or many computers, and is suited to multiple processor systems. It can incorporate any local database in FASTA format, and can search protein and DNA databases online. A key feature is a suite of visualisation tools (many available gratis), allowing facile manipulation of spectra, by hand annotation, reanalysis, and access to all procedures. We also described the use of Sonar MS/MS, a novel, rapid search engine requiring 40 MB RAM per process for searches against a genomic or EST database translated in all six reading frames. RADARS reduces the cost of analysis by its efficient algorithms: Sonar MS/MS can identifiy proteins without accurate knowledge of the parent ion mass and without protein tags. Statistical scoring methods provide close-to-expert accuracy and brings robust data analysis to the non-expert user.

Amino Acid Sequence↗

ExPASy: The proteomics server for in-depth protein knowledge and analysis.

The ExPASy (the Expert Protein Analysis System) World Wide Web server (http://www.expasy.org), is provided as a service to the life science community by a multidisciplinary team at the Swiss Institute of Bioinformatics (SIB). It provides access to a variety of databases and analytical tools dedicated to proteins and proteomics. ExPASy databases include SWISS-PROT and TrEMBL, SWISS-2DPAGE, PROSITE, ENZYME and the SWISS-MODEL repository. Analysis tools are available for specific tasks relevant to proteomics, similarity searches, pattern and profile searches, post-translational modification prediction, topology prediction, primary, secondary and tertiary structure analysis and sequence alignment. These databases and tools are tightly interlinked: a special emphasis is placed on integration of database entries with related resources developed at the SIB and elsewhere, and the proteomics tools have been designed to read the annotations in SWISS-PROT in order to enhance their predictions. ExPASy started to operate in 1993, as the first WWW server in the field of life sciences. In addition to the main site in Switzerland, seven mirror sites in different continents currently serve the user community.

Databases, Protein↗

[Establishment of 2-dE map of human low differentiation nasopharyngeal carcinoma cell line CNE-2 proteome].

OBJECTIVE: To establish a two-dimensional polyacrylamide gel electrophoresis (2-DE) map of human low differentiation nasopharyngeal carcinoma (NPC) cell line CNE-2 proteome. METHODS: Immobilized pH gradient 2-DE was applied to separate the total proteins of CNE-2 cells, matrix-assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF-MS) and database searching were used to indentify proteins. RESULTS: A good 2-DE pattern with high resolution and reproducibility was obtained. Seventy-eight protein spots were incised from sliver staining gel and digested in gel by trypisin. Of them 77 maps of peptide mass fingerprints (PMF) were obtained and 48 proteins were identified. CONCLUSION: A reference 2-DE map of CNE-2 proteome has been established and the data coupled with similar proteome analysis of other NPC cell lines will expand the human nasopharyngeal carcinoma proteome database.

Amino Acid Sequence↗

Comparative study of apoptosis-related gene loci in human, mouse and rat genomes.

Many genes are involved in mammalian cell apoptosis pathway. These apoptosis genes often contain characteristic functional domains, and can be classified into at least 15 functional groups, according to previous reports. Using an integrated bioinformatics platform for motif or domain search from three public mammalian proteomes (International Protein Index database for human, mouse, and rat), we systematically cataloged all of the proteins involved in mammalian apoptosis pathway. By localizing those proteins onto the genomes, we obtained a gene locus centric apoptosis gene catalog for human, mouse and rat. Further phylogenetic analysis showed that most of the apoptosis related gene loci are conserved among these three mammals. Interestingly, about one-third of apoptosis gene loci form gene clusters on mammal chromosomes, and exist in the three species, which indicated that mammalian apoptosis gene orders are also conserved. In addition, some tandem duplicated gene loci were revealed by comparing gene loci clusters in the three species. All data produced in this work were stored in a relational database and may be viewed at http://pcas.cbi.pku.edu.cn/database/apd.php.

Animals↗

GermOnline, a cross-species community knowledgebase on germ cell differentiation.

GermOnline provides information and microarray expression data for genes involved in mitosis and meiosis, gamete formation and germ line development across species. The database has been developed, and is being curated and updated, by life scientists in cooperation with bioinformaticists. Information is contributed through an online form using free text, images and the controlled vocabulary developed by the GeneOntology Consortium. Authors provide up to three references in support of their contribution. The database is governed by an international board of scientists to ensure a standardized data format and the highest quality of GermOnline's information content. Release 2.0 provides exclusive access to microarray expression data from Saccharomyces cerevisiae and Rattus norvegicus, as well as curated information on approximately 700 genes from various organisms. The locus report pages include links to external databases that contain relevant annotation, microarray expression and proteome data. Conversely, the Saccharomyces Genome Database (SGD), S.cerevisiae GeneDB and Swiss-Prot link to the budding yeast section of GermOnline from their respective locus pages. GermOnline, a fully operational prototype subject-oriented knowledgebase designed for community annotation and array data visualization, is accessible at http://www.germonline.org. The target audience includes researchers who work on mitotic cell division, meiosis, gametogenesis, germ line development, human reproductive health and comparative genomics.

Animals↗

Computational methods for comparison of large genomic and proteomic datasets reveal protein markers of metastatic cancer.

Large-scale genomic and proteomic analysis has provided a wealth of information on biologically relevant systems, and the ability to analyze this information is crucial to uncovering important biological relationships. However, it has proven difficult to compare large datasets from different sources due to different gene and protein identifiers assigned by individual laboratories and database systems. Here, we describe the design of a fully automated blast program (BlastPro) that facilitates rapid comparison of large protein-protein, nucleotide--nucleotide, or nucleotide--protein datasets from numerous, independent studies. Using this system, we compared several published genomic and proteomic databases for proteins that are upregulated in highly motile, metastatic tumor cells. Analysis of five independent studies comprised of greater than 1 x 10(6) genomic sequences and greater than 1,000 proteins revealed that the cytoskeletal-associated protein alpha-actinin is increased at both the mRNA and protein level in metastatic breast, prostate, and skin cancer cells. Interestingly, spatial analysis of alpha-actinin expression revealed that it is amplified 8-fold in the leading pseudopodium compared to the cell body compartment of migrating cells. These findings indicate that amplification of alpha-actinin and its localization to the leading pseudopodium are potential biomarkers of cancer progression to a more metastatic phenotype. Together, our results demonstrate that the BlastPro system can be used to compare large genomic and proteomic datasets to reveal important biological relationships including those associated with cancer progression.

Actinin↗

Coverage of whole proteome by structural genomics observed through protein homology modeling database.

We have been developing FAMSBASE, a protein homology-modeling database of whole ORFs predicted from genome sequences. The latest update of FAMSBASE ( http://daisy.nagahama-i-bio.ac.jp/Famsbase/ ), which is based on the protein three-dimensional (3D) structures released by November 2003, contains modeled 3D structures for 368,724 open reading frames (ORFs) derived from genomes of 276 species, namely 17 archaebacterial, 130 eubacterial, 18 eukaryotic and 111 phage genomes. Those 276 genomes are predicted to have 734,193 ORFs in total and the current FAMSBASE contains protein 3D structure of approximately 50% of the ORF products. However, cases that a modeled 3D structure covers the whole part of an ORF product are rare. When portion of an ORF with 3D structure is compared in three kingdoms of life, in archaebacteria and eubacteria, approximately 60% of the ORFs have modeled 3D structures covering almost the entire amino acid sequences, however, the percentage falls to about 30% in eukaryotes. When annual differences in the number of ORFs with modeled 3D structure are calculated, the fraction of modeled 3D structures of soluble protein for archaebacteria is increased by 5%, and that for eubacteria by 7% in the last 3 years. Assuming that this rate would be maintained and that determination of 3D structures for predicted disordered regions is unattainable, whole soluble protein model structures of prokaryotes without the putative disordered regions will be in hand within 15 years. For eukaryotic proteins, they will be in hand within 25 years. The 3D structures we will have at those times are not the 3D structure of the entire proteins encoded in single ORFs, but the 3D structures of separate structural domains. Measuring or predicting spatial arrangements of structural domains in an ORF will then be a coming issue of structural genomics.

Amino Acid Sequence↗

Proteomic resources: integrating biomedical information in humans.

Recent improvements in high-throughput proteomic technologies have unleashed the potential for generating vast amounts of data. Managing and sharing proteomic data is not an easy task. In this article, we will discuss some of the high-throughput proteomic techniques that are commonly used today. We will also review the major issues in sharing and dissemination of proteomic data and the recent community initiatives to standardize data formats and ontologies. An overview of the web-based resources and databases for analysis of proteomic data is also provided. Integration of disparate proteomic data sources with genomic and transcriptomic data should make systems biology type of approaches feasible in the near future.

Amino Acid Sequence↗

Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.

Protein identification via peptide mass fingerprinting (PMF) remains a key component of high-throughput proteomics experiments in post-genomic science. Candidate protein identifications are made using bioinformatic tools from peptide peak lists obtained via mass spectrometry (MS). These algorithms rely on several search parameters, including the number of potential uncut peptide bonds matching the primary specificity of the hydrolytic enzyme used in the experiment. Typically, up to one of these "missed cleavages" are considered by the bioinformatics search tools, usually after digestion of the in silico proteome by trypsin. Using two distinct, nonredundant datasets of peptides identified via PMF and tandem MS, a simple predictive method based on information theory is presented which is able to identify experimentally defined missed cleavages with up to 90% accuracy from amino acid sequence alone. Using this simple protocol, we are able to "mask" candidate protein databases so that confident missed cleavage sites need not be considered for in silico digestion. We show that that this leads to an improvement in database searching, with two different search engines, using the PMF dataset as a test set. In addition, the improved approach is also demonstrated on an independent PMF data set of known proteins that also has corresponding high-quality tandem MS data, validating the protein identifications. This approach has wider applicability for proteomics database searching, and the program for predicting missed cleavages and masking Fasta-formatted protein sequence databases has been made available via http:// ispider.smith.man.ac uk/MissedCleave.

Algorithms↗

LIP index for peptide classification using MS/MS and SEQUEST search via logistic regression.

This study addresses the issue of peptide identification resulting from tandem mass spectrometry proteomics analysis followed by database search. This work shows that the Logistic Identification of Peptides (LIP) Index achieves high sensitivity and specificity for peptide classification relative to a manually verified "gold" standard and also accurately estimates the probability of a correct peptide match. The LIP Index is a weighted average of SEQUEST output variables based on logistic regression models and is a transparent, easy to use, inclusive, extendable, and statistically sound approach to classify correct peptide identifications. Modifications, such as normalizing cross-correlations (Xcorr) for peptide length, adjusting for charge state, and the number of tryptic termini, significantly improve the fit the logistic regression models, as well as increase sensitivity and specificity. The LIP Index also incorporates earlier developed statistical models on spectral quality assessment and peptide identification, which further improves sensitivity and specificity.

Algorithms↗

SPLASH: systematic proteomics laboratory analysis and storage hub.

In the field of proteomics, the increasing difficulty to unify the data format, due to the different platforms/instrumentation and laboratory documentation systems, greatly hinders experimental data verification, exchange, and comparison. Therefore, it is essential to establish standard formats for every necessary aspect of proteomics data. One of the recently published data models is the proteomics experiment data repository [Taylor, C. F., Paton, N. W., Garwood, K. L., Kirby, P. D. et al., Nat. Biotechnol. 2003, 21, 247-254]. Compliant with this format, we developed the systematic proteomics laboratory analysis and storage hub (SPLASH) database system as an informatics infrastructure to support proteomics studies. It consists of three modules and provides proteomics researchers a common platform to store, manage, search, analyze, and exchange their data. (i) Data maintenance includes experimental data entry and update, uploading of experimental results in batch mode, and data exchange in the original PEDRo format. (ii) The data search module provides several means to search the database, to view either the protein information or the differential expression display by clicking on a gel image. (iii) The data mining module contains tools that perform biochemical pathway, statistics-associated gene ontology, and other comparative analyses for all the sample sets to interpret its biological meaning. These features make SPLASH a practical and powerful tool for the proteomics community.

Database Management Systems↗

Proteomics reveals protein profile changes in doxorubicin--treated MCF-7 human breast cancer cells.

MCF-7 cells are extensively used as a cell model to investigate human breast tumors and the cellular mechanism of antitumor drugs such as doxorubicin (DOX), an anthracycline antitumor drug widely used in clinical chemotherapy. To understand the effects of DOX on the protein expression, we perform a comprehensive proteomics to survey global changes in proteins after DOX treatment in MCF-7 cells. Exposure of MCF-7 cells to 0.1 microM DOX for 2 days induced a differentiation-like phenotype with prominent perinuclear autocatalytic vacuoles, abundant filamentous material, and irregular microvilli at the cell surface. In this study, we also present a proteome reference map of MCF-7 cells with 21 identified protein spots via analysis of N-terminal sequencing, mass spectrometry, immunoblot and/or computer matching with protein database. Based on the proteome map, we found that DOX causes a markedly decrease in the levels of three isoforms of heat shock protein 27 (HSP27) whereas the levels of other stress associated proteins including HSP60, calreticulin, and protein disulfide isomerase were not significantly altered in DOX-treated MCF-7 cells. Taken together, we suggest that that action of DOX on breast tumor cells may be partly related to dysregulation of HSP27 expression. Modulation of HSP27 levels may be a clinically useful potential target for design of antitumor drugs and controlling breast tumor growth.

Antineoplastic Agents↗

Comparative proteomics analysis of human lung squamous carcinoma.

Two-dimensional polyacrylamide gel electrophoresis (2-DE) profiles of human lung squamous carcinoma tissue and paired surrounding normal bronchial epithelial tissue were compared. Selected differential protein-spots were identified with peptide mass fingerprinting based on matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF-MS) and database searching. Well-resolved and reproducible 2-DE patterns of both the tumor and the normal tissues were acquired. The average deviations of spot position were 0.873+/-0.125mm in IEF direction and 1.025+/-0.213mm in SDS-PAGE direction, respectively. For the tumor tissues, a total of 1349+/-67 spots were detected and 1235+/-48 spots were matched with an average matching rate of 91.5%. For the corresponding normal tissues, a total of 1297+/-73 spots were detected and 1183+/-56 spots were matched with an average matching rate of 91.2%. A total of 1069+/-45 spots were matched between the tumor and the normal tissues. Forty differential proteins between tumor and normal tissues were characterized. Some proteins were the products of oncogenes and others were involved in the regulation of cell cycle and signal transduction. These data are valuable for mass identification of differentially expressed proteins involved in lung carcinogenesis, establishing human lung cancer proteome database and screening molecular marker to further study human lung squamous carcinoma.

Carcinoma, Squamous Cell↗

A role for Edman degradation in proteome studies.

Advances in protein database design and the software used to access the sequence data has led to progress in using protein attributes such as amino acid composition and peptide masses to identify proteins separated by two-dimensional electrophoresis. However, Edman degradation remains the principal technique for protein identification and it presents a significant bottleneck in the progress towards rapid protein identification. Simple modifications to the sequencing hardware, which automate the delivery of protein spots into the sequencer, and parallel sequencing of the protein spots represent a significant advance in the use of Edman degradation to rapidly generate the powerful protein attribute, an N-terminal sequence tag.

Amino Acid Sequence↗