Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Proteomic analysis of the human colon carcinoma cell line (LIM 1215): development of a membrane protein database.

The proteomic definition of plasma membrane proteins is an important initial step in searching for novel tumor marker proteins expressed during the different stages of cancer progression. However, due to the charge heterogeneity and poor solubility of membrane-associated proteins this subsection of the cell's proteome is often refractory to two-dimensional electrophoresis (2-DE), the current paradigm technology for studying protein expression profiles. Here, we describe a non-2-DE method for identifying membrane proteins. Proteins from an enriched membrane preparation of the human colorectal carcinoma cell line LIM1215 were initially fractionated by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE, 4-20%). The unstained gel was cut into 16 x 3 mm slices, and peptide mixtures resulting from in-gel tryptic digestion of each slice were individually subjected to capillary-column reversed phase-high performance liquid chromatography (RP-HPLC) coupled with electrospray ionization-ion trap-mass spectrometry (ESI-IT-MS). Interrogation of genomic databases with the resulting collision-induced dissociation (CID) generated peptide ion fragment data was used to identify the proteins in each gel slice. Over 284 proteins (including 92 membrane proteins) were identified, including many integral membrane proteins not previously identified by 2-DE, many proteins seen at the genomic level only, as well as several proteins identified by expressed sequence tags (ESTs) only. Additionally, a number of peptides, identified by de novo MS sequence analysis, have not been described in the databases. Further, a "targeted" ion approach was used to unambiguously identify known low-abundance plasma membrane proteins, using the membrane-associated A33 antigen, a gastrointestinal-specific epithelial cell protein, as an example. Following localization of the A33 antigen in the gel by immunoblotting, ions corresponding to the theoretical A33 antigen tryptic peptide masses were selected using an "inclusion" mass list for automated sequence analysis. Six peptides corresponding to the A33 antigen, present at levels well below those accessible using the standard automated "nontargeted" approach, were identified. The membrane protein database may be accessed via the World Wide Web (WWW) at http://www.ludwig. edu.au/jpsl/jpslhome.html.

Colonic Neoplasms↗

The PROTICdb database for 2-DE proteomics.

PROTICdb is a web-based database mainly designed to store and analyze plant proteome data obtained by 2D polyacrylamide gel electrophoresis (2D PAGE) and mass spectrometry (MS). The goals of PROTICdb are (1) to store, track, and query information related to proteomic experiments, i.e., from tissue sampling to protein identification and quantitative measurements; and (2) to integrate information from the user's own expertise and other sources into a knowledge base, used to support data interpretation (e.g., for the determination of allelic variants or products of posttranslational modifications). Data insertion into the relational database of PROTICdb is achieved either by uploading outputs from Mélanie, PDQuest, IM2d, ImageMaster(tm) 2D Platinum v5.0, Progenesis, Sequest, MS-Fit, and Mascot software, or by filling in web forms (experimental design and methods). 2D PAGE-annotated maps can be displayed, queried, and compared through the GelBrowser. Quantitative data can be easily exported in a tabulated format for statistical analyses with any third-party software. PROTICdb is based on the Oracle or the PostgreSQLDataBase Management System (DBMS) and is freely available upon request at http://cms.moulon.inra.fr/content/view/14/44/.

Databases, Protein↗

Protein analysis by mass spectrometry and sequence database searching: a proteomic approach to identify human lymphoblastoid cell line proteins.

Lymphoblastoid cell lines correspond to in vitro EBV-immortalized lymphocyte B-cells. These cells display a suitable model for experiments dealing with changes in protein expression occurring upon B-cell differentiation, after drug treatment, or after inhibition of some transcription factors. For all these reasons we have undertaken an effort aimed at developing a hematopoietic cell line protein two-dimensional electrophoresis (2-DE) database, containing B-lymphoblastoid 2-DE maps. In this work, matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF-MS) peptide mass fingerprinting analysis was adopted for protein identification. The peptide mass fingerprinting identification and the sequence coverage obtained on colloidal Coomassie blue (CBB) stained gel was close to that obtained using zinc-imidazole staining. Everything considered, CBB being more comfortable for subsequent spot manipulations, CBB staining was chosen for identification of a larger number of polypeptides. The results suggest that reticulation of the gel can interfere preventing the uptake of the enzyme during the in-gel digestion step. Consequently, low molecular mass proteins appear more difficult to identify by mass fingerprinting. Finally, the information provided in this study allows the construction of a new annoted reference map of human lymphoblastoid cell proteins. Among the identified proteins 60% were not yet positioned on 2-DE maps in three of the most important well-documented databases. The annoted map will be accessible via Internet on the LBPP server at URL:http:// www-smbh.univ-paris13.fr/lbtp/index.htm.

Acrylic Resins↗

Human plasma proteome analysis by reversed sequence database search and molecular weight correlation based on a bacterial proteome analysis.

In shotgun proteomics, proteins can be fractionated by 1-D gel electrophoresis and digested into peptides, followed by liquid chromatography to separate the peptide mixture. Mass spectrometry generates hundreds of thousands of tandem mass spectra from these fractions, and proteins are identified by database searching. However, the search scores are usually not sufficient to distinguish the correct peptides. In this study, we propose a confident protein identification method for high-throughput analysis of human proteome. To build a filtering protocol in database search, we chose Pseudomonas putida KT2440 as a reference because this bacterial proteome contains fewer modifications and is simpler than the human proteome. First, the P. putida KT2440 proteome was filtered by reversed sequence database search and correlated by the molecular weight in 1-D-gel band positions. The characterization protocol was then applied to determine the criteria for clustering of the human plasma proteome into three different groups. This protein filtering method, based on bacterial proteome data analysis, represents a rapid way to generate higher confidence protein list of the human proteome, which includes some of heavily modified and cleaved proteins.

Blood Proteins↗

The BPP (protein biochemistry and proteomics) two-dimensional electrophoresis database.

The BPP (protein biochemistry and proteomics) two-dimensional electrophoresis (2-DE) database (http://www-smbh.univ-paris13.fr/lbtp/Biochemistry/Biochimie/bque.htm) was established in 1998. The current release contains 11 reference maps from human hematopoietic and lymphoid cell line samples. These reference maps have now 255 identified spots, corresponding to 84 protein entries. The World Wide Web (WWW) presentation is designed to allow public access to the available 2-DE data together with logical connections to databases providing complementary information.

Databases, Factual↗

SPINE 2: a system for collaborative structural proteomics within a federated database framework.

We present version 2 of the SPINE system for structural proteomics. SPINE is available over the web at http://nesg.org. It serves as the central hub for the Northeast Structural Genomics Consortium, allowing collaborative structural proteomics to be carried out in a distributed fashion. The core of SPINE is a laboratory information management system (LIMS) for key bits of information related to the progress of the consortium in cloning, expressing and purifying proteins and then solving their structures by NMR or X-ray crystallography. Originally, SPINE focused on tracking constructs, but, in its current form, it is able to track target sample tubes and store detailed sample histories. The core database comprises a set of standard relational tables and a data dictionary that form an initial ontology for proteomic properties and provide a framework for large-scale data mining. Moreover, SPINE sits at the center of a federation of interoperable information resources. These can be divided into (i) local resources closely coupled with SPINE that enable it to handle less standardized information (e.g. integrated mailing and publication lists), (ii) other information resources in the NESG consortium that are inter-linked with SPINE (e.g. crystallization LIMS local to particular laboratories) and (iii) international archival resources that SPINE links to and passes on information to (e.g. TargetDB at the PDB).

Cooperative Behavior↗

AnoXcel: an Anopheles gambiae protein database.

The proteome of the mosquito Anopheles gambiae was organized on a hyperlinked spreadsheet format containing one protein per row and several pieces of information in each column. The information for each protein ranges from the presence or absence of signal peptide indicative of secretion, presence of transmembrane domains, similarities to several databases, chromosomal location, and relatedness to other An. gambiae proteins, etc. Hosted by AnoBase (http://www.anobase.org/), the whole spreadsheet or segments of it can be downloaded or searched from http://www.anobase.org/AnoBase/Genes/Ano-Xcel by the scientist dealing with the annotation of proteome subsets such as those deriving from transcriptomes, nucleotide microarrays or high throughput mass spectrometry data.

Animals↗

ProteoformDB: A Built-In Application to Generate Proteoform Database.

Proteins play essential functions through their complex regulations on cell-type-specific expression, localization, and molecular complexes. Protein complexity is further enhanced by proteoforms, which are the diverse molecular forms that each gene can produce through genomic alterations, transcriptional variations, translational regulations, and protein modifications. Profiling of proteoforms is a promising method for gaining a deeper understanding of the role of proteins in biological pathways and disease mechanisms. Here, we developed ProteoformDB, an application tool for generating proteoform databases, and we cataloged a total of over one million unique single-site human proteoforms. We showed that ProteoformDB can serve as a valuable resource to document the experimentally identified proteoforms in a database, supporting protein characterization in quantitative proteomics for both total protein abundances and modified protein forms.

Humans↗

Multidimensional liquid chromatography separation of intact proteins by chromatographic focusing and reversed phase of the human serum proteome: optimization and protein database.

In biomarker discovery, the detection of proteins with low abundance in the serum proteome can be achieved by optimization of protein separation methods as well as selective depletion of the higher abundance proteins such as immunoglobins (e.g. IgG) and albumin. A relative newcomer to the proteomic separation arena is the commercial instrument PF2D from Beckman Coulter that separates proteins in the first dimension using chromatofocusing followed in line by reversed phase chromatography in the second dimension, thereby separating intact proteins based on pI and hydrophobicity. In this study, assessment and optimization of serum separation (undepleted serum and albumin-IgG-depleted serum) by the PF2D is presented. Protein databases were created for serum obtained from a healthy individual under traditional and optimized methods and under different sample preparation protocols. Separation of the doubly depleted serum using the PF2D with 20% isopropanol present in the first dimension running buffer allowed us to unambiguously identify 150 non-redundant serum proteins (excluding all immunoglobulin and albumin, a minimum of two peptide matches with acceptable Mascot score) in which 81 have not been identified previously in serum. Among them, numerous cellular proteins were identified to be specifically the skeletal muscle isoform, such as skeletal muscle fast twitch isoforms of troponin T, myosin alkali light chain 1, and sarcoplasmic/endoplasmic reticulum calcium ATPase. The detection of specific skeletal muscle protein isoforms in the serum from healthy individuals reflects the physiological turnover that occurs in skeletal muscle, which will have an impact on the ability to use generic "cellular" proteins as biomarkers without further characterization of the precise isoforms or post-translational modifications present.

Albumins↗

Interactive InterPro-based comparisons of proteins in whole genomes.

MOTIVATION: The SWISS-PROT group at the EBI has developed the Proteome Analysis Database utilizing existing resources and providing comprehensive and integrated comparative analysis of the predicted protein coding sequences of the complete genomes of bacteria, archaea and eukaryotes. The Proteome Analysis Database is accompanied by a program that has been designed to carry out interactive InterPro proteome comparisons for any one proteome against any other one or more of the proteomes in the database.

Computational Biology↗

Two-dimensional map of the proteome of Haemophilus influenzae.

We have constructed a two-dimensional database of the proteome of Haemophilus influenzae, a bacterium of medical interest of which the complete genome, comprising about 1742 open reading frames, has been sequenced. The soluble protein fraction of the microorganism was analyzed by two-dimensional electrophoresis, using immobilized pH gradient strips of various pH regions, gels with different acrylamide concentrations and buffers with different trailing ions. In order to visualize low-copy-number gene products, we employed a series of protein extraction and sample application approaches and several chromatographic steps, including heparin chromatography, chromatofocusing and hydrophobic interaction chromatography. We have also analyzed the cell envelope-bound protein fraction using either immobilized pH gradient strips or a two-detergent system with a cationic detergent in the first and an anionic detergent in the second-dimensional separation. Different proteins (502) were identified by matrix-assisted laser desorption/ionization mass spectrometry and amino acid composition analysis. This is at present one of the largest two-dimensional proteome databases.

Databases, Factual↗

PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.

Peroxisomes are essential organelles of eukaryotic origin, ubiquitously distributed in cells and organisms, playing key roles in lipid and antioxidant metabolism. Loss or malfunction of peroxisomes causes more than 20 fatal inherited conditions. We have created a peroxisomal database (http://www.peroxisomeDB.org) that includes the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae, by gathering, updating and integrating the available genetic and functional information on peroxisomal genes. PeroxisomeDB is structured in interrelated sections 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases', that include hyperlinks to selected features of NCBI, ENSEMBL and UCSC databases. We have designed graphical depictions of the main peroxisomal metabolic routes and have included updated flow charts for diagnosis. Precomputed BLAST, PSI-BLAST, multiple sequence alignment (MUSCLE) and phylogenetic trees are provided to assist in direct multispecies comparison to study evolutionary conserved functions and pathways. Highlights of the PeroxisomeDB include new tools developed for facilitating (i) identification of novel peroxisomal proteins, by means of identifying proteins carrying peroxisome targeting signal (PTS) motifs, (ii) detection of peroxisomes in silico, particularly useful for screening the deluge of newly sequenced genomes. PeroxisomeDB should contribute to the systematic characterization of the peroxisomal proteome and facilitate system biology approaches on the organelle.

Animals↗

PEP: Predictions for Entire Proteomes.

PEP is a database of Predictions for Entire Proteomes. The database contains summaries of analyses of protein sequences from a range of organisms representing all three major kingdoms of life: eukaryotes, prokaryotes and archaea. All proteins publicly available for organisms were aligned against SWISS-PROT, TrEMBL and PDB. Additionally, the following annotations are provided: secondary structure, transmembrane helices, coiled coils, regions of low complexity, signal peptides, PROSITE motifs, nuclear localization signals and classes of cellular function. Proteins that contain long regions without regular secondary structure are also identified. We have produced a related database of structural domain-like fragments derived from PEP and clusters based on homology between all fragments. The PEP database, fragments and clusters are distributed freely as a set of flat files and have been integrated into SRS. The PEP group of databases can be accessed from: http://cubic.bioc.columbia.edu/pep.

Animals↗

Bioinformatics and mass spectrometry for microorganism identification: proteome-wide post-translational modifications and database search algorithms for characterization of intact H. pylori.

MALDI-TOF mass spectrometry has been coupled with Internet-based proteome database search algorithms in an approach for direct microorganism identification. This approach is applied here to characterize intact H. pylori (strain 26695) Gram-negative bacteria, the most ubiquitous human pathogen. A procedure for including a specific and common posttranslational modification, N-terminal Met cleavage, in the search algorithm is described. Accounting for posttranslational modifications in putative protein biomarkers improves the identification reliability by at least an order of magnitude. The influence of other factors, such as number of detected biomarker peaks, proteome size, spectral calibration, and mass accuracy, on the microorganism identification success rate is illustrated as well.

Algorithms↗

Proteomics in Drosophila melanogaster: first 2D database of larval hemolymph proteins.

A proteomic approach was used for the identification of larval hemolymph proteins of Drosophila melanogaster. We report the initial establishment of a two-dimensional gel electrophoresis reference map for hemolymph proteins of third instar larvae of D. melanogaster. We used immobilized pH gradients of pH 4-7 (linear) and a 12-14% linear gradient polyacrylamide gel. The protein spots were silver-stained and analyzed by nanoLC-Q-Tof MS/MS (on-line nanoscale liquid chromatography quadrupole time of flight tandem mass spectrometry) or by Matrix assisted laser desorption time of flight MS (MALDI-TOF MS). Querying the SWISSPROT database with the mass spectrometric data yielded the identity of the proteins in the spots. The presented proteome map lists those protein spots identified to date. This map will be updated continuously and will serve as a reference database for investigators, studying changes at the protein level in different physiological conditions.

Animals↗

IndexToolkit: an open source toolbox to index protein databases for high-throughput proteomics.

UNLABELLED: A software package, IndexToolkit, aimed at overcoming the disadvantage of FASTA-format databases for frequent searching, is developed to utilize an indexing strategy to substantially accelerate sequence queries. IndexToolkit includes user-friendly tools and an Application Programming Interface (API) to facilitate indexing, storage and retrieval of protein sequence databases. As open source, it provides a sequence-retrieval developing framework, which is easily extensible for high-speed-request proteomic applications, such as database searching or modification discovering. We applied IndexToolkit to database searching engine pFind to demonstrate its effect. Experimental studies show that IndexToolkit is able to support significantly faster searches of protein database. AVAILABILITY: The IndexToolkit is free to use under the open source GNU GPL license. The source code and the compiled binary can be freely accessed through the website http://pfind.jdl.ac.cn/IndexToolkit. In this website, the more detailed information including screenshots and documentations for users and developers is also available.

Database Management Systems↗

The mouse SWISS-2D PAGE database: a tool for proteomics study of diabetes and obesity.

A number of two-dimensional electrophoresis (2-DE) reference maps from mouse samples have been established and could be accessed through the internet. An up-to-date list can be found in WORLD-2D PAGE (http://www.expasy.ch/ch2d/2d- index.html), an index of 2-DE databases and services. None of them were established from mouse white and brown adipose tissues, pancreatic islets, liver nuclei and skeletal muscle. This publication describes the mouse SWISS-2D PAGE database. Proteins present in samples of mouse (C57BI/6J) liver, liver nuclei, muscle, white and brown adipose tissue and pancreatic islets are assembled and described in an accessible uniform format. SWISS-2D PAGE can be accessed through the World Wide Web (WWW) network on the ExPASy molecular biology server (http://www.expasy.ch/ ch2d/).

Animals↗

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49 h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning↗