Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

The BPP (protein biochemistry and proteomics) two-dimensional electrophoresis database.

The BPP (protein biochemistry and proteomics) two-dimensional electrophoresis (2-DE) database (http://www-smbh.univ-paris13.fr/lbtp/Biochemistry/Biochimie/bque.htm) was established in 1998. The current release contains 11 reference maps from human hematopoietic and lymphoid cell line samples. These reference maps have now 255 identified spots, corresponding to 84 protein entries. The World Wide Web (WWW) presentation is designed to allow public access to the available 2-DE data together with logical connections to databases providing complementary information.

Databases, Factual↗

SPINE 2: a system for collaborative structural proteomics within a federated database framework.

We present version 2 of the SPINE system for structural proteomics. SPINE is available over the web at http://nesg.org. It serves as the central hub for the Northeast Structural Genomics Consortium, allowing collaborative structural proteomics to be carried out in a distributed fashion. The core of SPINE is a laboratory information management system (LIMS) for key bits of information related to the progress of the consortium in cloning, expressing and purifying proteins and then solving their structures by NMR or X-ray crystallography. Originally, SPINE focused on tracking constructs, but, in its current form, it is able to track target sample tubes and store detailed sample histories. The core database comprises a set of standard relational tables and a data dictionary that form an initial ontology for proteomic properties and provide a framework for large-scale data mining. Moreover, SPINE sits at the center of a federation of interoperable information resources. These can be divided into (i) local resources closely coupled with SPINE that enable it to handle less standardized information (e.g. integrated mailing and publication lists), (ii) other information resources in the NESG consortium that are inter-linked with SPINE (e.g. crystallization LIMS local to particular laboratories) and (iii) international archival resources that SPINE links to and passes on information to (e.g. TargetDB at the PDB).

Cooperative Behavior↗

ProteoformDB: A Built-In Application to Generate Proteoform Database.

Proteins play essential functions through their complex regulations on cell-type-specific expression, localization, and molecular complexes. Protein complexity is further enhanced by proteoforms, which are the diverse molecular forms that each gene can produce through genomic alterations, transcriptional variations, translational regulations, and protein modifications. Profiling of proteoforms is a promising method for gaining a deeper understanding of the role of proteins in biological pathways and disease mechanisms. Here, we developed ProteoformDB, an application tool for generating proteoform databases, and we cataloged a total of over one million unique single-site human proteoforms. We showed that ProteoformDB can serve as a valuable resource to document the experimentally identified proteoforms in a database, supporting protein characterization in quantitative proteomics for both total protein abundances and modified protein forms.

Humans↗

Interactive InterPro-based comparisons of proteins in whole genomes.

MOTIVATION: The SWISS-PROT group at the EBI has developed the Proteome Analysis Database utilizing existing resources and providing comprehensive and integrated comparative analysis of the predicted protein coding sequences of the complete genomes of bacteria, archaea and eukaryotes. The Proteome Analysis Database is accompanied by a program that has been designed to carry out interactive InterPro proteome comparisons for any one proteome against any other one or more of the proteomes in the database.

Computational Biology↗

Two-dimensional map of the proteome of Haemophilus influenzae.

We have constructed a two-dimensional database of the proteome of Haemophilus influenzae, a bacterium of medical interest of which the complete genome, comprising about 1742 open reading frames, has been sequenced. The soluble protein fraction of the microorganism was analyzed by two-dimensional electrophoresis, using immobilized pH gradient strips of various pH regions, gels with different acrylamide concentrations and buffers with different trailing ions. In order to visualize low-copy-number gene products, we employed a series of protein extraction and sample application approaches and several chromatographic steps, including heparin chromatography, chromatofocusing and hydrophobic interaction chromatography. We have also analyzed the cell envelope-bound protein fraction using either immobilized pH gradient strips or a two-detergent system with a cationic detergent in the first and an anionic detergent in the second-dimensional separation. Different proteins (502) were identified by matrix-assisted laser desorption/ionization mass spectrometry and amino acid composition analysis. This is at present one of the largest two-dimensional proteome databases.

Databases, Factual↗

PEP: Predictions for Entire Proteomes.

PEP is a database of Predictions for Entire Proteomes. The database contains summaries of analyses of protein sequences from a range of organisms representing all three major kingdoms of life: eukaryotes, prokaryotes and archaea. All proteins publicly available for organisms were aligned against SWISS-PROT, TrEMBL and PDB. Additionally, the following annotations are provided: secondary structure, transmembrane helices, coiled coils, regions of low complexity, signal peptides, PROSITE motifs, nuclear localization signals and classes of cellular function. Proteins that contain long regions without regular secondary structure are also identified. We have produced a related database of structural domain-like fragments derived from PEP and clusters based on homology between all fragments. The PEP database, fragments and clusters are distributed freely as a set of flat files and have been integrated into SRS. The PEP group of databases can be accessed from: http://cubic.bioc.columbia.edu/pep.

Animals↗

Bioinformatics and mass spectrometry for microorganism identification: proteome-wide post-translational modifications and database search algorithms for characterization of intact H. pylori.

MALDI-TOF mass spectrometry has been coupled with Internet-based proteome database search algorithms in an approach for direct microorganism identification. This approach is applied here to characterize intact H. pylori (strain 26695) Gram-negative bacteria, the most ubiquitous human pathogen. A procedure for including a specific and common posttranslational modification, N-terminal Met cleavage, in the search algorithm is described. Accounting for posttranslational modifications in putative protein biomarkers improves the identification reliability by at least an order of magnitude. The influence of other factors, such as number of detected biomarker peaks, proteome size, spectral calibration, and mass accuracy, on the microorganism identification success rate is illustrated as well.

Algorithms↗

Proteomics in Drosophila melanogaster: first 2D database of larval hemolymph proteins.

A proteomic approach was used for the identification of larval hemolymph proteins of Drosophila melanogaster. We report the initial establishment of a two-dimensional gel electrophoresis reference map for hemolymph proteins of third instar larvae of D. melanogaster. We used immobilized pH gradients of pH 4-7 (linear) and a 12-14% linear gradient polyacrylamide gel. The protein spots were silver-stained and analyzed by nanoLC-Q-Tof MS/MS (on-line nanoscale liquid chromatography quadrupole time of flight tandem mass spectrometry) or by Matrix assisted laser desorption time of flight MS (MALDI-TOF MS). Querying the SWISSPROT database with the mass spectrometric data yielded the identity of the proteins in the spots. The presented proteome map lists those protein spots identified to date. This map will be updated continuously and will serve as a reference database for investigators, studying changes at the protein level in different physiological conditions.

Animals↗

The mouse SWISS-2D PAGE database: a tool for proteomics study of diabetes and obesity.

A number of two-dimensional electrophoresis (2-DE) reference maps from mouse samples have been established and could be accessed through the internet. An up-to-date list can be found in WORLD-2D PAGE (http://www.expasy.ch/ch2d/2d- index.html), an index of 2-DE databases and services. None of them were established from mouse white and brown adipose tissues, pancreatic islets, liver nuclei and skeletal muscle. This publication describes the mouse SWISS-2D PAGE database. Proteins present in samples of mouse (C57BI/6J) liver, liver nuclei, muscle, white and brown adipose tissue and pancreatic islets are assembled and described in an accessible uniform format. SWISS-2D PAGE can be accessed through the World Wide Web (WWW) network on the ExPASy molecular biology server (http://www.expasy.ch/ ch2d/).

Animals↗

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49 h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning↗

Comprehensive analyses of prostate gene expression: convergence of expressed sequence tag databases, transcript profiling and proteomics.

Several methods have been developed for the comprehensive analysis of gene expression in complex biological systems. Generally these procedures assess either a portion of the cellular transcriptome or a portion of the cellular proteome. Each approach has distinct conceptual and methodological advantages and disadvantages. We have investigated the application of both methods to characterize the gene expression pathway mediated by androgens and the androgen receptor in prostate cancer cells. This pathway is of critical importance for the development and progression of prostate cancer. Of clinical importance, modulation of androgens remains the mainstay of treatment for patients with advanced disease. To facilitate global gene expression studies we have first sought to define the prostate transcriptome by assembling and annotating prostate-derived expressed sequence tags (ESTs). A total of 55000 prostate ESTs were assembled into a set of 15953 clusters putatively representing 15953 distinct transcripts. These clusters were used to construct cDNA microarrays suitable for examining the androgen-response pathway at the level of transcription. The expression of 20 genes was found to be induced by androgens. This cohort included known androgen-regulated genes such as prostate-specific antigen (PSA) and several novel complementary DNAs (cDNAs). Protein expression profiles of androgen-stimulated prostate cancer cells were generated by two-dimensional electrophoresis (2-DE). Mass spectrometric analysis of androgen-regulated proteins in these cells identified the metastasis-suppressor gene NDKA/nm23, a finding that may explain a marked reduction in metastatic potential when these cells express a functional androgen receptor pathway.

DNA, Complementary↗

The Nuclear Protein Database (NPD): sub-nuclear localisation and functional annotation of the nuclear proteome.

The Nuclear Protein Database (NPD) is a curated database that contains information on more than 1300 vertebrate proteins that are thought, or are known, to localise to the cell nucleus. Each entry is annotated with information on predicted protein size and isoelectric point, as well as any repeats, motifs or domains within the protein sequence. In addition, information on the sub-nuclear localisation of each protein is provided and the biological and molecular functions are described using Gene Ontology (GO) terms. The database is searchable by keyword, protein name, sub-nuclear compartment and protein domain/motif. Links to other databases are provided (e.g. Entrez, SWISS-PROT, OMIM, PubMed, PubMed Central). Thus, NPD provides a gateway through which the nuclear proteome may be explored. The database can be accessed at http://npd.hgu.mrc.ac.uk and is updated monthly.

Amino Acid Sequence↗

Escherichia coli proteome analysis using the gene-protein database.

The gene-protein database of Escherichia coli is a collection of data, largely generated from the separation of complex mixtures of cellular proteins on two-dimensional (2-D) polyacrylamide gel electrophoresis. The database currently contains about 1600 protein spots. The data are comprised of both identification information for many of these proteins and data on how the level or synthesis rates of proteins vary under different growth conditions. Three projects are underway to further elucidate the E. coli proteome including a project to localize on 2-D gels all of the open reading framed encoded by the E. coli chromosome, a project to determine the condition(s) under which each open reading frame is expressed and a project to determine the abundance and location of each protein in the cell. Applications for proteome databases for cell modeling are discussed and examples of applications in therapeutic drug discovery are given.

Bacterial Proteins↗

Applications of InterPro in protein annotation and genome analysis.

The applications of InterPro span a range of biologically important areas that includes automatic annotation of protein sequences and genome analysis. In automatic annotation of protein sequences InterPro has been utilised to provide reliable characterisation of sequences, identifying them as candidates for functional annotation. Rules based on the InterPro characterisation are stored and operated through a database called RuleBase. RuleBase is used as the main tool in the sequence database group at the EBI to apply automatic annotation to unknown sequences. The annotated sequences are stored and distributed in the TrEMBL protein sequence database. InterPro also provides a means to carry out statistical and comparative analyses of whole genomes. In the Proteome Analysis Database, InterPro analyses have been combined with other analyses based on CluSTr, the Gene Ontology (GO) and structural information on the proteins.

Amino Acid Sequence↗

Functional discrimination of gene expression patterns in terms of the gene ontology.

The ever-growing amount of experimental data in molecular biology and genetics requires its automated analysis, by employing sophisticated knowledge discovery tools. We use an Inductive Logic Programming (ILP) learner to induce functional discrimination rules between genes studied using microarrays and found to be differentially expressed in three recently discovered subtypes of adenocarcinoma of the lung. The discrimination rules involve functional annotations from the Proteome HumanPSD database in terms of the Gene Ontology, whose hierarchical structure is essential for this task. While most of the lower levels of gene expression data (pre)processing have been automated, our work can be seen as a step toward automating the higher level functional analysis of the data. We view our application not just as a prototypical example of applying more sophisticated machine learning techniques to the functional analysis of genes, but also as an incentive for developing increasingly more sophisticated functional annotations and ontologies, that can be automatically processed by such learning algorithms.

Adenocarcinoma↗

The Dictyostelium discoideum proteome--the SWISS-2DPAGE database of the multicellular aggregate (slug).

The cellular slime mold Dictyostelium discoideum is a eukaryotic microorganism which has developmental life stages attractive to the cell and molecular biologist. By displaying the two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) protein map of different developmental stages, the key molecules can be identified and characterised, allowing a detailed understanding of the D. discoideum proteome. Here we describe the preparation of reference gel of the D. discoideum multicellular aggregate, the slug. Proteins were separated by 2-D PAGE with immobilised pH gradients (pH 3.5-10) in the first dimension and sodium dodecyl sulfate (SDS)-PAGE in the second dimension. Micropreparative gels were electroblotted onto polyvinylidene difluoride (PVDF) membranes and 150 spots were visualised by amido black staining. Protein spots were excised and 31 were putatively identified by matching their amino acid composition, estimated isoelectric point (pI) and molecular weight (M(r)) against the SWISS-PROT database with the ExPASy AAcompID tool (http:// expasy.hcuge.ch/ch2d/aacompi.html). A total of 25 proteins were identified by matching against database entries for D. discoideum, and another six by cross-species matching against database entries for Saccharomyces cerevisiae proteins. This map will be available in the SWISS-2DPAGE database.

Animals↗

Integrated genomic and proteomic analyses of a systematically perturbed metabolic network.

We demonstrate an integrated approach to build, test, and refine a model of a cellular pathway, in which perturbations to critical pathway components are analyzed using DNA microarrays, quantitative proteomics, and databases of known physical interactions. Using this approach, we identify 997 messenger RNAs responding to 20 systematic perturbations of the yeast galactose-utilization pathway, provide evidence that approximately 15 of 289 detected proteins are regulated posttranscriptionally, and identify explicit physical interactions governing the cellular response to each perturbation. We refine the model through further iterations of perturbation and global measurements, suggesting hypotheses about the regulation of galactose utilization and physical interactions between this and a variety of other metabolic pathways.

Computational Biology↗