Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Comprehensive analyses of prostate gene expression: convergence of expressed sequence tag databases, transcript profiling and proteomics.

Several methods have been developed for the comprehensive analysis of gene expression in complex biological systems. Generally these procedures assess either a portion of the cellular transcriptome or a portion of the cellular proteome. Each approach has distinct conceptual and methodological advantages and disadvantages. We have investigated the application of both methods to characterize the gene expression pathway mediated by androgens and the androgen receptor in prostate cancer cells. This pathway is of critical importance for the development and progression of prostate cancer. Of clinical importance, modulation of androgens remains the mainstay of treatment for patients with advanced disease. To facilitate global gene expression studies we have first sought to define the prostate transcriptome by assembling and annotating prostate-derived expressed sequence tags (ESTs). A total of 55000 prostate ESTs were assembled into a set of 15953 clusters putatively representing 15953 distinct transcripts. These clusters were used to construct cDNA microarrays suitable for examining the androgen-response pathway at the level of transcription. The expression of 20 genes was found to be induced by androgens. This cohort included known androgen-regulated genes such as prostate-specific antigen (PSA) and several novel complementary DNAs (cDNAs). Protein expression profiles of androgen-stimulated prostate cancer cells were generated by two-dimensional electrophoresis (2-DE). Mass spectrometric analysis of androgen-regulated proteins in these cells identified the metastasis-suppressor gene NDKA/nm23, a finding that may explain a marked reduction in metastatic potential when these cells express a functional androgen receptor pathway.

DNA, Complementary↗

The Nuclear Protein Database (NPD): sub-nuclear localisation and functional annotation of the nuclear proteome.

The Nuclear Protein Database (NPD) is a curated database that contains information on more than 1300 vertebrate proteins that are thought, or are known, to localise to the cell nucleus. Each entry is annotated with information on predicted protein size and isoelectric point, as well as any repeats, motifs or domains within the protein sequence. In addition, information on the sub-nuclear localisation of each protein is provided and the biological and molecular functions are described using Gene Ontology (GO) terms. The database is searchable by keyword, protein name, sub-nuclear compartment and protein domain/motif. Links to other databases are provided (e.g. Entrez, SWISS-PROT, OMIM, PubMed, PubMed Central). Thus, NPD provides a gateway through which the nuclear proteome may be explored. The database can be accessed at http://npd.hgu.mrc.ac.uk and is updated monthly.

Amino Acid Sequence↗

Escherichia coli proteome analysis using the gene-protein database.

The gene-protein database of Escherichia coli is a collection of data, largely generated from the separation of complex mixtures of cellular proteins on two-dimensional (2-D) polyacrylamide gel electrophoresis. The database currently contains about 1600 protein spots. The data are comprised of both identification information for many of these proteins and data on how the level or synthesis rates of proteins vary under different growth conditions. Three projects are underway to further elucidate the E. coli proteome including a project to localize on 2-D gels all of the open reading framed encoded by the E. coli chromosome, a project to determine the condition(s) under which each open reading frame is expressed and a project to determine the abundance and location of each protein in the cell. Applications for proteome databases for cell modeling are discussed and examples of applications in therapeutic drug discovery are given.

Bacterial Proteins↗

The use of proteotypic peptide libraries for protein identification.

This paper describes an algorithm to apply proteotypic peptide sequence libraries to protein identifications performed using tandem mass spectrometry (MS/MS). Proteotypic peptides are those peptides in a protein sequence that are most likely to be confidently observed by current MS-based proteomics methods. Libraries of proteotypic peptide sequences were compiled from the Global Proteome Machine Database for Homo sapiens and Saccharomyces cerevisiae model species proteomes. These libraries were used to scan through collections of tandem mass spectra to discover which proteins were represented by the data sets, followed by detailed analysis of the spectra with the full protein sequences corresponding to the discovered proteotypic peptides. This algorithm (Proteotypic Peptide Profiling, or P3) resulted in sequence-to-spectrum matches comparable to those obtained by conventional protein identification algorithms using only full protein sequences, with a 20-fold reduction in the time required to perform the identification calculations. The proteotypic peptide libraries, the open source code for the implementation of the search algorithm and a website for using the software have been made freely available. Approximately 4% of the residues in the H. sapiens proteome were required in the proteotypic peptide library to successfully identify proteins.

Algorithms↗

[Strategy for the protein identification of human proteome expression profile: selection of searching database].

Widely used method of protein identification for high-throughout proteome expression profile studies was database-dependent, so the selection of databases for the protein identification was very important. Despite the deficiency of available human protein databases, the complementarity of human proteins could be got mainly from human genome but not from the protein databases of other organisms. According to the comparison of the current protein databases from different aspects, IPI was recommended for the basic identification for the studies of human proteome expression profile, and other human protein or nucleic acid databases were needed for the complementary identification and novel protein mining.

Animals↗

Two-dimensional electrophoresis database of Listeria monocytogenes EGDe proteome and proteomic analysis of mid-log and stationary growth phase cells.

Listeria monocytogenes is the causative agent of listeriosis, one of the most significant foodborne diseases in industrialized countries. The complete genome of the L. monocytogenes EGDe strain, belonging to the serogroup 1/2a, has been sequenced and is comprised of 2853 open reading frames. The objective of the current study was to construct a two-dimensional (2-D) database of the proteome of this strain. The soluble protein fractions of the microorganism were recovered either in the mid-log or in the stationary phase of growth at 37 degrees C. These fractions were analyzed by 2-D electrophoresis (2-DE), using immobilized pH gradient strips of various pH values (3-10, 3-6, and 5-8) for the first-dimensional separations and 12.5% acrylamide gels for sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). 201 protein spots corresponding to 126 different proteins were identified by matrix assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS). The 2-DE maps presented here provide a first basis for further investigations of protein expression in L. monocytogenes. In this way, the comparison of proteome between cells in the exponential or stationary phase of growth at 37 degrees C allowed us to characterize 161 variations in protein spot intensity, of which 38 were identified. Among the differentially expressed proteins were ribosomal proteins (RpsF, RplJ, and RpmE), proteins involved in cellular metabolism (GlpD, PdhD, Pgm, Lmo1372, Lmo2696, and Lmo2743) or in stress adaptation (GroES and ferritin), a fructose-specific phosphotransferase enzyme IIB (Lmo0399) and different post-translational modified forms of listeriolysin (LLO).

Bacterial Proteins↗

A bioinformatics perspective on proteomics: data storage, analysis, and integration.

The field of proteomics is advancing rapidly as a result of powerful new technologies and proteomics experiments yield a vast and increasing amount of information. Data regarding protein occurrence, abundance, identity, sequence, structure, properties, and interactions need to be stored. Currently, a common standard has not yet been established and open access to results is needed for further development of robust analysis algorithms. Databases for proteomics will evolve from pure storage into knowledge resources, providing a repository for information (meta-data) which is mainly not stored in simple flat files. This review will shed light on recent steps towards the generation of a common standard in proteomics data storage and integration, but is not meant to be a comprehensive overview of all available databases and tools in the proteomics community.

Computational Biology↗

MS2Grouper: group assessment and synthetic replacement of duplicate proteomic tandem mass spectra.

Shotgun proteomics experiments require the collection of thousands of tandem mass spectra; these sets of data will continue to grow as new instruments become available that can scan at even higher rates. Such data contain substantial amounts of redundancy with spectra from a particular peptide being acquired many times during a single LC-MS/MS experiment. In this article, we present MS2Grouper, an algorithm that detects spectral duplication, assesses groups of related spectra, and replaces these groups with synthetic representative spectra. Errors in detecting spectral similarity are corrected using a paraclique criterion-spectra are only assessed as groups if they are part of a clique of at least three completely interrelated spectra or are subsequently added to such cliques by being similar to all but one of the clique members. A greedy algorithm constructs a representative spectrum for each group by iteratively removing the tallest peaks from the spectral collection and matching to peaks in the other spectra. This strategy is shown to be effective in reducing spectral counts by up to 20% in LC-MS/MS datasets from protein standard mixtures and proteomes, reducing database search times without a concomitant reduction in identified peptides.

Algorithms↗

Predicting co-complexed protein pairs using genomic and proteomic data integration.

BACKGROUND: Identifying all protein-protein interactions in an organism is a major objective of proteomics. A related goal is to know which protein pairs are present in the same protein complex. High-throughput methods such as yeast two-hybrid (Y2H) and affinity purification coupled with mass spectrometry (APMS) have been used to detect interacting proteins on a genomic scale. However, both Y2H and APMS methods have substantial false-positive rates. Aside from high-throughput interaction screens, other gene- or protein-pair characteristics may also be informative of physical interaction. Therefore it is desirable to integrate multiple datasets and utilize their different predictive value for more accurate prediction of co-complexed relationship. RESULTS: Using a supervised machine learning approach--probabilistic decision tree, we integrated high-throughput protein interaction datasets and other gene- and protein-pair characteristics to predict co-complexed pairs (CCP) of proteins. Our predictions proved more sensitive and specific than predictions based on Y2H or APMS methods alone or in combination. Among the top predictions not annotated as CCPs in our reference set (obtained from the MIPS complex catalogue), a significant fraction was found to physically interact according to a separate database (YPD, Yeast Proteome Database), and the remaining predictions may potentially represent unknown CCPs. CONCLUSIONS: We demonstrated that the probabilistic decision tree approach can be successfully used to predict co-complexed protein (CCP) pairs from other characteristics. Our top-scoring CCP predictions provide testable hypotheses for experimental validation.

Computational Biology↗

DBToolkit: processing protein databases for peptide-centric proteomics.

UNLABELLED: DBToolkit is a user-friendly, easily extensible tool that allows the processing of protein sequence databases to peptide-centric sequence databases. This processing is primarily aimed at enhancing the useful information content of these databases for use as optimized search spaces for efficient identification of peptide fragmentation spectra obtained by mass spectrometry. In addition, DBToolkit can be used to reliably solve a range of other typical tasks in processing sequence databases. AVAILABILITY: DBToolkit is open source under the GNU GPL license. The source code, full user and developer documentation and cross-platform binaries are freely downloadable from the project website at http://genesis.UGent.be/dbtoolkit/ CONTACT: lennart.martens@UGent.be

Database Management Systems↗

Anopheles gambiae genome reannotation through synthesis of ab initio and comparative gene prediction algorithms.

BACKGROUND: Complete genome annotation is a necessary tool as Anopheles gambiae researchers probe the biology of this potent malaria vector. RESULTS: We reannotate the A. gambiae genome by synthesizing comparative and ab initio sets of predicted coding sequences (CDSs) into a single set using an exon-gene-union algorithm followed by an open-reading-frame-selection algorithm. The reannotation predicts 20,970 CDSs supported by at least two lines of evidence, and it lowers the proportion of CDSs lacking start and/or stop codons to only approximately 4%. The reannotated CDS set includes a set of 4,681 novel CDSs not represented in the Ensembl annotation but with EST support, and another set of 4,031 Ensembl-supported genes that undergo major structural and, therefore, probably functional changes in the reannotated set. The quality and accuracy of the reannotation was assessed by comparison with end sequences from 20,249 full-length cDNA clones, and evaluation of mass spectrometry peptide hit rates from an A. gambiae shotgun proteomic dataset confirms that the reannotated CDSs offer a high quality protein database for proteomics. We provide a functional proteomics annotation, ReAnoXcel, obtained by analysis of the new CDSs through the AnoXcel pipeline, which allows functional comparisons of the CDS sets within the same bioinformatic platform. CDS data are available for download. CONCLUSION: Comprehensive A. gambiae genome reannotation is achieved through a combination of comparative and ab initio gene prediction algorithms.

Algorithms↗

Applications of InterPro in protein annotation and genome analysis.

The applications of InterPro span a range of biologically important areas that includes automatic annotation of protein sequences and genome analysis. In automatic annotation of protein sequences InterPro has been utilised to provide reliable characterisation of sequences, identifying them as candidates for functional annotation. Rules based on the InterPro characterisation are stored and operated through a database called RuleBase. RuleBase is used as the main tool in the sequence database group at the EBI to apply automatic annotation to unknown sequences. The annotated sequences are stored and distributed in the TrEMBL protein sequence database. InterPro also provides a means to carry out statistical and comparative analyses of whole genomes. In the Proteome Analysis Database, InterPro analyses have been combined with other analyses based on CluSTr, the Gene Ontology (GO) and structural information on the proteins.

Amino Acid Sequence↗

Implications of structural genomics target selection strategies: Pfam5000, whole genome, and random approaches.

Structural genomics is an international effort to determine the three-dimensional shapes of all important biological macromolecules, with a primary focus on proteins. Target proteins should be selected according to a strategy that is medically and biologically relevant, of good value, and tractable. As an option to consider, we present the "Pfam5000" strategy, which involves selecting the 5000 most important families from the Pfam database as sources for targets. We compare the Pfam5000 strategy to several other proposed strategies that would require similar numbers of targets. These strategies include complete solution of several small to moderately sized bacterial proteomes, partial coverage of the human proteome, and random selection of approximately 5000 targets from sequenced genomes. We measure the impact that successful implementation of these strategies would have upon structural interpretation of the proteins in Swiss-Prot, TrEMBL, and 131 complete proteomes (including 10 of eukaryotes) from the Proteome Analysis database at the European Bioinformatics Institute (EBI). Solving the structures of proteins from the 5000 largest Pfam families would allow accurate fold assignment for approximately 68% of all prokaryotic proteins (covering 59% of residues) and 61% of eukaryotic proteins (40% of residues). More fine-grained coverage that would allow accurate modeling of these proteins would require an order of magnitude more targets. The Pfam5000 strategy may be modified in several ways, for example, to focus on larger families, bacterial sequences, or eukaryotic sequences; as long as secondary consideration is given to large families within Pfam, coverage results vary only slightly. In contrast, focusing structural genomics on a single tractable genome would have only a limited impact in structural knowledge of other proteomes: A significant fraction (about 30-40% of the proteins and 40-60% of the residues) of each proteome is classified in small families, which may have little overlap with other species of interest. Random selection of targets from one or more genomes is similar to the Pfam5000 strategy in that proteins from larger families are more likely to be chosen, but substantial effort would be spent on small families.

Animals↗

2-DE proteomic profiling of neuronal stem cells.

Proteomics has become a powerful tool in neuroscience studies. Although numerous human neural stem cells are available for research purposes since many years, there exists only limited information on proteomic data from stable neural stem cell lines. Profiling and functional proteome studies of neuronal stem cells will help to describe the protein inventory as well as protein activity and interactions, subcellular localization and posttranslational modifications. The proteomic analysis of neuronal differentiation processes will elucidate the complex events leading to the generation of different phenotypes via distinctive developmental programs that control self-renewal, differentiation, and plasticity. Using the ReNcell VM197 model, a cell line derived from human fetal ventral mesencephalon stem cells, we studied the protein inventory of the stem cells by 2-DE gel electrophoresis and mass spectrometric protein identification and constructed a 2-DE protein map consisting of more than 400 identified protein spots. This proteome reference database constitutes the basis for further investigations of differential protein expression during differentiation. A profiling of the neuronal differentiation-associated changes displayed the large rearrangement of the proteome during this process, and the proteomic techniques proved to be a valuable tool for the elucidation of neuronal differentiation process and for target protein screening.

Animals↗

PHOG: a database of supergenomes built from proteome complements.

BACKGROUND: Orthologs and paralogs are widely used terms in modern comparative genomics. Existing procedures for resolving orthologous/paralogous relationships are often based on manual revision of clusters of orthologous groups and/or lack any rigorous evolutionary base. DESCRIPTION: We developed a completely automated procedure that creates clusters of orthologous groups at each node of the taxonomy tree (PHOGs--Phylogenetic Orthologous Groups). As a result of this procedure, a tree of orthologous groups was obtained. Each cluster is a "supergene" and it is represented by an "ancestral" sequence obtained from the multiple alignment of orthologous and paralogous genes. The procedure has been applied to the taxonomy tree of organisms from all three domains of life. Protein complements from 50 bacterial, archaeal and eukaryotic species were used to create PHOGs at all tree nodes. 51367 PHOGs were obtained at the root node. CONCLUSION: The PHOG database demonstrates that it is possible to automatically process any number of sequenced genomes and to reconstruct orthologous and paralogous relationships between genomes using a rigorous evolutionary approach. This database can become a very useful tool in various areas of comparative genomics.

Databases, Genetic↗

Two-dimensional maps and databases of the human macrophage proteome and secretome.

Macrophages exert a crucial, but still incompletely known, role in complex disorders such inflammatory, immunological, and infectious diseases. A differential proteomic approach should help to elucidate the macrophage dysfunctions involved in these diseases. With this goal in mind, we established the first two-dimensional maps of the human macrophage proteome and secretome. Intracellular and secreted proteins were extracted from monocyte-derived macrophages obtained from healthy donors (n = 16), and separated by two-dimensional gel electrophoresis. Silver-stained gels were analyzed using Progenesis software. A high level of between-gel reproducibility was obtained, allowing us to generate two patterns specific of the macrophage proteome and secretome, respectively. A total of 127 and 66 distinct intracellular and secreted polypeptide spots, corresponding to 100 and 38 different proteins, respectively, were identified by matrix assisted laser desorption/ionisation-mass spectrometry. The two-dimensional reference maps and databases resulting from this study confirm that macrophages are involved in a wide range of biological functions, and that they provide a useful tool for a wide array of investigators involved in macrophage biology, allowing to investigate the macrophage protein changes associated with various disorders or environmental stimuli.

Databases, Protein↗

Functional discrimination of gene expression patterns in terms of the gene ontology.

The ever-growing amount of experimental data in molecular biology and genetics requires its automated analysis, by employing sophisticated knowledge discovery tools. We use an Inductive Logic Programming (ILP) learner to induce functional discrimination rules between genes studied using microarrays and found to be differentially expressed in three recently discovered subtypes of adenocarcinoma of the lung. The discrimination rules involve functional annotations from the Proteome HumanPSD database in terms of the Gene Ontology, whose hierarchical structure is essential for this task. While most of the lower levels of gene expression data (pre)processing have been automated, our work can be seen as a step toward automating the higher level functional analysis of the data. We view our application not just as a prototypical example of applying more sophisticated machine learning techniques to the functional analysis of genes, but also as an incentive for developing increasingly more sophisticated functional annotations and ontologies, that can be automatically processed by such learning algorithms.

Adenocarcinoma↗

The Dictyostelium discoideum proteome--the SWISS-2DPAGE database of the multicellular aggregate (slug).

The cellular slime mold Dictyostelium discoideum is a eukaryotic microorganism which has developmental life stages attractive to the cell and molecular biologist. By displaying the two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) protein map of different developmental stages, the key molecules can be identified and characterised, allowing a detailed understanding of the D. discoideum proteome. Here we describe the preparation of reference gel of the D. discoideum multicellular aggregate, the slug. Proteins were separated by 2-D PAGE with immobilised pH gradients (pH 3.5-10) in the first dimension and sodium dodecyl sulfate (SDS)-PAGE in the second dimension. Micropreparative gels were electroblotted onto polyvinylidene difluoride (PVDF) membranes and 150 spots were visualised by amido black staining. Protein spots were excised and 31 were putatively identified by matching their amino acid composition, estimated isoelectric point (pI) and molecular weight (M(r)) against the SWISS-PROT database with the ExPASy AAcompID tool (http:// expasy.hcuge.ch/ch2d/aacompi.html). A total of 25 proteins were identified by matching against database entries for D. discoideum, and another six by cross-species matching against database entries for Saccharomyces cerevisiae proteins. This map will be available in the SWISS-2DPAGE database.

Animals↗