Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Integr8: enhanced inter-operability of European molecular biology databases.

OBJECTIVES: The increasing production of molecular biology data in the post-genomic era, and the proliferation of databases that store it, require the development of an integrative layer in database services to facilitate the synthesis of related information. The solution of this problem is made more difficult by the absence of universal identifiers for biological entities, and the breadth and variety of available data. METHODS: Integr8 was modelled using UML (Universal Modelling Language). Integr8 is being implemented as an n-tier system using a modern object-oriented programming language (Java). An object-relational mapping tool, OJB, is being used to specify the interface between the upper layers and an underlying relational database. RESULTS: The European Bioinformatics Institute is launching the Integr8 project. Integr8 will be an automatically populated database in which we will maintain stable identifiers for biological entities, describe their relationships with each other (in accordance with the central dogma of biology), and store equivalences between identified entities in the source databases. Only core data will be stored in Integr8, with web links to the source databases providing further information. CONCLUSIONS: Integr8 will provide the integrative layer of the next generation of bioinformatics services from the EBI. Web-based interfaces will be developed to offer gene-centric views of the integrated data, presenting (where known) the links between genome, proteome and phenotype.

Computational Biology↗

Identification of tumor-associated antigens using proteomics.

In the post-genomic era, the identification of tumor-associated antigens that elicit a humoral response is allowed at the protein level using proteomics. Indeed, the screening of autoantibodies using 2-D Western blot experiments with sera from cancer patients, followed by the subsequent identification of the target protein by mass spectrometry and database search has permitted the exploitation of the B-cell repertoire of patients with cancer. Applied to several types of cancer, a proteomic-based approach has revealed a high frequency of autoantibodies in sera from patients. Several of the antigenic proteins identified may constitute novel cancer markers and may have clinical utility in diagnosis or in establishing prognosis. Furthermore, the approach has allowed to distinguish isoforms that may help to define epitopes. On the other hand, the analysis of the expression levels of some of the antigenic proteins has revealed differential expression in tumors as compared with healthy tissues that might explain antigenicity.

Antigens, Neoplasm↗

Improvement of an in-gel tryptic digestion method for matrix-assisted laser desorption/ionization-time of flight mass spectrometry peptide mapping by use of volatile solubilizing agents.

The combination of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS), in-gel enzymatic digestion of proteins separated by two-dimensional gel electrophoresis and searches of molecular weight in peptide-mass databases is a powerful and well established method for protein identification in proteomics analysis. For successful protein identification by MALDI-TOF mass spectrometry of peptide mixtures, critical parameters include highly specific enzymatic cleavage, high mass accuracy and sufficient numbers and sequence coverage of the peptides which can be analyzed. For in-gel digestion with trypsin, the method employed should be compatible both with enzymatic cleavage and subsequent MALDI-TOF MS analysis. We report here an improved method for preparation of peptides for MALDI-TOF MS mass fingerprinting by using volatile solubilizing agents during the in-gel digestion procedure. Our study clearly demonstrates that modification of the in-gel digestion protocols by addition of dimethyl formamide (DMF) or a mixture of DMF/N,N-dimethyl acetamide at various concentrations can significantly increase the recovery of peptides. These higher yields of peptides resulted in more effective protein identification.

Amino Acid Sequence↗

PRIME: a graphical interface for integrating genomic/proteomic databases.

Data mining, finding and integration of information about proteins of interest, is an essential component in modern biological and biomedical research. Even when focusing on a single organism and only on a small number of proteins, there are often dozens fo data sources containing relevant information. We are developing PRIME, a protein information environment, to serve as a virtual central database which integrates distributed heterogeneous information about proteins (linked by common identifier). PRIME has powerful capabilities to visualize all kinds of protein annotation in specialized views. These views can be displayed side by side at the same time and can be synchronized in order to show simultaneously different aspects of identical proteins. These features allow a quick and comprehensive overview of properties of single proteins or protein sets.

Computational Biology↗

Gaining knowledge from previously unexplained spectra-application of the PTM-Explorer software to detect PTM in HUPO BPP MS/MS data.

A novel software tool named PTM-Explorer has been applied to LC-MS/MS datasets acquired within the Human Proteome Organisation (HUPO) Brain Proteome Project (BPP). PTM-Explorer enables automatic identification of peptide MS/MS spectra that were not explained in typical sequence database searches. The main focus was detection of PTMs, but PTM-Explorer detects also unspecific peptide cleavage, mass measurement errors, experimental modifications, amino acid substitutions, transpeptidation products and unknown mass shifts. To avoid a combinatorial problem the search is restricted to a set of selected protein sequences, which stem from previous protein identifications using a common sequence database search. Prior to application to the HUPO BPP data, PTM-Explorer was evaluated on excellently manually characterized and evaluated LC-MS/MS data sets from Alpha-A-Crystallin gel spots obtained from mouse eye lens. Besides various PTMs including phosphorylation, a wealth of experimental modifications and unspecific cleavage products were successfully detected, completing the primary structure information of the measured proteins. Our results indicate that a large amount of MS/MS spectra that currently remain unidentified in standard database searches contain valuable information that can only be elucidated using suitable software tools.

Amino Acid Sequence↗

Link test--A statistical method for finding prostate cancer biomarkers.

We present a new method, link-test, to select prostate cancer biomarkers from SELDI mass spectrometry and microarray data sets. Biomarkers selected by link-test are supported by data sets from both mRNA and protein levels, and therefore results in improved robustness. Link-test determines the level of significance of the association between a microarray marker and a specific mass spectrum marker by constructing background mass spectra distributions estimated by all human protein sequences in the SWISS-PROT database. The data set consist of both microarray and mass spectrometry data from prostate cancer patients and healthy controls. A list of statistically justified prostate cancer biomarkers is reported by link-test. Cross-validation results show high prediction accuracy using the identified biomarker panel. We also employ a text-mining approach with OMIM database to validate the cancer biomarkers. The study with link-test represents one of the first cross-platform studies of cancer biomarkers.

Algorithms↗

The transcriptional regulatory network of the amino acid producer Corynebacterium glutamicum.

The complete nucleotide sequence of the Corynebacterium glutamicum ATCC 13032 genome was previously determined and allowed the reliable prediction of 3002 protein-coding genes within this genome. Using computational methods, we have defined 158 genes, which form the minimal repertoire for proteins that presumably act as transcriptional regulators of gene expression. Most of these regulatory proteins have a direct role as DNA-binding transcriptional regulator, while others either have less well-defined functions in transcriptional regulation or even more general functions, such as the sigma factors. Recent advances in genome-wide transcriptional profiling of C. glutamicum generated a huge amount of data on regulation of gene expression. To understand transcriptional regulation of gene expression from the perspective of systems biology, rather than from the analysis of an individual regulatory protein, we compiled the current knowledge on the defined DNA-binding transcriptional regulators and their physiological role in modulating transcription in response to environmental signals. This comprehensive data collection provides a solid basis for database-guided reconstructions of the gene regulatory network of C. glutamicum, currently comprising 56 transcriptional regulators that exert 411 regulatory interactions to control gene expression. A graphical reconstruction revealed first insights into the functional modularity, the hierarchical architecture and the topological design principles of the transcriptional regulatory network of C. glutamicum.

Amino Acids↗

Bioinformatics: harvesting information for plant and crop science.

Bioinformatics is an integral aspect of plant and crop science research. Developments in data management and analytical software are reviewed with an emphasis on applications in functional genomics. This includes information resources for Arabidopsis and crop species, and tools available for analysis and visualisation of comparative genomic data. Approaches used to explore relationships between plant genes and expressed sequences are compared, including use of ontologies. The impact of bioinformatics in forward and reverse genetics is described, together with the potential from data mining. The role of bioinformatics is explored in the wider context of plant and crop science.

Algorithms↗

PIGOK: Linking protein identity to gene ontology and function.

Here we introduce a computer database that allows for the rapid retrieval of physicochemical properties, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes information about a protein or a list of proteins. We applied PIGOK analyzing Schizosaccharomyces pombe proteins displaying differential expression under oxidative stress and identified their biological functions and pathways. The database is available on the Internet at http://pc4-133.ludwig.ucl.ac.uk/pigok.html.

Computational Biology↗

Protein expression in a Drosophila model of Parkinson's disease.

Liquid chromatographies coupled to mass spectrometry and database analysis techniques are used to carry out a large-scale proteome characterization for a Drosophila model of Parkinson's disease. Semiquantitative analysis is performed on A30P alpha-synuclein expressing transgenic Drosophila and a control lacking the gene at presymptomatic, early, and advanced disease stages. Changes in gene expression at the level of the proteome are compared with changes reported from published transcriptome measurements. A summary of the comparison indicates that approximately 44% of transcripts that show changes can also be observed as proteins. However, the patterns of change in protein expression vary substantially compared with the patterns of change observed for corresponding transcripts. In addition, the expression changes of many genes are observed for only transcripts or proteins. Proteome measurements provide evidence for dysregulation of a group of proteins associated with the actin cytoskeleton and mitochondrion at presymptomatic and early disease stages that may presage the development of later symptoms. Overall, the proteome measurements provide a view of gene expression that is highly complementary to the insights obtained from the transcriptome.

Animals↗

The ABC's (and XYZ's) of peptide sequencing.

Proteomics is an increasingly powerful and indispensable technology in molecular cell biology. It can be used to identify the components of small protein complexes and large organelles, to determine post-translational modifications and in sophisticated functional screens. The key - but little understood - technology in mass-spectrometry-based proteomics is peptide sequencing, which we describe and review here in an easily accessible format.

Amino Acid Sequence↗

YPD, PombePD and WormPD: model organism volumes of the BioKnowledge library, an integrated resource for protein information.

The BioKnowledge Library is a relational database and web site (http://www.proteome.com) composed of protein-specific information collected from the scientific literature. Each Protein Report on the web site summarizes and displays published information about a single protein, including its biochemical function, role in the cell and in the whole organism, localization, mutant phenotype and genetic interactions, regulation, domains and motifs, interactions with other proteins and other relevant data. This report describes four species-specific volumes of the BioKnowledge Library, concerned with the model organisms Saccharomyces cerevisiae (YPD), Schizosaccharomyces pombe (PombePD) and Caenorhabditis elegans (WormPD), and with the fungal pathogen Candida albicans (CalPD). Protein Reports of each species are unified in format, easily searchable and extensively cross-referenced between species. The relevance of these comprehensively curated resources to analysis of proteins in other species is discussed, and is illustrated by a survey of model organism proteins that have similarity to human proteins involved in disease.

Animals↗

New molecular research technologies in the study of muscle disease.

PURPOSE OF REVIEW: To describe the current technologies and progress in DNA polymorphism association studies, mRNA expression profiling (microarrays), and proteomics with respect to muscle disease, and the increasing impact of public-access databases of genome-wide information. RECENT FINDINGS: mRNA expression profiling is becoming the most mature of the highly parallel molecular technologies, with microarrays now able to query the large majority of all genes using 1 million oligonucleotide probes built on 1.2-cm2 glass substrates. Applications of microarrays to normal muscle physiology and muscle disease are discussed. Single nucleotide polymorphism association studies promise to determine the predisposition of individuals to acquired muscle disease, including sarcopenia and atrophy, although such studies are in their infancy. Proteomics technologies do not enjoy the sensitivity and specificity of hybridization, and must instead rely on mass spectrometers. Mass spectrometry technology is advancing rapidly, although the sensitivity and throughput is far behind that of mRNA expression profiling. SUMMARY: As the gene mutations responsible for many types of muscular dystrophy and myopathy have been discovered, protein and gene testing has been integrated into the standard patient diagnostic workup. Future developments will include simpler and less expensive molecular diagnostics, advances in the understanding of downstream consequences of these defects, and the genetic predispositions underlying acquired muscle disease.

Female↗

The status of structural genomics defined through the analysis of current targets and structures.

Structural genomics--large-scale macromolecular 3-dimenional structure determination--is unique in that major participants report scientific progress on a weekly basis. The target database (TargetDB) maintained by the Protein Data Bank (http://targetdb.pdb.org) reports this progress through the status of each protein sequence (target) under consideration by the major structural genomics centers worldwide. Hence, TargetDB provides a unique opportunity to analyze the potential impact that this major initiative provides to scientists interested in the sequence-structure-function-disease paradigm. Here we report such an analysis with a focus on: (i) temporal characteristics--how is the project doing and what can we expect in the future? (ii) target characteristics--what are the predicted functions of the proteins targeted by structural genomics and how biased is the target set when compared to the PDB and to predictions across complete genomes? (iii) structures solved--what are the characteristics of structures solved thus far and what do they contribute? The analysis required a more extensive database of structure predictions using different methods integrated with data from other sources. This database, associated tools and related data sources are available from http://spam.sdsc.edu.

Computational Biology↗

A methodology to migrate the gene ontology to a description logic environment using DAML+OIL.

The Gene Ontology Next Generation Project (GONG) is developing a staged methodology to evolve the current representation of the Gene Ontology into DAML+OIL in order to take advantage of the richer formal expressiveness and the reasoning capabilities of the underlying description logic. Each stage provides a step level increase in formal explicit semantic content with a view to supporting validation, extension and multiple classification of the Gene Ontology. The paper introduces DAML+OIL and demonstrates the activity within each stage of the methodology and the functionality gained.

Classification↗