Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

iProClass: an integrated database of protein family, function and structure information.

The iProClass database provides comprehensive, value-added descriptions of proteins and serves as a framework for data integration in a distributed networking environment. The protein information in iProClass includes family relationships as well as structural and functional classifications and features. The current version consists of about 830 000 non-redundant PIR-PSD, SWISS-PROT, and TrEMBL proteins organized with more than 36 000 PIR superfamilies, 145 000 families, 4000 domains, 1300 motifs and 550 000 FASTA similarity clusters. It provides rich links to over 50 database of protein sequences, families, functions and pathways, protein-protein interactions, post-translational modifications, protein expressions, structures and structural classifications, genes and genomes, ontologies, literature and taxonomy. Protein and superfamily summary reports present extensive annotation information and include membership statistics and graphical display of domains and motifs. iProClass employs an open and modular architecture for interoperability and scalability. It is implemented in the Oracle object-relational database system and is updated biweekly. The database is freely accessible from the web site at http://pir.georgetown.edu/iproclass/ and searchable by sequence or text string. The data integration in iProClass supports exploration of protein relationships. Such knowledge is fundamental to the understanding of protein evolution, structure and function and crucial to functional genomic and proteomic research.

Amino Acid Motifs↗

PHYTOPROT: a database of clusters of plant proteins.

All the protein sequences from plants (including Arabidopsis thaliana) available from SwissProt/TrEMBL have been the subject of an all-by-all systematic comparison and grouped into clusters of related proteins. Within each cluster, the sequences have been submitted to pyramidal classification; in the case where two or several subfamilies have been grouped together, the pyramidal tree helps in finding which sequences make the links between subfamilies. In addition, the 'domains' that are common to two or more sequences within a cluster were determined and displayed à la ProDom. The resulting graphical representations proved to be quite efficient in pinpointing those protein sequences suffering from a probable error in the annotation of their genes. The clusters can be searched through various criteria and their pyramidal classifications and their domain representations can be displayed by querying http://genoplante-info. infobiogen.fr/phytoprot. The user can also launch a BLAST search of a query sequence against all the clusters.

Arabidopsis Proteins↗

ApiEST-DB: analyzing clustered EST data of the apicomplexan parasites.

ApiEST-DB (http://www.cbil.upenn.edu/paradbs-servlet/) provides integrated access to publicly available EST data from protozoan parasites in the phylum Apicomplexa. The database currently incorporates a total of nearly 100,000 ESTs from several parasite species of clinical and/or veterinary interest, including Eimeria tenella, Neospora caninum, Plasmodium falciparum, Sarcocystis neurona and Toxoplasma gondii. To facilitate analysis of these data, EST sequences were clustered and assembled to form consensus sequences for each organism, and these assemblies were then subjected to automated annotation via similarity searches against protein and domain databases. The underlying relational database infrastructure, Genomics Unified Schema (GUS), enables complex biologically based queries, facilitating validation of gene models, identification of alternative splicing, detection of single nucleotide polymorphisms, identification of stage-specific genes and recognition of phylogenetically conserved and phylogenetically restricted sequences.

Animals↗

WebGestalt: an integrated system for exploring gene sets in various biological contexts.

High-throughput technologies have led to the rapid generation of large-scale datasets about genes and gene products. These technologies have also shifted our research focus from 'single genes' to 'gene sets'. We have developed a web-based integrated data mining system, WebGestalt (http://genereg.ornl.gov/webgestalt/), to help biologists in exploring large sets of genes. WebGestalt is composed of four modules: gene set management, information retrieval, organization/visualization, and statistics. The management module uploads, saves, retrieves and deletes gene sets, as well as performs Boolean operations to generate the unions, intersections or differences between different gene sets. The information retrieval module currently retrieves information for up to 20 attributes for all genes in a gene set. The organization/visualization module organizes and visualizes gene sets in various biological contexts, including Gene Ontology, tissue expression pattern, chromosome distribution, metabolic and signaling pathways, protein domain information and publications. The statistics module recommends and performs statistical tests to suggest biological areas that are important to a gene set and warrant further investigation. In order to demonstrate the use of WebGestalt, we have generated 48 gene sets with genes over-represented in various human tissue types. Exploration of all the 48 gene sets using WebGestalt is available for the public at http://genereg.ornl.gov/webgestalt/wg_enrich.php.

Computer Graphics↗

MAO: a Multiple Alignment Ontology for nucleic acid and protein sequences.

The application of high-throughput techniques such as genomics, proteomics or transcriptomics means that vast amounts of heterogeneous data are now available in the public databases. Bioinformatics is responding to the challenge with new integrated management systems for data collection, validation and analysis. Multiple alignments of genomic and protein sequences provide an ideal environment for the integration of this mass of information. In the context of the sequence family, structural and functional data can be evaluated and propagated from known to unknown sequences. However, effective integration is being hindered by syntactic and semantic differences between the different data resources and the alignment techniques employed. One solution to this problem is the development of an ontology that systematically defines the terms used in a specific domain. Ontologies are used to share data from different resources, to automatically analyse information and to represent domain knowledge for non-experts. Here, we present MAO, a new ontology for multiple alignments of nucleic and protein sequences. MAO is designed to improve interoperation and data sharing between different alignment protocols for the construction of a high quality, reliable multiple alignment in order to facilitate knowledge extraction and the presentation of the most pertinent information to the biologist.

Databases, Genetic↗

The mouse secretome: functional classification of the proteins secreted into the extracellular environment.

We have developed a computational strategy to identify the set of soluble proteins secreted into the extracellular environment of a cell. Within the protein sequences predominantly derived from the RIKEN representative transcript and protein set, we identified 2033 unique soluble proteins that are potentially secreted from the cell. These proteins contain a signal peptide required for entry into the secretory pathway and lack any transmembrane domains or intracellular localization signals. This class of proteins, which we have termed the mouse secretome, included >500 novel proteins and 92 proteins <100 amino acids in length. Functional analysis of the secretome included identification of human orthologs, functional units based on InterPro and SCOP Superfamily predictions, and expression of the proteins within the RIKEN READ microarray database. To highlight the utility of this information, we discuss the CUB domain-containing protein family.

Animals↗

Towards cooperative frameworks for modeling and integrating biological processes knowledge.

Data organization has become a strategic target for biologists due to the increasing volume of genomic data available for them. For this purpose, we need a complete knowledge model for representing biological system. In this paper, we deal with both processes for the creation and integration of shareable, reusable domain models within biology, which is a critical issue. In particular, this work introduces a new cooperative development approach for biology ontologies. This approach is based on the integration of the ontologies supplied by different human experts. Two experiments in biological domains are presented and their results discussed.

Artificial Intelligence↗

Rat Phenome Project: the untapped potential of existing rat strains.

The National Bio Resource Project for the Rat in Japan collects, preserves, and distributes rat strains. More than 250 inbred strains have been deposited thus far into the National Bio Resource Project for the Rat and are maintained as specific pathogen-free rats or cryopreserved embryos. We are now comprehensively characterizing deposited strains as part of the Rat Phenome Project to reevaluate their value as models of human diseases. Phenotypic data are being collected for 7 categories and 109 parameters: functional observational battery (neurobehavior), behavior studies, blood pressure, biochemical blood tests, hematology, urology, and anatomy. Furthermore, genotypes are being determined for 370 simple sequence-length polymorphism markers distributed through the whole rat genome. Here, we report these large-scale, high-throughput screening data that have already been collected for 54 rat strains. This comprehensive, original phenotypic data can be systematically viewed by "strain ranking" for each parameter. This allows investigators to explore the relationship between several rat strains, to identify new rat models, and to select the most suitable strains for specific experiments. The discovery of several potential models for human diseases, such as hypertension, hypotension, renal diseases, hyperlipemia, hematological disorders, and neurological disorders, illustrates the potential of many existing rat strains. All deposited strains and obtained data are freely available for any interested researcher worldwide at http://www.anim.med.kyoto-u.ac.jp/nbr.

Animals↗

Assessing protein similarity with Gene Ontology and its use in subnuclear localization prediction.

BACKGROUND: The accomplishment of the various genome sequencing projects resulted in accumulation of massive amount of gene sequence information. This calls for a large-scale computational method for predicting protein localization from sequence. The protein localization can provide valuable information about its molecular function, as well as the biological pathway in which it participates. The prediction of localization of a protein at subnuclear level is a challenging task. In our previous work we proposed an SVM-based system using protein sequence information for this prediction task. In this work, we assess protein similarity with Gene Ontology (GO) and then improve the performance of the system by adding a module of nearest neighbor classifier using a similarity measure derived from the GO annotation terms for protein sequences. RESULTS: The performance of the new system proposed here was compared with our previous system using a set of proteins resided within 6 localizations collected from the Nuclear Protein Database (NPD). The overall MCC (accuracy) is elevated from 0.284 (50.0%) to 0.519 (66.5%) for single-localization proteins in leave-one-out cross-validation; and from 0.420 (65.2%) to 0.541 (65.2%) for an independent set of multi-localization proteins. The new system is available at http://array.bioengr.uic.edu/subnuclear.htm. CONCLUSION: The prediction of protein subnuclear localizations can be largely influenced by various definitions of similarity for a pair of proteins based on different similarity measures of GO terms. Using the sum of similarity scores over the matched GO term pairs for two proteins as the similarity definition produced the best predictive outcome. Substantial improvement in predicting protein subnuclear localizations has been achieved by combining Gene Ontology with sequence information.

Algorithms↗

Pfarao: a web application for protein family analysis customized for cytoskeletal and motor proteins (CyMoBase).

BACKGROUND: Annotation of protein sequences of eukaryotic organisms is crucial for the understanding of their function in the cell. Manual annotation is still by far the most accurate way to correctly predict genes. The classification of protein sequences, their phylogenetic relation and the assignment of function involves information from various sources. This often leads to a collection of heterogeneous data, which is hard to track. Cytoskeletal and motor proteins consist of large and diverse superfamilies comprising up to several dozen members per organism. Up to date there is no integrated tool available to assist in the manual large-scale comparative genomic analysis of protein families. DESCRIPTION: Pfarao (Protein Family Application for Retrieval, Analysis and Organisation) is a database driven online working environment for the analysis of manually annotated protein sequences and their relationship. Currently, the system can store and interrelate a wide range of information about protein sequences, species, phylogenetic relations and sequencing projects as well as links to literature and domain predictions. Sequences can be imported from multiple sequence alignments that are generated during the annotation process. A web interface allows to conveniently browse the database and to compile tabular and graphical summaries of its content. CONCLUSION: We implemented a protein sequence-centric web application to store, organize, interrelate, and present heterogeneous data that is generated in manual genome annotation and comparative genomics. The application has been developed for the analysis of cytoskeletal and motor proteins (CyMoBase) but can easily be adapted for any protein.

Amino Acid Sequence↗

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning↗

GoSurfer: a graphical interactive tool for comparative analysis of large gene sets in Gene Ontology space.

UNLABELLED: The analysis of complex patterns of gene regulation is central to understanding the biology of cells, tissues and organisms. Patterns of gene regulation pertaining to specific biological processes can be revealed by a variety of experimental strategies, particularly microarrays and other highly parallel methods, which generate large datasets linking many genes. Although methods for detecting gene expression have improved substantially in recent years, understanding the physiological implications of complex patterns in gene expression data is a major challenge. This article presents GoSurfer, an easy-to-use graphical exploration tool with built-in statistical features that allow a rapid assessment of the biological functions represented in large gene sets. GoSurfer takes one or two list(s) of gene identifiers (Affymetrix probe set ID) as input and retrieves all the Gene Ontology (GO) terms associated with the input genes. GoSurfer visualises these GO terms in a hierarchical tree format. With GoSurfer, users can perform statistical tests to search for the GO terms that are enriched in the annotations of the input genes. These GO terms can be highlighted on the GO tree. Users can manipulate the GO tree in various ways and interactively query the genes associated with any GO term. The user-generated graphics can be saved as graphics files, and all the GO information related to the input genes can be exported as text files. AVAILABILITY: GoSurfer is a Windows-based program freely available for noncommercial use and can be downloaded at http://www.gosurfer.org. Datasets used to construct the trees shown in the figures in this article are available at http://www.gosurfer.org/download/GoSurfer.zip.

Computer Graphics↗

The HUPO PSI's molecular interaction format--a community standard for the representation of protein interaction data.

A major goal of proteomics is the complete description of the protein interaction network underlying cell physiology. A large number of small scale and, more recently, large-scale experiments have contributed to expanding our understanding of the nature of the interaction network. However, the necessary data integration across experiments is currently hampered by the fragmentation of publicly available protein interaction data, which exists in different formats in databases, on authors' websites or sometimes only in print publications. Here, we propose a community standard data model for the representation and exchange of protein interaction data. This data model has been jointly developed by members of the Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organization (HUPO), and is supported by major protein interaction data providers, in particular the Biomolecular Interaction Network Database (BIND), Cellzome (Heidelberg, Germany), the Database of Interacting Proteins (DIP), Dana Farber Cancer Institute (Boston, MA, USA), the Human Protein Reference Database (HPRD), Hybrigenics (Paris, France), the European Bioinformatics Institute's (EMBL-EBI, Hinxton, UK) IntAct, the Molecular Interactions (MINT, Rome, Italy) database, the Protein-Protein Interaction Database (PPID, Edinburgh, UK) and the Search Tool for the Retrieval of Interacting Genes/Proteins (STRING, EMBL, Heidelberg, Germany).

Database Management Systems↗

Bioinformatics support for high-throughput proteomics.

In the "post-genome" era, mass spectrometry (MS) has become an important method for the analysis of proteome data. The rapid advancement of this technique in combination with other methods used in proteomics results in an increasing number of high-throughput projects. This leads to an increasing amount of data that needs to be archived and analyzed. To cope with the need for automated data conversion, storage, and analysis in the field of proteomics, the open source system ProDB was developed. The system handles data conversion from different mass spectrometer software, automates data analysis, and allows the annotation of MS spectra (e.g. assign gene names, store data on protein modifications). The system is based on an extensible relational database to store the mass spectra together with the experimental setup. It also provides a graphical user interface (GUI) for managing the experimental steps which led to the MS data. Furthermore, it allows the integration of genome and proteome data. Data from an ongoing experiment was used to compare manual and automated analysis. First tests showed that the automation resulted in a significant saving of time. Furthermore, the quality and interpretability of the results was improved in all cases.

Algorithms↗

Bioinformatics and cellular signaling.

The understanding of cellular function requires an integrated analysis of context-specific, spatiotemporal data from diverse sources. Recent advances in describing the genomic and proteomic 'parts list' of the cell and deciphering the interrelationship of these parts are described, including genome-wide location analysis, standards for microarray data analysis, and two-hybrid and mass spectrometry approaches. This information is being collected and curated in databases such as the Alliance for Cellular Signaling (AfCS) Molecule Pages, which will serve as vital tools for the reconstruction and analysis of cellular signaling networks.

Cell Physiological Phenomena↗

Comparison of protein and peptide prefractionation methods for the shotgun proteomic analysis of Synechocystis sp. PCC 6803.

Proteome analysis by gel-free "shotgun" proteomics relies on the simplification of a peptide mixture before it is analyzed in a mass spectrometer. While separation on a reverse-phase (RP) liquid chromatographic column is widely employed, a variety of other methods have been used to fractionate both proteins and peptides before this step. We compared six different protein and peptide fractionation workflows, using Synechocystis sp. PCC 6803, a useful model cyanobacterium for potential exploitation to improve its production of hydrogen and other secondary metabolites. Pre-digestion protein separation was performed by strip-based isoelectric focusing, one-dimensional polyacrylamide gel electrophoresis, or weak anion exchange chromatography, while pre-RP peptide separation was accomplished by isoelectric focusing (IEF) or strong cation exchange chromatography. Peptides were identified using electrospray ionization quadrupole time of flight-tandem mass spectrometry. Mass spectrometry (MS) and tandem mass spectra were analyzed using ProID software employing both a single organism database and the entire NCBI non-redundant database, and a total of 776 proteins were identified using a stringent set of selection criteria. Method comparisons were made on the basis of the results obtained (number and types of proteins identified), as well as ease of use and other practical aspects. IEF-IEF protein and peptide fractionation prior to RP gave the best overall performance.

Chromatography, High Pressure Liquid↗

Autophagy in chronically ischemic myocardium.

We tested the hypothesis that chronically ischemic (IS) myocardium induces autophagy, a cellular degradation process responsible for the turnover of unnecessary or dysfunctional organelles and cytoplasmic proteins, which could protect against the consequences of further ischemia. Chronically instrumented pigs were studied with repetitive myocardial ischemia produced by one, three, or six episodes of 90 min of coronary stenosis (30% reduction in baseline coronary flow followed by reperfusion every 12 h) with the non-IS region as control. In this model, wall thickening in the IS region was chronically depressed by approximately 37%. Using a nonbiased proteomic approach combining 2D gel electrophoresis with in-gel proteolysis, peptide mapping by MS, and sequence database searches for protein identification, we demonstrated increased expression of cathepsin D, a protein known to mediate autophagy. Additional autophagic proteins, cathepsin B, heat shock cognate protein Hsc73 (a key protein marker for chaperone-mediated autophagy), beclin 1 (a mammalian autophagy gene), and the processed form of microtubule-associated protein 1 light chain 3 (a marker for autophagosomes), were also increased. These changes, not evident after one episode, began to appear after two or three episodes and were most marked after six episodes of ischemia, when EM demonstrated autophagic vacuoles in chronically IS myocytes. Conversely, apoptosis, which was most marked after three episodes, decreased strikingly after six episodes, when autophagy had increased. Immunohistochemistry staining for cathepsin B was more intense in areas where apoptosis was absent. Thus, autophagy, triggered by ischemia, could be a homeostatic mechanism, by which apoptosis is inhibited and the deleterious effects of chronic ischemia are limited.

Animals↗