Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Prediction of protein function using protein-protein interaction data.

Assigning functions to novel proteins is one of the most important problems in the postgenomic era. Several approaches have been applied to this problem, including the analysis of gene expression patterns, phylogenetic profiles, protein fusions, and protein-protein interactions. In this paper, we develop a novel approach that employs the theory of Markov random fields to infer a protein's functions using protein-protein interaction data and the functional annotations of protein's interaction partners. For each function of interest and protein, we predict the probability that the protein has such function using Bayesian approaches. Unlike other available approaches for protein annotation in which a protein has or does not have a function of interest, we give a probability for having the function. This probability indicates how confident we are about the prediction. We employ our method to predict protein functions based on "biochemical function," "subcellular location," and "cellular role" for yeast proteins defined in the Yeast Proteome Database (YPD, www.incyte.com), using the protein-protein interaction data from the Munich Information Center for Protein Sequences (MIPS, mips.gsf.de). We show that our approach outperforms other available methods for function prediction based on protein interaction data. The supplementary data is available at www-hto.usc.edu/~msms/ProteinFunction.

Bayes Theorem↗

Multiple urinary peptides are associated with hypertension: a link to molecular pathophysiology.

OBJECTIVES: Hypertension is a common condition worldwide; however, its underlying mechanisms remain largely unknown. This study aimed to identify urinary peptides associated with hypertension to further explore the relevant molecular pathophysiology. METHODS: Peptidome data from 2876 individuals without end-organ damage were retrieved from the Human Urinary Proteome Database, belonging to general population (discovery) or type 2 diabetic (validation) cohorts. Participants were divided based on systolic blood pressure (SBP) and diastolic BP (DBP) into hypertensive (SBP &#x2265;140&#x200a;mmHg and/or DBP &#x2265;90&#x200a;mmHg) and normotensive (SBP <120&#x200a;mmHg and DBP <80&#x200a;mmHg, without antihypertensive treatment) groups. Differences in peptide abundance between the two groups were confirmed using an external cohort ( n &#x200a;=&#x200a;420) of participants without end-organ damage, matched for age, BMI, eGFR, sex, and the presence of diabetes. Furthermore, the association of the peptides with BP as a continuous variable was investigated. The findings were compared with peptide biomarkers of chronic diseases and bioinformatic analyses were conducted to highlight the underlying molecular mechanisms. RESULTS: Between hypertensive and normotensive individuals, 96 (mostly COL1A1 and COL3A1) peptides were found to be significantly different in both the discovery (adjusted) and validation (nominal significance) cohorts, with consistent regulation. Of these, 83 were consistently regulated in the matched cohort. A weak, yet significant, association between their abundance and standardized BP was also observed. CONCLUSION: Hypertension is associated with an altered urinary peptide profile with evident differential regulation of collagen-derived peptides. Peptides related to vascular calcification and sodium regulation were also affected. Whether these modifications reflect the pathophysiology of hypertension and/or early subclinical organ damage requires further investigation.

Humans↗

Comparing low coverage random shotgun sequence data from Brassica oleracea and Oryza sativa genome sequence for their ability to add to the annotation of Arabidopsis thaliana.

Since the completion of the Arabidopsis thaliana genome sequence, there is an ongoing effort to annotate the genome as accurately as possible. Comparing genome sequences of related species complements the current annotation strategies by identifying genes and improving gene structure. A total of 595,321 Brassica oleracea shotgun reads were sequenced by TIGR (The Institute for Genome Research) and the collaboration of Washington University and Cold Spring Harbor. Vicogenta (a genome viewer based on GMOD and GBrowse) was created to view the current annotation and sequence alignments for Arabidopsis. Brassica reads were compared with the Arabidopsis genome and proteome databases using BLAST. Hypothetical genes and conserved unannotated regions on the short arm of chromosome 4 from Arabidopsis were experimentally verified using RT-PCR. We were able to improve the Arabidopsis annotation by identifying 25 genes that were missed, and confirming expression of 43 hypothetical genes in Arabidopsis. We were also able to detect conservation in genes whose transcription is normally suppressed due to methylation. We also examined how useful the O. sativa genome and ESTs from other species are, compared with Brassica, in improving the Arabidopsis annotation.

Amino Acid Sequence↗

Cell cycle progression in G1 and S phases is CCR4 dependent following ionizing radiation or replication stress in Saccharomyces cerevisiae.

To identify new nonessential genes that affect genome integrity, we completed a screening for diploid mutant Saccharomyces cerevisiae strains that are sensitive to ionizing radiation (IR) and found 62 new genes that confer resistance. Along with those previously reported (Bennett et al., Nat. Genet. 29:426-434, 2001), these genes bring to 169 the total number of new IR resistance genes identified. Through the use of existing genetic and proteomic databases, many of these genes were found to interact in a damage response network with the transcription factor Ccr4, a core component of the CCR4-NOT and RNA polymerase-associated factor 1 (PAF1)-CDC73 transcription complexes. Deletions of individual members of these two complexes render cells sensitive to the lethal effects of IR as diploids, but not as haploids, indicating that the diploid G1 cell population is radiosensitive. Consistent with a role in G1, diploid ccr4Delta cells irradiated in G1 show enhanced lethality compared to cells exposed as a synchronous G2 population. In addition, a prolonged RAD9-dependent G1 arrest occurred following IR of ccr4Delta cells and CCR4 is a member of the RAD9 epistasis group, thus confirming a role for CCR4 in checkpoint control. Moreover, ccr4Delta cells that transit S phase in the presence of the replication inhibitor hydroxyurea (HU) undergo prolonged cell cycle arrest at G2 followed by cellular lysis. This S-phase replication defect is separate from that seen for rad52 mutants, since rad52Delta ccr4Delta cells show increased sensitivity to HU compared to rad52Delta or ccr4Delta mutants alone. These results indicate that cell cycle transition through G1 and S phases is CCR4 dependent following radiation or replication stress.

Cell Cycle Proteins↗

Integrative plant biology: role of phloem long-distance macromolecular trafficking.

Recent studies have revealed the operation of a long-distance communication network operating within the vascular system of higher plants. The evolutionary development of this network reflects the need to communicate environmental inputs, sensed by mature organs, to meristematic regions of the plant. One consequence of such a long-distance signaling system is that newly forming organs can develop properties optimized for the environment into which they will emerge, mature, and function. The phloem translocation stream of the angiosperms contains, in addition to photosynthate and other small molecules, a variety of macromolecules, including mRNA, small RNA, and proteins. This review highlights recent progress in the characterization of phloem-mediated transport of macromolecules as components of an integrated long-distance signaling network. Attention is focused on the role played by these proteins and RNA species in coordination of developmental programs and the plant's response to both environmental cues and pathogen challenge. Finally, the importance of developing phloem transcriptome and proteomic databases is discussed within the context of advances in plant systems biology.

Biological Transport↗

A combined in vitro/bioinformatic investigation of redox regulatory mechanisms governing cell cycle progression.

The intracellular reduction-oxidation (redox) environment influences cell cycle progression; however, underlying mechanisms are poorly understood. To examine potential mechanisms, the intracellular redox environment was characterized per cell cycle phase in Chinese hamster ovary fibroblasts via flow cytometry by measuring reduced glutathione (GSH), reactive oxygen species (ROS), and DNA content with monochlorobimane, 2',7'-dichlorohydrofluorescein diacetate (H2DCFDA), and DRAQ5, respectively. GSH content was significantly greater in G2/M compared with G1 phase cells, whereas GSH was intermediate in S phase cells. ROS content was similar among phases. Together, these data demonstrate that G2/M cells are more reduced than G1 cells. Conventional approaches to define regulatory mechanisms are subjective in nature and focus on single proteins/pathways. Proteome databases provide a means to overcome these inherent limitations. Therefore, a novel bioinformatic approach was developed to exhaustively identify putative redox-regulated cell cycle proteins containing redox-sensitive protein motifs. Using the InterPro (http://www.ebi.ac.uk/interpro/) database, we categorized 536 redox-sensitive motifs as: 1) active/functional-site cysteines, 2) electron transport, 3) heme, 4) iron binding, 5) zinc binding, 6) metal binding (non-Fe/Zn), and 7) disulfides. Comparing this list with 1,634 cell cycle-associated proteins from Swiss-Prot and SpTrEMBL (http://us.expasy.org/sprot/) revealed 92 candidate proteins. Three-fourths (69 of 92) of the candidate proteins function in the central cell cycle processes of transcription, nucleotide metabolism, (de)phosphorylation, and (de)ubiquitinylation. The majority of oxidant-sensitive candidate proteins (68.9%) function during G2/M phase. As the G2/M phase is more reduced than the G1 phase, oxidant-sensitive proteins may be temporally regulated by oscillation of the intracellular redox environment. Combined with evidence of intracellular redox compartmentalization, we propose a spatiotemporal mechanism that functionally links an oscillating intracellular redox environment with cell cycle progression.

Amino Acid Motifs↗

Model-driven user interfaces for bioinformatics data resources: regenerating the wheel as an alternative to reinventing it.

BACKGROUND: The proliferation of data repositories in bioinformatics has resulted in the development of numerous interfaces that allow scientists to browse, search and analyse the data that they contain. Interfaces typically support repository access by means of web pages, but other means are also used, such as desktop applications and command line tools. Interfaces often duplicate functionality amongst each other, and this implies that associated development activities are repeated in different laboratories. Interfaces developed by public laboratories are often created with limited developer resources. In such environments, reducing the time spent on creating user interfaces allows for a better deployment of resources for specialised tasks, such as data integration or analysis. Laboratories maintaining data resources are challenged to reconcile requirements for software that is reliable, functional and flexible with limitations on software development resources. RESULTS: This paper proposes a model-driven approach for the partial generation of user interfaces for searching and browsing bioinformatics data repositories. Inspired by the Model Driven Architecture (MDA) of the Object Management Group (OMG), we have developed a system that generates interfaces designed for use with bioinformatics resources. This approach helps laboratory domain experts decrease the amount of time they have to spend dealing with the repetitive aspects of user interface development. As a result, the amount of time they can spend on gathering requirements and helping develop specialised features increases. The resulting system is known as Pierre, and has been validated through its application to use cases in the life sciences, including the PEDRoDB proteomics database and the e-Fungi data warehouse. CONCLUSION: MDAs focus on generating software from models that describe aspects of service capabilities, and can be applied to support rapid development of repository interfaces in bioinformatics. The Pierre MDA is capable of supporting common database access requirements with a variety of auto-generated interfaces and across a variety of repositories. With Pierre, four kinds of interfaces are generated: web, stand-alone application, text-menu, and command line. The kinds of repositories with which Pierre interfaces have been used are relational, XML and object databases.

Computational Biology↗

Gene expression analysis reveals that histone deacetylation sites may serve as partitions of chromatin gene expression domains.

BACKGROUND: It has been a long-term puzzle whether chromatin can be further divided into distinct gene expression domains. Because histone deacetylation affects chromatin structure, that in turn may affect the expression of nearby genes, histone deacetylation sites may act to partition chromatin into different gene expression domains. In this article, we explore the relationship between histone deacetylation sites and gene expression patterns on the genome scale using different data sources, including microarray data measuring gene expression levels, microarray data measuring histone deacetylation sites, and information on regulatory targets of transcription factors. RESULTS: Using 269 Saccharomyces cerevisiae microarray datasets, histone deacetylation datasets, and regulatory targets of transcription factors assembled from the Yeast Proteome Database and ChIP-chip data, we found that histone deacetylation sites can reduce the level of co-expression of neighboring genes. CONCLUSION: Histone deacetylation sites may serve as possible partition sites for chromatin domains and affect gene expression.

Binding Sites↗

Two-dimensional gel proteome reference map of blood monocytes.

BACKGROUND: Blood monocytes play a central role in regulating host inflammatory processes through chemotaxis, phagocytosis, and cytokine production. However, the molecular details underlying these diverse functions are not completely understood. Understanding the proteomes of blood monocytes will provide new insights into their biological role in health and diseases. RESULTS: In this study, monocytes were isolated from five healthy donors. Whole monocyte lysates from each donor were then analyzed by 2D gel electrophoresis, and proteins were detected using Sypro Ruby fluorescence and then examined for phosphoproteomes using ProQ phospho-protein fluorescence dye. Between 1525 and 1769 protein spots on each 2D gel were matched, analyzed, and quantified. Abundant protein spots were then subjected to analysis by mass spectrometry. This report describes the protein identities of 231 monocyte protein spots, which represent 164 distinct proteins and their respective isoforms or subunits. Some of these proteins had not been previously characterized at the protein level in monocytes. Among the 231 protein spots, 19 proteins revealed distinct modification by protein phosphorylation. CONCLUSION: The results of this study offer the most detailed monocyte proteomic database to date and provide new perspectives into the study of monocyte biology.

Journal Article↗

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus↗

A flexible integration and visualisation system for biomarker discovery.

Biological data have accumulated at an unprecedented pace as a result of improvements in molecular technologies. However, the translation of data into information, and subsequently into knowledge, requires the intricate interplay of data access, visualisation and interpretation. Biological data are complex and are organised either hierarchically or non-hierarchically. For non-hierarchically organised data, it is difficult to view relationships among biological facts. In addition, it is difficult to make changes in underlying data storage without affecting the visualisation interface. Here, we demonstrate a platform where non-hierarchically organised data can be visualised through the application of a customised hierarchy incorporating medical subject headings (MeSH) classifications. This platform gives users flexibility in updating and manipulation. It can also facilitate fresh scientific insight by highlighting biological impacts across different hierarchical branches. An example of the integration of biomarker information from the curated Proteome database using MeSH and the StarTree visualisation tool is presented.

Algorithms↗

Antibacterial activities of peptides from the water-soluble extracts of Italian cheese varieties.

Water-soluble extracts of 9 Italian cheese varieties that differed mainly for type of cheese milk, starter, technology, and time of ripening were fractionated by reversed-phase fast protein liquid chromatography, and the antimicrobial activity of each fraction was first assayed toward Lactobacillus sakei A15 by well-diffusion assay. Active fractions were further analyzed by HPLC coupled to electrospray ionization-ion trap mass spectrometry, and peptide sequences were identified by comparison with a proteomic database. Parmigiano Reggiano, Fossa, and Gorgonzola water-soluble extracts did not show antibacterial peptides. Fractions of Pecorino Romano, Canestrato Pugliese, Crescenza, and Caprino del Piemonte contained a mixture of peptides with a high degree of homology. Pasta filata cheeses (Caciocavallo and Mozzarella) also had antibacterial peptides. Peptides showed high levels of homology with N-terminal, C-terminal, or whole fragments of well known antimicrobial or multifunctional peptides reported in the literature: alphaS1-casokinin (e.g., sheep alphaS1-casein (CN) f22-30 of Pecorino Romano and cow alphaS1-CN f24-33 of Canestrato Pugliese); isracidin (e.g., sheep alphaS1-CN f10-21 of Pecorino Romano); kappacin and casoplatelin (e.g., cow kappa-CN f106-115 of Canestrato Pugliese and Crescenza); and beta-casomorphin-11 (e.g., goat beta-CN f60-68 of Caprino del Piemonte). As shown by the broth microdilution technique, most of the water-soluble fractions had a large spectrum of inhibition (minimal inhibitory concentration of 20 to 200 microg/mL) toward gram-positive and gram-negative bacterial species, including potentially pathogenic bacteria of clinical interest. Cheeses manufactured from different types of cheese milk (cow, sheep, and goat) have the potential to generate similar peptides with antimicrobial activity.

Amino Acid Sequence↗

Prediction of protein function using protein-protein interaction data.

Assigning functions to novel proteins is one of the most important problems in the post-genomic era. Several approaches have been applied to this problem, including analyzing gene expression patterns, phylogenetic profiles, protein fusions and protein-protein interactions. We develop a novel approach that applies the theory of Markov random fields to infer a protein's functions using protein-protein interaction data and the functional annotations of its interaction protein partners. For each function of interest and a protein, we predict the probability that the protein has that function using Bayesian approaches. Unlike in other available approaches for protein annotation where a protein has or does not have a function of interest, we give a probability for having the function. This probability indicates how confident we are about the prediction. We apply our method to predict cellular functions (43 categories including a category "others") for yeast proteins defined in the Yeast Proteome Database (YPD), using the protein-protein interaction data from the Munich Information Center for Protein Sequences (MIPS, http://mips.gsf.de). We show that our approach outperforms other available methods for function prediction based on protein interaction data.

Amino Acid Sequence↗

The wildcat toolbox: a set of perl script utilities for use in peptide mass spectral database searching and proteomics experiments.

We describe in this communication a set of functional perl script utilities for use in peptide mass spectral database searching and proteomics experiments, known as the Wildcat Toolbox. These are all freely available for download from our laboratory Web site (http://proteomics.arizona.edu/toolbox.html) as a combined zip file, and can also be accessed via the Proteome Commons Web site (www.proteomecommons.org) in the tools section. We make them available to other potential users in the spirit of open source software development; we do not have the resources to provide any significant technical support for them, but we hope users will share both bugs and improvements with the community at large.

Algorithms↗

Tissue Molecular Anatomy Project (TMAP): an expression database for comparative cancer proteomics.

By mining publicly accessible databases, we have developed a collection of tissue-specific predictive protein expression maps as a function of cancer histological state. Data analysis is applied to the differential expression of gene products in pooled libraries from the normal to the altered state(s). We wish to report the initial results of our survey across different tissues and explore the extent to which this comparative approach may help uncover panels of potential biomarkers of tumorigenesis which would warrant further examination in the laboratory.

Databases, Protein↗

Yeast Protein database (YPD): a database for the complete proteome of Saccharomyces cerevisiae.

The Yeast Protein Database (YPD) is a database for the proteins of the budding yeast,Saccharomyces cerevisiae. YPD is the first annotated database for the complete proteome of any organism. Now that the complete genome sequence of yeast is available, YPD contains entries for each of the characterized proteins and for each of the uncharacterized proteins predicted from the sequence. Contained in YPD are the calculated properties of each protein such as molecular weight and isoelectric point, experimentally determined properties such as subcellular localization and post-translational modifications, and extensive annotations from the yeast literature. YPD contains 25 000 lines of textual annotation that describe the known functions, mutant phenotypes, interactions, and other properties for the approximately 6000 proteins in the yeast proteome. The information in YPD is updated daily, and it is available on the World Wide Web at http://www.proteome.com/YPDhome.html .

Amino Acid Sequence↗

A dynamic two-dimensional polyacrylamide gel electrophoresis database: the mycobacterial proteome via Internet.

Proteome analysis by two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) and mass spectrometry, in combination with protein chemical methods, is a powerful approach for the analysis of the protein composition of complex biological samples. Data organization is imperative for efficient handling of the vast amount of information generated. Thus we have constructed a 2-D PAGE database to store and compare protein patterns of cell-associated and culture-supernatant proteins of different mycobacterial strains. In accordance with the guidelines for federated 2-DE databases, we developed a program that generates a dynamic 2-D PAGE database for the World-Wide-Web to organise and publish, via the internet, our results from proteome analysis of different Mycobacterium tuberculosis as well as Mycobacterium bovis BCG strains. The uniform resource locator for the database is http://www.mpiib-berlin.mpg.de/2D-PAGE and can be read with a Java compatible browser. The interactive hypertext markup language documents displayed are generated dynamically in each individual session from a rational data file, a 2-D gel image file and a map file describing the protein spots as polygons. The program consists of common gateway interface scripts written in PERL, minimizing the administrative workload of the database. Furthermore, the database facilitates not only interactive use, but also worldwide active participation of other scientific groups with their own data, requiring only minimal computer hardware and knowledge of information technology.

Bacterial Proteins↗