Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

An integrated strategy for the optimization of microarray data interpretation.

The completion of a microarray experiment represents just a starting point toward understanding the biology of interest. A follow-up strategy is needed to fully elucidate the functional significance of microarray-derived measurements of differential expression. Given the fact that no single approach can fully unravel the fundamental biology that is typically quite complex, the follow-up strategy must be integrated at multiple levels encompassing bioinformatics, genomics, and proteomics. In this review, we discuss an integrative approach, which can be used to prioritize microarray-derived candidate genes, define their functions, and place them in the context of the biological system being studied.

Animals↗

ProtoBee: hierarchical classification and annotation of the honey bee proteome.

The recently sequenced genome of the honey bee (Apis mellifera) has produced 10,157 predicted protein sequences, calling for a computational effort to extract biological insights from them. We have applied an unsupervised hierarchical protein-clustering method, which was previously used in the ProtoNet system, to nearly 200,000 proteins consisting of the predicted honey bee proteins, the SWISS-PROT protein database, and the complete set of proteins of the mouse (Mus musculus) and the fruit fly (Drosophila melanogaster). The hierarchy produced by this method has been entitled ProtoBee. In ProtoBee, the proteins are hierarchically organized into 18,936 separate tree hierarchies, each representing a protein functional family. By using the mouse and Drosophila complete proteomes as reference, we are able to highlight functional groups of putative gene-loss events, putative novel proteins of unique functionality, and bee-specific paralogs. We have studied some of the ProtoBee findings and suggest their biological relevance. Examples include novel opsin genes and intriguing nuclear matches of mitochondrial genes. The organization of bee sequences into functional clusters suggests a natural way of automatically inferring functional annotation. Following this notion, we were able to assign functional annotation to about 70% of the sequences. ProtoBee is available at http://www.protobee.cs.huji.ac.il.

Animals↗

The human pituitary proteome: the characterization of differentially expressed proteins in an adenoma compared to a control.

In order to clarify the basic molecular mechanisms that participate in the formation of human pituitary macroadenomas, this study, for the first time, describes the comparative proteomics between a pituitary adenoma tissue and a control tissue. A vertical, two-dimensional polyacrylamide gel electrophoresis system and PDQuest image analysis software were used to provide a high level of between-gel reproducibility and electrophoretic separation to accurately locate each differentially expressed protein. Mass spectrometry (MALDI-TOF and LC-ESI-Q-IT) and protein databases were used to characterize each differentially expressed protein. A total of 137 differential gel spots (37 increased spot volumes, 39 decreased, 19 new and 42 lost) were found when we compared an adenoma proteome to a control proteome. Seventy-one spots (20 increased, 27 decreased, 13 new, 11 lost), representing 39 differentially regulated proteins, were identified. Five differentially regulated proteins (prolactin, cellular retinoic acid-binding protein II, G-protein beta subunit 3, secretagogin and calreticulin) were also validated with results from a comparative transcriptomics study of pituitary adenomas and controls. The functional characteristics of these differentially expressed proteins provide a differential proteomic profile between a pituitary adenoma and a control.

Adenoma↗

CIBEX: center for information biology gene expression database.

We describe the current status of the gene expression database CIBEX (Center for Information Biology gene EXpression database, http://cibex.nig.ac.jp), with a data retrieval system in compliance with MIAME, a standard that the MGED Society has developed for comparing and data produced in microarray experiments at different laboratories worldwide. CIBEX serves as a public repository for a wide range of high-throughput experimental data in gene expression research, including microarray-based experiments measuring mRNA, serial analysis of gene expression (SAGE tags), and mass spectrometry proteomic data.

Computational Biology↗

The human urinary proteome contains more than 1500 proteins, including a large proportion of membrane proteins.

BACKGROUND: Urine is a desirable material for the diagnosis and classification of diseases because of the convenience of its collection in large amounts; however, all of the urinary proteome catalogs currently being generated have limitations in their depth and confidence of identification. Our laboratory has developed methods for the in-depth characterization of body fluids; these involve a linear ion trap-Fourier transform (LTQ-FT) and a linear ion trap-orbitrap (LTQ-Orbitrap) mass spectrometer. Here we applied these methods to the analysis of the human urinary proteome. RESULTS: We employed one-dimensional sodium dodecyl sulfate polyacrylamide gel electrophoresis and reverse phase high-performance liquid chromatography for protein separation and fractionation. Fractionated proteins were digested in-gel or in-solution, and digests were analyzed with the LTQ-FT and LTQ-Orbitrap at parts per million accuracy and with two consecutive stages of mass spectrometric fragmentation. We identified 1543 proteins in urine obtained from ten healthy donors, while essentially eliminating false-positive identifications. Surprisingly, nearly half of the annotated proteins were membrane proteins according to Gene Ontology (GO) analysis. Furthermore, extracellular, lysosomal, and plasma membrane proteins were enriched in the urine compared with all GO entries. Plasma membrane proteins are probably present in urine by secretion in exosomes. CONCLUSION: Our analysis provides a high-confidence set of proteins present in human urinary proteome and provides a useful reference for comparing datasets obtained using different methodologies. The urinary proteome is unexpectedly complex and may prove useful in biomarker discovery in the future.

Chromatography, High Pressure Liquid↗

SpecAlign--processing and alignment of mass spectra datasets.

SUMMARY: Pre-processing of chromatographic profile or mass spectral data is an important aspect of many types of proteomics and biomarker discovery experiments. Here we present a graphical computational tool, SpecAlign, that enables simultaneous visualization and manipulation of multiple datasets. SpecAlign not only provides all common processing functions, but also uniquely implements an algorithm that enables the complete alignment of each mass spectrum within a loaded dataset. We demonstrate its utility by aligning two datasets each containing six spectra; one set was acquired prior to instrument calibration and the other following calibration. AVAILABILITY: The software is free of charge and available for download from http://ptcl.chem.ox.ac.uk/~jwong/specalign. Supports Windows operating systems including Windows 9X/NT/2000/XP.

Algorithms↗

Evaluation of algorithms for protein identification from sequence databases using mass spectrometry data.

In this work, the commonly used algorithms for mass spectrometry based protein identification, Mascot, MS-Fit, ProFound and SEQUEST, were studied in respect to the selectivity and sensitivity of their searches. The influence of various search parameters were also investigated. Approximately 6600 searches were performed using different search engines with several search parameters to establish a statistical basis. The applied mass spectrometric data set was chosen from a current proteome study. The huge amount of data could only be handled with computational assistance. We present a software solution for fully automated triggering of several peptide mass fingerprinting (PMF) and peptide fragmentation fingerprinting (PFF) algorithms. The development of this high-throughput method made an intensive evaluation based on data acquired in a typical proteome project possible. Previous evaluations of PMF and PFF algorithms were mainly based on simulations.

Algorithms↗

Chromosome mapping and identification of amphiphilic proteins of hexaploid wheat kernels.

Amphiphilic proteomic analysis was carried out on the ITMI (International Triticae Mapping Population) population resulting from a cross between "Synthetic", i.e.: "W7984" and "Opata". Out of a total of 446 spots, 170 were specific to either of the two parents, and 276 were common to both. Preliminary analysis, which was performed on 80 progenies (Amiour et al. 2002a), was completed here using a total of 101 selfed lines. Seventy two Loci of amphiphilic spots placed at LOD = 5 were conclusively assigned to 15 chromosomes. Some spots mapped during the first analysis were eliminated because of the significant distortion segregation observed in the second analysis. Group-1 chromosomes had by far the greatest number of mapped spots (51). Using the Quantitative Trait Loci (QTLs) approach, analysis of the quantitative variation of each spot revealed that 96 spots out of the 170 specific ones showed at least one Protein Quantity Locus (PQL). These PQLs were distributed throughout the genome. With Matrix Laser Desorption Ionisation Time Of Flight (MALDI-TOF) spectrometry and Database interrogation, a total of 93 specific and 41 common spots were identified. This enabled us to show that the majority of these proteins are associated with membranes and/or play a role in plant defence against external invasions. Using multiple-regression analysis, other amphiphilic proteins, in addition to puroindolines, were shown to be involved in variation in kernel hardness in the ITMI population.

Chromosome Mapping↗

Insights into the evolution of the nucleolus by an analysis of its protein domain repertoire.

Recently, the first investigation of nucleoli using mass spectrometry led to the identification of 271 proteins. This represents a rich resource for a comprehensive investigation of nucleolus evolution. We applied a protocol for the identification of known and novel conserved protein domains of the nucleolus, resulting in the identification of 115 known and 91 novel domain profiles. The phyletic distribution of nucleolar protein domains in a collection of complete proteomes of selected organisms from all domains of life confirms the archaebacterial origin of the core machinery for ribosome maturation and assembly, but also reveals substantial eubacterial and eukaryotic contributions to nucleolus evolution. We predict that, in different phases of nucleolus evolution, protein domains with different biochemical functions were recruited to the nucleolus. We suggest a model for the late and continuous evolution of the nucleolus in early eukaryotes and argue against an endosymbiotic origin of the nucleolus and the nucleus. Supplementary material for this article can be found on the BioEssays website at http://www.interscience.wiley.com/jpages/0265-9247/suppmat/index.html.

Archaeal Proteins↗

GLYCOSCIENCES.de: an Internet portal to support glycomics and glycobiology research.

The development of glycan-related databases and bioinformatics applications is considerably lagging behind compared with the wealth of available data and software tools in genomics and proteomics. Because the encoding of glycan structures is more complex, most of the bioinformatics approaches cannot be applied to glycan structures. No standard procedures exist where glycan structures found in various species, organs, tissues or cells can be routinely deposited. In this article the concepts of the GLYCOSCIENCES.de portal are described. It is demonstrated how an efficient structure-based cross-linking of various glycan-related data originating from different resources can be accomplished using a single user interface. The structure oriented retrieval options-exact structure, substructure, motif, composition and sugar components-are discussed. The types of available data-references, composition, spatial structures, nuclear magnetic resonance (NMR) shifts (experimental and estimated), theoretically calculated fragments and Protein Database (PDB) entries-are exemplified for Man(3.) The free availability and unrestricted use of glycan-related data is an absolute prerequisite to efficiently share distributed resources. Additionally, there is an urgent need to agree to a generally accepted exchange format as well as to a common software interface. An open access repository for glyco-related experimental data will secure that the loss of primary data will be considerably reduced.

Computational Biology↗

Proteome analysis of mouse primary astrocytes.

Astrocytes play a role in energy metabolism, neuronal homeostasis and release of neuronal growth factors and several neurotransmitters. They also relate to a variety of brain diseases and contribute to restore brain dysfunction. Although current research has revealed several roles for astrocytes, knowledge on astrocytic protein expression is limited and a systematic and comprehensive proteome study of astrocytes has not been reported so far. We applied a proteomics technique based on two-dimensional gel electrophoresis coupled with mass spectrometry (MALDI-TOF/TOF) and unambiguously identified 301 spots corresponding to 191 individual proteins in primary mouse astrocytes. The identified proteins were from antioxidant, chaperone, cytoskeleton, nucleic acid binding, signaling, proteasomal, hypothetical and miscellaneous proteins. A reference database is provided and proteins were identified in astrocytes specifically and unambiguously for the first time. A reliable analytical tool independent of antibody availability and specificity along with tentative astrocytic marker proteins is described.

Animals↗

Experimental and bioinformatic approaches for interrogating protein-protein interactions to determine protein function.

An ambitious goal of proteomics is to elucidate the structure, interactions and functions of all proteins within cells and organisms. One strategy to determine protein function is to identify the protein-protein interactions. The increasing use of high-throughput and large-scale bioinformatics-based studies has generated a massive amount of data stored in a number of different databases. A challenge for bioinformatics is to explore this disparate data and to uncover biologically relevant interactions and pathways. In parallel, there is clearly a need for the development of approaches that can predict novel protein-protein interaction networks in silico. Here, we present an overview of different experimental and bioinformatic methods to elucidate protein-protein interactions.

Animals↗

Two-hybrid systematic screening of the yeast proteome.

The yeast two-hybrid system is a genetic method that detects protein-protein interactions. One application is the detection by library screening of new interactors of a protein of known function. In the August issue of Nature Genetics, Fromont-Racine et al. showed for the first time that the construction of the protein interaction map of a complex pathway, such as that of the mRNA splicing machinery, is now possible, because of the combination of recent technical improvements elaborated in several laboratories. With a yeast cell mating procedure that increases screen efficiency, they used their complex yeast genomic library of 5 x 10(6) clones to test 700 x 10(6) interactions against 15 proteins. They identified and classified 170 potential interactors, including approximately 70 proteins of previously unknown function. More than 25% of the interactors are probably biologically relevant. The achievements of Fromont-Racine et al. have opened the way to the systematic analysis of the protein interaction networks of the 6,000 open reading frames-yeast proteome. This task requires, however, automation of the library screens and creation of a two-hybrid library database.

Fungal Proteins↗

Short-term culturing of low-grade superficial bladder transitional cell carcinomas leads to changes in the expression levels of several proteins involved in key cellular activities.

Fresh, superficial transitional cell carcinomas (TCCs) of low-grade atypia (3 grade I, Ta; 6 grade II, Ta), as well as primary cultures derived from them were labeled with [35S]methionine for 16 h, between 2 and 6 days after inoculation. Whole protein extracts were subjected to IEF (isoelectric focusing) two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) followed by autoradiography. Proteins were identified by a combination of proteomic technologies that included microsequencing, mass spectrometry, 2-D PAGE immunoblotting and comparison with the bladder TCC protein database available on the internet (http://biobase.dk/cgi-bin/celis). Comparison of the IEF 2-D gel protein profiles of fresh tumors and their primary cultures showed that the overall expression profiles were strikingly similar, although differing significantly in the levels of several proteins whose rate of synthesis was differentially regulated in at least 85% of the tumor/culture pairs as a result of the short-term culturing. Most of the proteins affected by culturing were upregulated and among them we identified components of the cytoskeleton (keratin 18, gelsolin and tropomyosin 3), a molecular chaperone (hsp 28), aldose reductase, GST pi, metastasin, synuclein, the calreticulin precursor and three polypeptides of unknown identity. Only four major proteins were downregulated, and these included two fatty acid-binding proteins (FABP:FABP5 and A-FABP) which are thought to play a role in growth control, the differentiation-associated keratin 20, and the calcium-binding protein annexin V. Proteins that were differentially regulated in only some of the cultured tumors included alpha-enolase, triosphosphate isomerase, members of the 14-3-3 family, hnRNPs F and H, PGDH, hsp (heat-shock protein) 60, BIP, the interleukin-1 receptor antagonist, the nucleolar protein B23, as well as several proteins of yet unknown identity. The suitability of in vitro bladder tumor culture models to study complex biological phenomena such as malignancy and invasion is discussed.

Carcinoma, Transitional Cell↗

Protein interaction networks of Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster: large-scale organization and robustness.

High-throughput screens have begun to reveal protein interaction networks in several organisms. To understand the general properties of these protein interaction networks, a systematic analysis of topological structure and robustness was performed on the protein interaction networks of Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster. It shows that the three protein interaction networks have a scale-free and high-degree clustering nature as the consequence of their hierarchical organization. It also shows that they have the small-world property with similar diameter at 4-5. Evaluation of the consequences of random removal of both proteins and interactions from the protein interaction networks suggests their high degree of robustness. Simulation of a protein's removal shows that the protein interaction network's error tolerance is accompanied by attack vulnerability. These fundamental analyses of the networks might serve as a starting point for further exploring complex biological networks and the coming research of "systems biology".

Algorithms↗

Binding MOAD (Mother Of All Databases).

Binding MOAD (Mother of All Databases) is the largest collection of high-quality, protein-ligand complexes available from the Protein Data Bank. At this time, Binding MOAD contains 5331 protein-ligand complexes comprised of 1780 unique protein families and 2630 unique ligands. We have searched the crystallography papers for all 5000+ structures and compiled binding data for 1375 (26%) of the protein-ligand complexes. The binding-affinity data ranges 13 orders of magnitude. This is the largest collection of binding data reported to date in the literature. We have also addressed the issue of redundancy in the data. To create a nonredundant dataset, one protein from each of the 1780 protein families was chosen as a representative. Representatives were chosen by tightest binding, best resolution, etc. For the 1780 "best" complexes that comprise the nonredundant version of Binding MOAD, 475 (27%) have binding data. This significant collection of protein-ligand complexes will be very useful in elucidating the biophysical patterns of molecular recognition and enzymatic regulation. The complexes with binding-affinity data will help in the development of improved scoring functions and structure-based drug discovery techniques. The dataset can be accessed at http://www.BindingMOAD.org.

Biophysics↗

Automated interpretation of mass spectra of complex mixtures by matching of isotope peak distributions.

Mass spectrometry is now firmly established as a powerful technique for the identification and characterization of proteins when used in conjunction with sequence databases. Various approaches involving stable-isotope labeling have been developed for quantitative comparisons between paired samples in proteomic expression analysis by mass spectrometry. However, interpretation of such mass spectra is far from being fully automated, mainly due to the difficulty of analyzing complex patterns resulting from the overlap of multiple peaks arising from the assortment of natural isotopes. In order to facilitate the interpretation of a complex mass spectrum of such a mixture, such as an MS spectrum of a stable-isotope-enriched ion species, we report on the development of a software application, 'Matching' (web accessible), that enables the automatic matching of theoretical isotope envelopes to multiple ion peaks in a raw spectrum. It is particularly useful for resolving the relative abundances of narrow-split paired peaks caused by enrichment with a stable isotope, such as 18O, 13C, 2H, or 15N.

Algorithms↗