Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Project management system for structural and functional proteomics: Sesame.

A computing infrastructure (Sesame) has been designed to manage and link individual steps in complex projects. Sesame is being developed to support a large-scale structural proteomics pilot project. When complete, the system is expected to manage all steps from target selection to data-bank deposition and report writing. We report here on the design criteria of the Sesame system and on results demonstrating successful achievement of the basic goals of its architecture. The Sesame software package, which follows the client/server paradigm, consists of a framework, which supports secure interactions among the three tiers of the system (the client, server, and database tiers), and application modules that carry out specific tasks. The framework utilizes industry standards. The client tier is written in Java2 and can be accessed anywhere through the Internet. All the development on the server tier is also carried out in Java2 so as to accommodate a wide variety of computer platforms. The database tier employs a commercial database management system. Each Sesame application module consists of a simple user interface in the client tier, corresponding objects in the server tier, and relevant data stored in the centralized database. For security, access to stored data is controlled by access privileges. The system facilitates both local and remote collaborations. Because users interact with the system using Java Web Start or through a web browser, access is limited only by the availability of an Internet connection. We describe several Sesame modules that have been developed to the point where they are being utilized routinely to support steps involved in structural and functional proteomics. This software is available to parties interested in using it and assisting to guide its further development.

Database Management Systems↗

Functions of the tegument of schistosomes: clues from the proteome and lipidome.

The tegumental outer-surface of schistosomes is a unique double membrane structure that is of crucial importance for modulation of the host response and parasite survival. Although several tegumental proteins had been identified by classical biochemical approaches, knowledge on the entire molecular composition of the tegument was limited. The Schistosoma mansoni genome project, together with recently developed proteomic and lipidomic techniques, allowed studies on detailed characterisation of the proteins and lipids of the tegumental membranes. These studies identified tegumental proteins and lipids that confirm the function of the tegument in nutrient uptake and immune evasion. However, these studies also demonstrated that compared to the complete worm, the tegument is enriched in lipids that are absent in the host. The tegument is also enriched in proteins that share no sequence similarity to any sequence present in databases of species other than schistosomes. These results suggest that the unique tegumental structures comprise multiple unique components that are likely to fulfil yet unknown functions. The tegumental proteome and lipidome, therefore, imply that many unknown molecular mechanisms are employed by schistosomes to survive within their host.

Animals↗

A structure-based anatomy of the E.coli metabolome.

The Escherichia coli metabolome has been characterised using the two-dimensional structures of 745 metabolites, obtained from the EcoCyc and KEGG databases. Physicochemical properties of the metabolome have been calculated to provide an overview of this set of cognate ligands. A library of fragments commonly found among these molecules has been employed to reveal the main constituents of metabolites, and to assist a broad classification of the metabolome into biochemically relevant classes. Fragment-based fingerprints reveal the metabolome as a continuum in the two-dimensional structural space, where clusters of molecules sharing similar scaffolds can be identified, but are generally overlapping. Nucleotide, carbohydrate and amino acid-like molecules are the most prominent, but at high levels of similarity, a more detailed classification is possible. Classification schemes for the metabolome are a promising tool for understanding the chemical diversity of the metabolome. When used in conjunction with existing classifications of the proteome, they can help to elucidate the binding preferences and promiscuity of proteins and their cognate substrates.

Computational Biology↗

Generalized modeling of enzyme-ligand interactions using proteochemometrics and local protein substructures.

Modeling and understanding protein-ligand interactions is one of the most important goals in computational drug discovery. To this end, proteochemometrics uses structural and chemical descriptors from several proteins and several ligands to induce interaction-models. Here, we present a new and generalized approach in which proteins varying greatly in terms of sequence and structure are represented by a library of local substructures. Using linear regression and rule-based learning, we combine such local substructures with chemical descriptors from the ligands to model binding affinity for a training set of hydrolase and lyase enzymes. We evaluate the predictive performance of these models using cross validation and sets of unseen ligand with unknown three-dimensional structure. The models are shown to generalize by outperforming models using descriptors from only proteins or only ligands, or models using global structure similarities rather than local similarities. Thus, we demonstrate that this approach is capable of describing dependencies between local structural properties and ligands in otherwise dissimilar protein structures. These dependencies are often, but not always, associated with local substructures that are in contact with the ligands. Finally, we show that strongly bound enzyme-ligand complexes require the presence of particular local substructures, while weakly bound complexes may be described by the absence of certain properties. The results demonstrate that the alignment-independent approach using local substructures is capable of describing protein-ligand interaction for largely different proteins and hence opens up for proteochemometrics-analysis of the interaction-space of entire proteomes. Current approaches are limited to families of closely related proteins. families of closely related proteins.

Algorithms↗

Differential expression of proteins in response to ceramide-mediated stress signal in colon cancer cells by 2-D gel electrophoresis and MALDI-TOF-MS.

Comparative cancer cell proteome analysis is a strategy to study the implication of ceramides in the transmission of stress signals. To better understand the mechanisms by which ceramide regulate some physiological or pathological events and the response to the pharmacological treatment of cancer, we performed a differential analysis of the proteome of HCT-116 (human colon carcinoma) cells in response to these substances. We first established the first 2-dimensional map of the HCT-116 proteome. Then, HCT116 cell proteome treated or not with C6-ceramide have been compared using two-dimensional electrophoresis, matrix-assisted laser desorption/ionization-mass spectrometry and bioinformatic (genomic databases). 2-DE gel analysis revealed more than fourty proteins that were differentially expressed in control cells and cells treated with ceramide. Among them, we confirmed the differential expression of proteins involved in apoptosis and cell adhesion.

Apoptosis↗

AGML Central: web based gel proteomic infrastructure.

SUMMARY: AGML Central is a web-based open-source public infrastructure for dissemination of two-dimensional Gel Electrophoresis (2-DE) proteomics data in AGML format (Annotated Gel Markup Language). It includes a growing collection of converters from proprietary formats such as those produced by PDQUEST (BioRad), PHORETIX 2-D (Nonlinear Dynamics) and Melanie (GenBio SA). The resulting unifying AGML formatted entry, with or without the raw gel images, is optionally stored in a database for future reference. AGML Central was developed to provide a common platform for data dissemination and development of 2-DE data analysis tools. This resource responds to an increasing use of AGML for 2-DE public source data representation which requires automated tools for conversion from proprietary formats. Conversion and short-term storage is made publicly available, permanent storage requires prior registering. A JAVA applet visualizer was developed to visualize the AGML data with cross-reference links. In order to facilitate automated access a SOAP web service is also included in the AGML Central infrastructure. AVAILABILITY: http://bioinformatics.musc.edu/agmlcentral.

Database Management Systems↗

Defining the mandate of proteomics in the post-genomics era: workshop report.

Research in proteomics is the next step after genomics in understanding life processes at the molecular level. In the largest sense proteomics encompasses knowledge of the structure, function and expression of all proteins in the biochemical or biological contexts of all organisms. Since that is an impossible goal to achieve, at least in our lifetimes, it is appropriate to set more realistic, achievable goals for the field. Up to now, primarily for reasons of feasibility, scientists have tended to concentrate on accumulating information about the nature of proteins and their absolute and relative levels of expression in cells (the primary tools for this have been 2D gel electrophoresis and mass spectrometry). Although these data have been useful and will continue to be so, the information inherent in the broader definition of proteomics must also be obtained if the true promise of the growing field is to be realized. Acquiring this knowledge is the challenge for researchers in proteomics and the means to support these endeavors need to be provided. An attempt has been made to present the major issues confronting the field of proteomics and two clear messages come through in this report. The first is that the mandate of proteomics is and should be much broader than is frequently recognized. The second is that proteomics is much more complicated than sequencing genomes. This will require new technologies but it is highly likely that many of these will be developed. Looking back 10 to 20 years from now, the question is: Will we have done the job wisely or wastefully? This report summarizes the presentations made at a symposium at the National Academy of Sciences on February 25, 2002.

Computational Biology↗

Use of a proteome strategy for tagging proteins present at the plasma membrane.

A plasma membrane (PM) fraction was purified from Arabidopsis thaliana using a standard procedure and analyzed by two-dimensional (2D) gel electrophoresis. The proteins were classified according to their relative abundance in PM or cell membrane supernatant fractions. Eighty-two of the 700 spots detected on the PM 2D gels were microsequenced. More than half showed sequence similarity to proteins of known function. Of these, all the spots in the PM-specific and PM-enriched fractions, together with half of the spots with similar abundance in PM fraction and supernatant, have previously been found at the PM, supporting the validity of this approach. Extrapolation from this analysis indicates that (i) approximately 550 polypeptides found at the PM could be resolved on 2D gels; (ii) that numerous proteins with multiple locations are found at the PM; and (iii) that approximately 80% of PM-specific spots correspond to proteins with unknown function. Among the later, half are represented by ESTs or cDNAs in databases. In this way, several unknown gene products were potentially localized to the PM. These data are discussed with respect to the efficiency of organelle proteome approaches to link systematically genomic data to genome expression. It is concluded that generalized proteomes can constitute a powerful resource, with future completion of Arabidopsis genome sequencing, for genome-wide exploration of plant function.

Amino Acid Sequence↗

In vitro and in silico processes to identify differentially expressed proteins.

We present an integrated proteomics platform designed for performing differential analyses. Since reproducible results are essential for comparative studies, we explain how we improved reproducibility at every step of our laboratory processes, e.g. by taking advantage of the powerful laboratory information management system we developed. The differential capacity of our platform is validated by detecting known markers in a real sample and by a spiking experiment. We introduce an innovative two-dimensional (2-D) plot for displaying identification results combined with chromatographic data. This 2-D plot is very convenient for detecting differential proteins. We also adapt standard multivariate statistical techniques to show that peptide identification scores can be used for reliable and sensitive differential studies. The interest of the protein separation approach we generally apply is justified by numerous statistics, complemented by a comparison with a simple shotgun analysis performed on a small volume sample. By introducing an automatic integration step after mass spectrometry data identification, we are able to search numerous databases systematically, including the human genome and expressed sequence tags. Finally, we explain how rigorous data processing can be combined with the work of human experts to set high quality standards, and hence obtain reliable (false positive < 0.35%) and nonredundant protein identifications.

Body Fluids↗

Visualization of comparative genomic analyses by BLAST score ratio.

BACKGROUND: The first microbial genome sequence, Haemophilus influenzae, was published in 1995. Since then, more than 400 microbial genome sequences have been completed or commenced. This massive influx of data provides the opportunity to obtain biological insights through comparative genomics. However few tools are available for this scale of comparative analysis. RESULTS: The BLAST Score Ratio (BSR) approach, implemented in a Perl script, classifies all putative peptides within three genomes using a measure of similarity based on the ratio of BLAST scores. The output of the BSR analysis enables global visualization of the degree of proteome similarity between all three genomes. Additional output enables the genomic synteny (conserved gene order) between each genome pair to be assessed. Furthermore, we extend this synteny analysis by overlaying BSR data as a color dimension, enabling visualization of the degree of similarity of the peptides being compared. CONCLUSIONS: Combining the degree of similarity, synteny and annotation will allow rapid identification of conserved genomic regions as well as a number of common genomic rearrangements such as insertions, deletions and inversions. The script and example visualizations are available at: http://www.microbialgenomics.org/BSR/.

Algorithms↗

High-throughput proteomics for alcohol research.

This report summarizes the proceedings of a satellite symposium of the 2003 Research Society on Alcoholism meeting held on June 21, 2003, in Fort Lauderdale, FL. The goal of this symposium, sponsored by the NIAAA, was to identify new proteomic directions in alcohol research that will (1) enable studies that focus on characterizing protein function, biochemical pathways, and networks to understand alcohol-related illnesses; (2) identify protein-protein interactions, posttranslational modifications, and subcellular localizations; (3) identify molecular targets for medication development; (4) develop biomarkers for susceptibility, dependence, consumption, and relapse, as well as alcohol-induced pathologies; and (5) develop high-throughput drug screens to test the efficacy of therapeutics that control alcohol-induced diseases. The purpose of the symposium was also to promote the application of high-throughput proteomic approaches, including isolation of membrane-bound proteins, in situ proteomics, large-scale two-dimensional separations, protein microarray platforms, mass spectrometry, matrix-assisted laser desorption/ionization, matrix-assisted laser desorption/ionization time-of-flight, liquid chromatography-tandem mass spectrometry, and isotope-coded affinity tags. In addition, the development of protein network maps by using new bioinformatics approaches for database mining was also discussed.

Alcohol Drinking↗

ASAP: the Alternative Splicing Annotation Project.

Recently, genomics analyses have demonstrated that alternative splicing is widespread in mammalian genomes (30-60% of genes reported to have multiple isoforms), and may be one of their most important mechanisms of functional regulation. However, by comparison with other genomics data such as genome annotation, SNPs, or gene expression, there exists relatively little database infrastructure for the study of alternative splicing. We have constructed an online database ASAP (the Alternative Splicing Annotation Project) for biologists to access and mine the enormous wealth of alternative splicing information coming from genomics and proteomics. ASAP is based on genome-wide analyses of alternative splicing in human (30 793 alternative splice relationships found) from detailed alignment of expressed sequences onto the genomic sequence. ASAP provides precise gene exon-intron structure, alternative splicing, tissue specificity of alternative splice forms, and protein isoform sequences resulting from alternative splicing. Moreover, it can help biologists design probe sequences for distinguishing specific mRNA isoforms. ASAP is intended to be a community resource for collaborative annotation of alternative splice forms, their regulation, and biological functions. The URL for ASAP is http://www.bioinformatics.ucla.edu/ASAP.

Alternative Splicing↗

Global proteome discovery using an online three-dimensional LC-MS/MS.

We have developed a proteomics technology featuring on-line three-dimensional liquid chromatography coupled to tandem mass spectrometry (3D LC-MS/MS). Using 3D LC-MS/MS, the yeast-soluble, urea-solubilized peripheral membrane and SDS-solubilized membrane protein samples collectively yielded 3019 unique yeast protein identifications with an average of 5.5 peptides per protein from the 6300-gene Saccharomyces Genome Database searched with SEQUEST. A single run of the urea-solubilized sample yielded 2255 unique protein identifications, suggesting high peak capacity and resolving power of 3D LC-MS/MS. After precipitation of SDS from the digested membrane protein sample, 3D LC-MS/MS allowed the analysis of membrane proteins. Among 1221 proteins containing two or more predicted transmembrane domains, 495 such proteins were identified. The improved yeast proteome data allowed the mapping of many metabolic pathways and functional categories. The 3D LC-MS/MS technology provides a suitable tool for global proteome discovery.

Amino Acid Sequence↗

The serum proteome of Equus caballus.

We constructed a reference two-dimensional protein map for horse (Equus caballus) serum. The serum proteins were separated by two-dimensional electrophoresis (2-DE); 29 different gene products were identified. Proteins represented by 25 spots/spot groups were identified by tandem nanoelectrospray mass spectrometry (MS), four by matrix-assisted laser desorption ionization time-of-flight (TOF) MS and one was sequenced by TOF-TOF technology. The identities of four proteins were deduced by similarity to the human plasma protein database. In selected cases, i.e. the immunoglobulins, immunoblotting with specific antibodies provided additional information about the respective proteins. Albumin was detected as the full-length protein and as fragments of various sizes. Spots representing products of different mass and charge were also detected for alpha1-antitrypsin, haptoglobin and transthyretin. Thus, despite the fact that the Equus caballus genome is incompletely characterized, we were able to identify almost all moderate to high abundance proteins stained in the serum 2-DE pattern.

Albumins↗

Searching for hypothetical proteins: theory and practice based upon original data and literature.

A large part of mammalian proteomes is represented by hypothetical proteins (HP), i.e. proteins predicted from nucleic acid sequences only and protein sequences with unknown function. Databases are far from being complete and errors are expected. The legion of HP is awaiting experiments to show their existence at the protein level and subsequent bioinformatic handling in order to assign proteins a tentative function is mandatory. Two-dimensional gel-electrophoresis with subsequent mass spectrometrical identification of protein spots is an appropriate tool to search for HP in the high-throughput mode. Spots are identified by MS or by MS/MS measurements (MALDI-TOF, MALDI-TOF-TOF) and subsequent software as e.g. Mascot or ProFound. In many cases proteins can thus be unambiguously identified and characterised; if this is not the case, de novo sequencing or Q-TOF analysis is warranted. If the protein is not identified, the sequence is being sent to databases for BLAST searches to determine identities/similarities or homologies to known proteins. If no significant identity to known structures is observed, the protein sequence is examined for the presence of functional domains (databases PROSITE, PRINTS, InterPro, ProDom, Pfam and SMART), subjected to searches for motifs (ELM) and finally protein-protein interaction databases (InterWeaver, STRING) are consulted or predictions from conformations are performed. We here provide information about hypothetical proteins in terms of protein chemical analysis, independent of antibody availability and specificity and bioinformatic handling to contribute to the extension/completion of protein databases and include original work on HP in the brain to illustrate the processes of HP identification and functional assignment.

Amino Acid Sequence↗

De novo peptide sequencing via tandem mass spectrometry.

Peptide sequencing via tandem mass spectrometry (MS/MS) is one of the most powerful tools in proteomics for identifying proteins. Because complete genome sequences are accumulating rapidly, the recent trend in interpretation of MS/MS spectra has been database search. However, de novo MS/MS spectral interpretation remains an open problem typically involving manual interpretation by expert mass spectrometrists. We have developed a new algorithm, SHERENGA, for de novo interpretation that automatically learns fragment ion types and intensity thresholds from a collection of test spectra generated from any type of mass spectrometer. The test data are used to construct optimal path scoring in the graph representations of MS/MS spectra. A ranked list of high scoring paths corresponds to potential peptide sequences. SHERENGA is most useful for interpreting sequences of peptides resulting from unknown proteins and for validating the results of database search algorithms in fully automated, high-throughput peptide sequencing.

Algorithms↗

BRIDGE: an interactive application for multi-omics data analysis, visualization and integration.

SUMMARY: BRIDGE is a Shiny-based application that provides an accessible, modular platform for individual and integrative multi-omics analysis. Using an independent SQLite database backend, it offers a local, private, and user-friendly environment that requires no prior computational expertise. The application supports proteomics, phospho-proteomics, and RNA-seq analyses through a comprehensive suite of visualization and analytical modules, together with an integrated multi-omics analysis pipeline. Built-in caching and asynchronous processing improve responsiveness, enabling efficient exploration, analysis, and visualization of multi-omics datasets on moderate hardware. AVAILABILITY AND IMPLEMENTATION: BRIDGE is implemented in R using Shiny and is freely available as a Docker container at https://ghcr.io/paulilab/bridge. A public demonstration server with example datasets is available at https://bridge.imp.ac.at. Code and datasets are also available at https://github.com/paulilab/BRIDGE and under DOI: https://doi.org/10.5281/zenodo.20215824.

Multiomics↗

Discovering motif pairs at interaction sites from protein sequences on a proteome-wide scale.

MOTIVATION: Protein-protein interaction, mediated by protein interaction sites, is intrinsic to many functional processes in the cell. In this paper, we propose a novel method to discover patterns in protein interaction sites. We observed from protein interaction networks that there exist a kind of significant substructures called interacting protein group pairs, which exhibit an all-versus-all interaction between the two protein-sets in such a pair. The full-interaction between the pair indicates a common interaction mechanism shared by the proteins in the pair, which can be referred as an interaction type. Motif pairs at the interaction sites of the protein group pairs can be used to represent such interaction type, with each motif derived from the sequences of a protein group by standard motif discovery algorithms. The systematic discovery of all pairs of interacting protein groups from large protein interaction networks is a computationally challenging problem. By a careful and sophisticated problem transformation, the problem is solved using efficient algorithms for mining frequent patterns, a problem extensively studied in data mining. RESULTS: We found 5349 pairs of interacting protein groups from a yeast interaction dataset. The expected value of sequence identity within the groups is only 7.48%, indicating non-homology within these protein groups. We derived 5343 motif pairs from these group pairs, represented in the form of blocks. Comparing our motifs with domains in the BLOCKS and PRINTS databases, we found that our blocks could be mapped to an average of 3.08 correlated blocks in these two databases. The mapped blocks occur 4221 out of total 6794 domains (protein groups) in these two databases. Comparing our motif pairs with iPfam consisting of 3045 interacting domain pairs derived from PDB, we found 47 matches occurring in 105 distinct PDB complexes. Comparing with another putative domain interaction database InterDom, we found 203 matches. AVAILABILITY: http://research.i2r.a-star.edu.sg/BindingMotifPairs/resources. SUPPLEMENTARY INFORMATION: http://research.i2r.a-star.edu.sg/BindingMotifPairs and Bioinformatics online.

Algorithms↗