Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Identifying biological themes within lists of genes with EASE.

EASE is a customizable software application for rapid biological interpretation of gene lists that result from the analysis of microarray, proteomics, SAGE and other high-throughput genomic data. The biological themes returned by EASE recapitulate manually determined themes in previously published gene lists and are robust to varying methods of normalization, intensity calculation and statistical selection of genes. EASE is a powerful tool for rapidly converting the results of functional genomics studies from 'genes' to 'themes'.

Computational Biology↗

Identification of human whole saliva protein components using proteomics.

The determination of salivary biomarkers as a means of monitoring general health and for the early diagnosis of disease is of increasing interest in clinical research. Based on the linkage between salivary proteins and systemic diseases, the aim of this work was the identification of saliva proteins using proteomics. Salivary proteins were separated using two-dimensional (2-D) gel electrophoresis over a pH range between 3-10, digested, and then analyzed by matrix assisted laser desorption/ionization-time of flight (MALDI-TOF)-TOF mass spectrometry (MS) and tandem mass spectrometry (MS/MS). Proteins were identified using automated MS and MS/MS data acquisition. The resulting data were searched against a protein database using an internal Mascot search routine. Ninety spots give identifications with high statistical reliability. Of the identified proteins, 11 were separated and identified in saliva for the first time using proteomics tools. Moreover, three proteins that have not been previously identified in saliva, PLUNC, cystatin A, and cystatin B were identified.

Cystatin B↗

An XML standard for the dissemination of annotated 2D gel electrophoresis data complemented with mass spectrometry results.

BACKGROUND: Many proteomics initiatives require a seamless bioinformatics integration of a range of analytical steps between sample collection and systems modeling immediately assessable to the participants involved in the process. Proteomics profiling by 2D gel electrophoresis to the putative identification of differentially expressed proteins by comparison of mass spectrometry results with reference databases, includes many components of sample processing, not just analysis and interpretation, are regularly revisited and updated. In order for such updates and dissemination of data, a suitable data structure is needed. However, there are no such data structures currently available for the storing of data for multiple gels generated through a single proteomic experiments in a single XML file. This paper proposes a data structure based on XML standards to fill the void that exists between data generated by proteomics experiments and storing of data. RESULTS: In order to address the resulting procedural fluidity we have adopted and implemented a data model centered on the concept of annotated gel (AG) as the format for delivery and management of 2D Gel electrophoresis results. An eXtensible Markup Language (XML) schema is proposed to manage, analyze and disseminate annotated 2D Gel electrophoresis results. The structure of AG objects is formally represented using XML, resulting in the definition of the AGML syntax presented here. CONCLUSION: The proposed schema accommodates data on the electrophoresis results as well as the mass-spectrometry analysis of selected gel spots. A web-based software library is being developed to handle data storage, analysis and graphic representation. Computational tools described will be made available at http://bioinformatics.musc.edu/agml. Our development of AGML provides a simple data structure for storing 2D gel electrophoresis data.

Computational Biology↗

Proteomics-grade de novo sequencing approach.

The conventional approach in modern proteomics to identify proteins from limited information provided by molecular and fragment masses of their enzymatic degradation products carries an inherent risk of both false positive and false negative identifications. For reliable identification of even known proteins, complete de novo sequencing of their peptides is desired. The main problems of conventional sequencing based on tandem mass spectrometry are incomplete backbone fragmentation and the frequent overlap of fragment masses. In this work, the first proteomics-grade de novo approach is presented, where the above problems are alleviated by the use of complementary fragmentation techniques CAD and ECD. Implementation of a high-current, large-area dispenser cathode as a source of low-energy electrons provided efficient ECD of doubly charged peptides, the most abundant species (65-80%), in a typical trypsin-based proteomics experiment. A new linear de novo algorithm is developed combining efficiency and speed, processing on a conventional 3 GHz PC, 1000 MS/MS data sets in 60 s. More than 6% of all MS/MS data for doubly charged peptides yielded complete sequences, and another 13% gave nearly complete sequences with a maximum gap of two amino acid residues. These figures are comparable with the typical success rates (5-15%) of database identification. For peptides reliably found in the database (Mowse score > or = 34), the agreement with de novo-derived full sequences was >95%. Full sequences were derived in 67% of the cases when full sequence information was present in MS/MS spectra. Thus the new de novo sequencing approach reached the same level of efficiency and reliability as conventional database-identification strategies.

Algorithms↗

Evaluating preparative isoelectric focusing of complex peptide mixtures for tandem mass spectrometry-based proteomics: a case study in profiling chromatin-enriched subcellular fractions in Saccharomyces cerevisiae.

We have evaluated the use of free-flow electrophoresis, an emerging separation method for preparative isoelectric focusing of complex peptide mixtures, as a tool for high-throughput tandem mass spectrometry-based proteomic analysis. In this study, we investigated the ability of free-flow electrophoresis to resolve and fractionate complex peptide mixtures and also the effectiveness of using peptide isoelectric point in conjunction with peptide match probability scoring in sequence database searching. As a model system for this study, we analyzed a chromatin-enriched fraction from the yeast Saccharomyces cerevisiae. This mixture was fractionated using preparative isoelectric focusing by free-flow electrophoresis, followed by online capillary liquid chromatography electrospray tandem mass spectrometry and sequence database searching. Our results demonstrate that (1) FFE effectively resolves and fractionates complex peptide mixtures on the basis of peptide isoelectric point and (2) the introduction of peptide pI is effective in minimizing both false positive and false negative sequence matches in sequence database searching of tandem mass spectrometry data.

Amino Acid Sequence↗

Data-dependent electron capture dissociation FT-ICR mass spectrometry for proteomic analyses.

Electron capture dissociation (ECD) offers many benefits for the analysis of peptides and proteins, and consequently shows great potential for the field of proteomics. Recent developments have reduced the time scale required for ECD to milliseconds resulting in the technique's compatibility with on-line separation techniques, e.g., HPLC. Here, we demonstrate incorporation of ECD into a high-throughput data-dependent LC-MS/MS approach for the analysis of proteomic samples. The approach is applied to analysis of the protein Fc-ROR2 isolated from chondrocytes and is the first example of LC-ECD-MS/MS of such a sample. Protein sequence coverage was 29%. Within that coverage, fifteen peptides were isolated and subjected to ECD. In most cases, the sequence tag generated by ECD was over 70% (in terms of the number of peptide backbone cleavages). The ECD data were searched against the nonredundant human NCBI database using the SEQUEST algorithm. Protein ROR2 was assigned, as was IgG (Fc domain). The results demonstrate the suitability of ECD as an integral technique in high-throughput proteomic strategies.

Algorithms↗

Proteomic analysis of antigens from Leishmania infantum promastigotes.

Leishmaniasis is a zoonotic disease caused by the species of the genus Leishmania, flagellated protozoa that multiply inside mammalian macrophages and are transmitted by the bite of the sandfly. The disease is widespread and due to the lack of fully effective treatment and vaccination the search for new drugs and immune targets is needed. Proteomics seems to be a suitable strategy because the annotated sequenced genome of L. major is available. Here, we present a high-resolution proteome for L. infantum promastigotes comprising of around 700 spots. Western blot with rabbit hyperimmune serum raised against L. infantum promastiogote extracts and further analysis by MALDI-TOF and MALDI-TOF/TOF MS allowed the identification of various relevant functional antigenic proteins. Major antigenic proteins were identified as propionil carboxilasa, ATPase beta subunit, transketolase, proteasome subunit, succinyl-diaminopimelate desuccinylase, a probable tubulin alpha chain, the full-size heat shock protein 70, and several proteins of unknown function. In addition, one enzyme from the ergosterol biosynthesis pathway (adrenodoxin reductase) and the structural paraflagellar rod protein 3 (PAR3) were found among non-antigenic proteins. This study corroborates the usefulness of proteomics in identifying new proteins with crucial biological functions in Leishmania parasites.

Animals↗

GARBAN II: an integrative framework for extracting biological information from proteomic and genomic data.

Genomic and proteomic analyses generate a massive amount of data that requires specific bioinformatic tools for its management and interpretation. GARBAN II, developed from the previous GARBAN platform, provides an integrated framework to simultaneously analyse and compare multiple datasets from DNA microarrays and proteomic studies. The general architecture, gene classification and comparison, and graphical representation have been redesigned to ensure a user-friendly feature and to improve the capabilities and efficiency of this system. Additionally, GARBAN II has been extended with new applications to display networks of coexpressed genes and to integrate access to BioRag and MotifScanner so as to facilitate the holistic analysis of users' data.

Animals↗

Understanding protein trafficking in plant cells through proteomics.

The functions of approximately one-third of the proteins encoded by the Arabidopsis thaliana genome are completely unknown. Moreover, many annotations of the remainder of the genome supply tentative functions, at best. Knowing the ultimate localization of these proteins, as well as the pathways used for getting there, may provide clues as to their functions. The putative localization of most proteins currently relies on in silico-based bioinformatics approaches, which, unfortunately, often result in erroneous predictions. Emerging proteomics techniques coupled with other systems biology approaches now provide researchers with a plethora of methods for elucidating the final location of these proteins on a large scale, as well as the ability to dissect protein-sorting pathways in plants.

Databases, Protein↗

A versatile method for deciphering plant membrane proteomes.

Proteomics is a very powerful approach to link the information contained in sequenced genomes, like that of Arabidopsis, to the functional knowledge provided by studies of plant cell compartments. This article summarizes the different steps of a versatile strategy that has been developed to decipher plant membrane proteomes. Initiated with envelope membranes from spinach chloroplasts, this strategy has been adapted to thylakoids, and further extended to a series of membranes from the model plant Arabidopsis: chloroplast envelope membranes, plasma membrane, and mitochondrial membranes. The first step is the preparation of highly purified membrane fractions from plant tissues. The second step in the strategy is the fractionation of membrane proteins on the basis of their physico-chemical properties. Chloroform/methanol extraction and washing of membranes with NaOH, NaCl or any other agent led to the simplification of the protein content of the fraction to be analysed. The next step is the genuine proteomic step, i.e. the separation of proteins by 1D-gel electrophoresis followed by in-gel proteolytic digestion of the polypeptides, analysis of the proteolytic peptides using mass spectrometry, and protein identification by searching through databases. The last step is the validation of the procedure by checking the subcellular location. The results obtained by using this strategy demonstrate that a combination of different proteomics approaches, together with bioinformatics, indeed provide a better understanding of the biochemical machinery of the different plant membranes at the molecular level.

Alkalies↗

The serine carboxypeptidase like gene family of rice (Oryza sativa L. ssp. japonica).

Serine carboxypeptidases (SCPs) comprise a large family of protein hydrolyzing enzymes and have roles ranging from protein turnover and C-terminal processing to wound responses and xenobiotic metabolism. The proteins can be classified into three groups, namely carboxypeptidase I, II and III, based on their coding protein sequences and the fact that each family is characterized by a central catalytic domain of unique topology designated as the "alpha/beta hydrolase fold". The available SCP protein sequences have been utilized as datasets to build a HMM (hidden Markov model) profile, which is used to search the rice (Oryza sativa L. ssp. japonica) proteome. A total of 71 SCP and serine carboxypeptidase-like (SCPL) protein-coding genes exist in rice. The intron-exon structure, chromosome localization, expression and characteristics of encoded protein sequences of the 71 putative genes are reviewed.

Amino Acid Sequence↗

Proteomic analysis of proteins expressed by Helicobacter pylori under oxidative stress.

Helicobacter pylori is a spiral, slow growing gram-negative microaerophilic bacterium. It has been shown to be the etiological agent of gastroduodenal diseases, such as chronic gastritis, gastric and duodenal ulcers, and gastric cancer. To address the influence of oxidative stress and its underlying mechanisms, we have compared proliferation, urease activity and protein expression profile of H. pylori incubated under normal microaerophilic (5% O2) and aerobic stress (20% O2) conditions. Oxidative-stress cells displayed coccoid morphology and time-dependent decrease in proliferation. The urease activity was completely abrogated after 32 h. We have further compared the protein expression profiles of H. pylori under normal growing and oxidative-stress conditions by a global proteomic analysis, which includes high-resolution 2-DE followed by MALDI-TOF-MS and bioinformatic databases search/peptide-mass comparison. The results revealed that more than ten proteins were differentially expressed under oxidative stress. Most notably, the protein expression levels of urease accessory protein E (UreE, an essential metallochaperone for urease activity) and alkylhydroperoxide reductase (AhpC) with antioxidant potential are greatly decreased under stress conditions. Measurements of messenger RNA transcription level by performing RT-PCR on total mRNA also confirmed that gene expressions for these two proteins are consistently repressed under oxygen tension. These changes form a firm basis to account for the loss of urease activity and anti-oxidative ability of H. pylori after long-term exposure to reactive oxygen. Conceivably, UreE and AhpC may thus be listed as potential targets for the development of therapeutic drugs against H. pylori.

Animals↗

Current chemical tagging strategies for proteome analysis by mass spectrometry.

Proteomics, the analysis of the protein complement of a cell or an organism, has grown rapidly as a subdiscipline of the life sciences. Mass spectrometry (MS) is one of the central detection techniques in proteome analysis, yet it has to rely on prior sample preparation steps that reduce the enormous complexity of the protein mixtures obtained from biological systems. For that reason, a number of so-called tagging (or labeling) strategies have been developed that target specific amino acid residues or post-translational modifications, enabling the enrichment of subfractions via affinity clean-up, resulting in the identification of an ever increasing number of proteins. In addition, the attachment of stable-isotope-labeled tags now allows the relative quantitation of protein levels of two samples, e.g. those representing different cell states, which is of great significance for drug discovery and molecular biology. Finally, tagging schemes also serve to facilitate interpretation of MS/MS spectra, therefore assisting in de novo elucidation of protein sequences and automated database searching. This review summarizes the different application fields for tagging strategies for today's MS-based proteome analysis. Advantages and drawbacks of the numerous strategies that have appeared in the literature in the last years are highlighted, and an outlook on emerging tagging techniques is given.

Amino Acids↗

PARIS: a proteomic analysis and resources indexation system.

UNLABELLED: We developed a system for managing data from two-dimensional electrophoresis-based proteomic experiments. Named PARIS, the system stores gel image and information about experiments and analysis procedures, allows the user to search and navigate in genomic and proteomic data, supports visual verification and validation of the analysis results, and provides tools for cross multi-experiment and multi-experimenter data validation and exploration. AVAILABILITY: The software is freely available from http://www.inra.fr/bia/J/imaste/Projets/PARIS/index.html

Abstracting and Indexing↗

Global analysis of the Ralstonia metallidurans proteome: prelude for the large-scale study of heavy metal response.

A proteome map of Ralstonia metallidurans strain CH34 was constructed using two-dimensional (2-D) gel electrophoresis in combination with automated Edman degradation and mass spectrometry (MS). R. metallidurans CH34 is the type-strain of a family of highly related strains characterized by their multiple resistance to millimolar amounts of heavy metals, conferred by two large plasmids. The protein content of this bacterium grown in minimal medium was separated by 2-D gel electrophoresis using various pH gradients. Protein identification was carried out via N-terminal amino acid sequencing, matrix assisted laser desorption/ionisation-time of flight-mass spectrometry (MALDI-TOF-MS) and tandem MS. So far, 224 different proteins were characterized from 352 protein spots. Although the proteome map is still not complete, one could appraise the importance of proteomics for genome analyses through (i). the identification of previously undetected open reading frames, (ii). the identification of proteins not encoded by the already sequenced genome fragments, (iii). the characterization of protein-encoding genes spanning two different contigs, enabling their merging, and (iv). the precise delineation of the N-terminus of several proteins. Finally, this map will prove a useful tool in the identification of proteins differentially expressed in the presence of different heavy metals.

Databases, Protein↗

Exploring the range of protein flexibility, from a structural proteomics perspective.

Changes in protein conformation play a vital role in biochemical processes, from biopolymer synthesis to membrane transport. Initial systematizations of protein flexibility, in a database framework, concentrated on the movement of domains and linkers. Movements were described in terms of simple sliding and hinging mechanisms of individual secondary structural elements. Recently, the accelerated pace and sophistication of methods for structural characterization of proteins has allowed high-resolution studies of increasingly complex assemblies and conformational changes. New data emphasize a breadth of possible structural mechanisms, particularly the ability to drastically alter protein architecture and the native flexibility of many structures.

Pliability↗

The UCSC Known Genes.

The University of California Santa Cruz (UCSC) Known Genes dataset is constructed by a fully automated process, based on protein data from Swiss-Prot/TrEMBL (UniProt) and the associated mRNA data from Genbank. The detailed steps of this process are described. Extensive cross-references from this dataset to other genomic and proteomic data were constructed. For each known gene, a details page is provided containing rich information about the gene, together with extensive links to other relevant genomic, proteomic and pathway data. As of July 2005, the UCSC Known Genes are available for human, mouse and rat genomes. The Known Genes serves as a foundation to support several key programs: the Genome Browser, Proteome Browser, Gene Sorter and Table Browser offered at the UCSC website. All the associated data files and program source code are also available. They can be accessed at http://genome.ucsc.edu. The genomic coverage of UCSC Known Genes, RefSeq, Ensembl Genes, H-Invitational and CCDS is analyzed. Although UCSC Known Genes offers the highest genomic and CDS coverage among major human and mouse gene sets, more detailed analysis suggests all of them could be further improved.

Base Sequence↗