Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

The rat brain hippocampus proteome.

The hippocampus is crucial in memory storage and retrieval and plays an important role in stress response. In humans, the CA1 area of hippocampus is one of the first brain areas to display pathology in Alzheimer's disease. A comprehensive analysis of the hippocampus proteome has not been accomplished yet. We applied proteomics technologies to construct a two-dimensional database for rat brain hippocampus proteins. Hippocampus samples from eight months old animals were analyzed by two-dimensional electrophoresis and the proteins were identified by matrix-assisted laser desorption ionization time-of-flight mass spectrometry. The database comprises 148 different gene products, which are in the majority enzymes, structural proteins and heat shock proteins. It also includes 39 neuron specific gene products. The database may be useful in animal model studies of neurological disorders.

Animals↗

GFSWeb: a web tool for genome-based identification of proteins from mass spectrometric samples.

The interpretation of mass spectrometry data for protein identification has become a vital component of proteomics research. However, since most existing software tools rely on protein databases, their success is limited, especially as the pace of annotation efforts fails to keep pace with sequencing. We present a publicly available, web-based version of a software tool that maps peptide mass fingerprint data directly to their genomic origin, allowing for genome-based, annotation-independent protein identification.

Databases, Protein↗

ICDS database: interrupted CoDing sequences in prokaryotic genomes.

Unrecognized frameshifts, in-frame stop codons and sequencing errors lead to Interrupted CoDing Sequence (ICDS) that can seriously affect all subsequent steps of functional characterization, from in silico analysis to high-throughput proteomic projects. Here, we describe the Interrupted CoDing Sequence database containing ICDS detected by a similarity-based approach in 80 complete prokaryotic genomes. ICDS can be retrieved by species browsing or similarity searches via a web interface (http://www-bio3d-igbmc.u-strasbg.fr/ICDS/). The definition of each interrupted gene is provided as well as the ICDS genomic localization with the surrounding sequence. Furthermore, to facilitate the experimental characterization of ICDS, we propose optimized primers for re-sequencing purposes. The database will be regularly updated with additional data from ongoing sequenced genomes. Our strategy has been validated by three independent tests: (i) ICDS prediction on a benchmark of artificially created frameshifts, (ii) comparison of predicted ICDS and results obtained from the comparison of the two genomic sequences of Bacillus licheniformis strain ATCC 14580 and (iii) re-sequencing of 25 predicted ICDS of the recently sequenced genome of Mycobacterium smegmatis. This allows us to estimate the specificity and sensitivity (95 and 82%, respectively) of our program and the efficiency of primer determination.

Bacillus↗

Study of human laryngeal muscle protein using two-dimensional electrophoresis and mass spectrometry.

Proteomic analysis was performed to construct a protein database for human laryngeal muscle. Thyroarytenoid (TA) muscle specimens were obtained from six post mortem cases within 24 h of death. Isoelectric focusing was performed by using immobilized pH gradient strips followed by 12% sodium dodecyl sulfate-polyacrylamide gel electrophoresis. Silver stained gels were then analyzed using PDQuest software to locate, quantify and match spots. Proteins were identified by matrix-assisted laser desorption/ionization-mass spectrometry on the basis of peptide mass fingerprinting following in-gel digestion with trypsin. Comparison of protein distribution between broad and narrow pH range gels demonstrated that 75% of all protein spots from human TA muscle were located within the pH range 5-8, and between mass 15-120 kDa. Based on peptide mass fingerprinting, 75 proteins were identified and classified into six functional groups. These include membrane proteins (8.5%), cytoskeletal and myofibrillar proteins (14.6%), energy production proteins (28%), proteins associated with stress responses (8.5%), and protein associated with transcription regulation (10.9%). Approximately one-third (29%) were categorized as "other proteins". This data provides an initial reference map for comparative studies of protein expression in human and laryngeal muscle. Further development of this database will provide a valuable resource for molecular analysis of normal and pathologic conditions affecting human striated muscle.

Databases as Topic↗

AgBase: a functional genomics resource for agriculture.

BACKGROUND: Many agricultural species and their pathogens have sequenced genomes and more are in progress. Agricultural species provide food, fiber, xenotransplant tissues, biopharmaceuticals and biomedical models. Moreover, many agricultural microorganisms are human zoonoses. However, systems biology from functional genomics data is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation and agricultural research communities are smaller with limited funding compared to many model organism communities. DESCRIPTION: To facilitate systems biology in these traditionally agricultural species we have established "AgBase", a curated, web-accessible, public resource http://www.agbase.msstate.edu for structural and functional annotation of agricultural genomes. The AgBase database includes a suite of computational tools to use GO annotations. We use standardized nomenclature following the Human Genome Organization Gene Nomenclature guidelines and are currently functionally annotating chicken, cow and sheep gene products using the Gene Ontology (GO). The computational tools we have developed accept and batch process data derived from different public databases (with different accession codes), return all existing GO annotations, provide a list of products without GO annotation, identify potential orthologs, model functional genomics data using GO and assist proteomics analysis of ESTs and EST assemblies. Our journal database helps prevent redundant manual GO curation. We encourage and publicly acknowledge GO annotations from researchers and provide a service for researchers interested in GO and analysis of functional genomics data. CONCLUSION: The AgBase database is the first database dedicated to functional genomics and systems biology analysis for agriculturally important species and their pathogens. We use experimental data to improve structural annotation of genomes and to functionally characterize gene products. AgBase is also directly relevant for researchers in fields as diverse as agricultural production, cancer biology, biopharmaceuticals, human health and evolutionary biology. Moreover, the experimental methods and bioinformatics tools we provide are widely applicable to many other species including model organisms.

Agriculture↗

Effect of 2MEGA labeling on membrane proteome analysis using LC-ESI QTOF MS.

One of the challenges associated with large-scale proteome analysis using tandem mass spectrometry (MS/MS) and automated database searching is to reduce the number of false positive identifications without sacrificing the number of true positives found. In this work, a systematic investigation of the effect of 2MEGA labeling (N-terminal dimethylation after lysine guanidination) on the proteome analysis of a membrane fraction of an Escherichia coli cell extract by 2-dimensional liquid chromatography MS/MS is presented. By a large-scale comparison of MS/MS spectra of native peptides with those from the 2MEGA-labeled peptides, the labeled peptides were found to undergo facile fragmentation with enhanced a1 or a1-related (a(1)-17 and a(1)-45) ions derived from all N-terminal amino acids in the MS/MS spectra; these ions are usually difficult to detect in the MS/MS spectra of nonderivatized peptides. The 2MEGA labeling alleviated the biased detection of arginine-terminated peptides that is often observed in MALDI and ESI MS experiments. 2MEGA labeling was found not only to increase the number of peptides and proteins identified but also to generate enhanced a1 or a1-related ions as a constraint to reduce the number of false positive identifications. In total, 640 proteins were identified from the E. coli membrane fraction, with each protein identified based on peptide mass and sequence match of one or more peptides using MASCOT database search algorithm from the MS/MS spectra generated by a quadrupole time-of-flight mass spectrometer. Among them, the subcellular locations of 336 proteins are presently known, including 258 membrane and membrane-associated proteins (76.8%). Among the classified proteins, there was a dramatic increase in the total number of integral membrane proteins identified in the 2MEGA-labeled sample (153 proteins) versus the unlabeled sample (77 proteins).

Amino Acid Sequence↗

Review: prediction of in vivo fates of proteins in the era of genomics and proteomics.

Even after a nascent protein emerges from the ribosome, its fate is still controlled by its own amino acid sequence information. Namely, it may be co-/posttranslationally modified (e.g., phosphorylated, N-/O-glycosylated, and lipidated); it may be inserted into the membrane, translocated to an organelle, or secreted to the outside milieu; it may be processed for maturation or selective degradation; finally, its fragment may be presented on the cell surface as an antigen. Here, prediction methods of such protein fates from their amino acid sequences are reviewed. In many cases, artificial neural network techniques have been effectively used. The prediction of in vivo fates of proteins will be useful for characterizing newly identified candidate genes in a genome or for interpreting multiple spots in proteome analyses.

Animals↗

Methods for peptide identification by spectral comparison.

BACKGROUND: Tandem mass spectrometry followed by database search is currently the predominant technology for peptide sequencing in shotgun proteomics experiments. Most methods compare experimentally observed spectra to the theoretical spectra predicted from the sequences in protein databases. There is a growing interest, however, in comparing unknown experimental spectra to a library of previously identified spectra. This approach has the advantage of taking into account instrument-dependent factors and peptide-specific differences in fragmentation probabilities. It is also computationally more efficient for high-throughput proteomics studies. RESULTS: This paper investigates computational issues related to this spectral comparison approach. Different methods have been empirically evaluated over several large sets of spectra. First, we illustrate that the peak intensities follow a Poisson distribution. This implies that applying a square root transform will optimally stabilize the peak intensity variance. Our results show that the square root did indeed outperform other transforms, resulting in improved accuracy of spectral matching. Second, different measures of spectral similarity were compared, and the results illustrated that the correlation coefficient was most robust. Finally, we examine how to assemble multiple spectra associated with the same peptide to generate a synthetic reference spectrum. Ensemble averaging is shown to provide the best combination of accuracy and efficiency. CONCLUSION: Our results demonstrate that when combined, these methods can boost the sensitivity and specificity of spectral comparison. Therefore they are capable of enhancing and complementing existing tools for consistent and accurate peptide identification.

Journal Article↗

Peptide and protein identification by matrix-assisted laser desorption ionization (MALDI) and MALDI-post-source decay time-of-flight mass spectrometry.

The potential of matrix-assisted laser desorption ionization (MALDI) and MALDI-post-source decay (PSD) time-of-flight mass spectrometry for the characterization of peptides and proteins is discussed. Recent instrumental developments provide for levels of sensitivity and accuracy that make these techniques major analytical tools for proteome analysis. New software developments employing protein database searches have greatly enhanced the fields of application of MALDI-PSD. Peptides and proteins can be easily identified even if only a partial sequence information is determined. Derivatization procedures have been optimized for MALDI-PSD to increase the structural information and to obtain a complete peptide sequence even in critical cases. They are fast, simple and can be performed on target. MALDI-PSD is also a very powerful tool to characterize or elucidate post-translational or chemically induced modifications. In association with database searches, proteins issued from electrophoretic gels can be identified after specific enzymatic cleavages and peptide mapping.

Amino Acid Sequence↗

Protein family classification and functional annotation.

With the accelerated accumulation of genomic sequence data, there is a pressing need to develop computational methods and advanced bioinformatics infrastructure for reliable and large-scale protein annotation and biological knowledge discovery. The Protein Information Resource (PIR) provides an integrated public resource of protein informatics to support genomic and proteomic research. PIR produces the Protein Sequence Database of functionally annotated protein sequences. The annotation problems are addressed by a classification-driven and rule-based method with evidence attribution, coupled with an integrated knowledge base system being developed. The approach allows sensitive identification, consistent and rich annotation, and systematic detection of annotation errors, as well as distinction of experimentally verified and computationally predicted features. The knowledge base consists of two new databases, sequence analysis tools, and graphical interfaces. PIR-NREF, a non-redundant reference database, provides a timely and comprehensive collection of all protein sequences, totaling more than 1,000,000 entries. iProClass, an integrated database of protein family, function, and structure information, provides extensive value-added features for about 830,000 proteins with rich links to over 50 molecular databases. This paper describes our approach to protein functional annotation with case studies and examines common identification errors. It also illustrates that data integration in PIR supports exploration of protein relationships and may reveal protein functional associations beyond sequence homology.

Amino Acid Motifs↗

Role of biomarkers in monitoring exposures to chemicals: present position, future prospects.

Biomarkers are becoming increasingly important in toxicology and human health. Many research groups are carrying out studies to develop biomarkers of exposure to chemicals and apply these for human monitoring. There is considerable interest in the use and application of biomarkers to identify the nature and amounts of chemical exposures in occupational and environmental situations. Major research goals are to develop and validate biomarkers that reflect specific exposures and permit the prediction of the risk of disease in individuals and groups. One important objective is to prevent human cancer. This review presents a commentary and consensus views about the major developments on biomarkers for monitoring human exposure to chemicals. A particular emphasis is on monitoring exposures to carcinogens. Significant developments in the areas of new and existing biomarkers, analytical methodologies, validation studies and field trials together with auditing and quality assessment of data are discussed. New developments in the relatively young field of toxicogenomics possibly leading to the identification of individual susceptibility to both cancer and non-cancer endpoints are also considered. The construction and development of reliable databases that integrate information from genomic and proteomic research programmes should offer a promising future for the application of these technologies in the prediction of risks and prevention of diseases related to chemical exposures. Currently adducts of chemicals with macromolecules are important and useful biomarkers especially for certain individual chemicals where there are incidences of occupational exposure. For monitoring exposure to genotoxic compounds protein adducts, such as those formed with haemoglobin, are considered effective biomarkers for determining individual exposure doses of reactive chemicals. For other organic chemicals, the excreted urinary metabolites can also give a useful and complementary indication of exposure for acute exposures. These methods have revealed 'backgrounds' in people not knowingly exposed to chemicals and the sources and significance of these need to be determined, particularly in the context of their contribution to background health risks.

Biomarkers↗

Machine learning approaches for the prediction of signal peptides and other protein sorting signals.

Prediction of protein sorting signals from the sequence of amino acids has great importance in the field of proteomics today. Recently, the growth of protein databases, combined with machine learning approaches, such as neural networks and hidden Markov models, have made it possible to achieve a level of reliability where practical use in, for example automatic database annotation is feasible. In this review, we concentrate on the present status and future perspectives of SignalP, our neural network-based method for prediction of the most well-known sorting signal: the secretory signal peptide. We discuss the problems associated with the use of SignalP on genomic sequences, showing that signal peptide prediction will improve further if integrated with predictions of start codons and transmembrane helices. As a step towards this goal, a hidden Markov model version of SignalP has been developed, making it possible to discriminate between cleaved signal peptides and uncleaved signal anchors. Furthermore, we show how SignalP can be used to characterize putative signal peptides from an archaeon, Methanococcus jannaschii. Finally, we briefly review a few methods for predicting other protein sorting signals and discuss the future of protein sorting prediction in general.

Algorithms↗

EST mining and functional expression assays identify extracellular effector proteins from the plant pathogen Phytophthora.

Plant pathogenic microbes have the remarkable ability to manipulate biochemical, physiological, and morphological processes in their host plants. These manipulations are achieved through a diverse array of effector molecules that can either promote infection or trigger defense responses. We describe a general functional genomics approach aimed at identifying extracellular effector proteins from plant pathogenic microorganisms by combining data mining of expressed sequence tags (ESTs) with virus-based high-throughput functional expression assays in plants. PexFinder, an algorithm for automated identification of extracellular proteins from EST data sets, was developed and applied to 2147 ESTs from the oomycete plant pathogen Phytophthora infestans. The program identified 261 ESTs (12.2%) corresponding to a set of 142 nonredundant Pex (Phytophthora extracellular protein) cDNAs. Of these, 78 (55%) Pex cDNAs were novel with no significant matches in public databases. Validation of PexFinder was performed using proteomic analysis of secreted protein of P. infestans. To identify which of the Pex cDNAs encode effector proteins that manipulate plant processes, high-throughput functional expression assays in plants were performed on 63 of the identified cDNAs using an Agrobacterium tumefaciens binary vector carrying the potato virus X (PVX) genome. This led to the discovery of two novel necrosis-inducing cDNAs, crn1 and crn2, encoding extracellular proteins that belong to a large and complex protein family in Phytophthora. Further characterization of the crn genes indicated that they are both expressed in P. infestans during colonization of the host plant tomato and that crn2 induced defense-response genes in tomato. Our results indicate that combining data mining using PexFinder with PVX-based functional assays can facilitate the discovery of novel pathogen effector proteins. In principle, this strategy can be applied to a variety of eukaryotic plant pathogens, including oomycetes, fungi, and nematodes.

Algal Proteins↗

Gene expression analyzed by high-resolution state array analysis and quantitative proteomics: response of yeast to mating pheromone.

The transcriptome provides the database from which a cell assembles its collection of proteins. Translation of individual mRNA species into their encoded proteins is regulated, producing discrepancies between mRNA and protein levels. Using a new modeling approach to data analysis, a striking diversity is revealed in association of the transcriptome with the translational machinery. Each mRNA has its own pattern of ribosome loading, a circumstance that provides an extraordinary dynamic range of regulation, above and beyond actual transcript levels. Using this approach together with quantitative proteomics, we explored the immediate changes in gene expression in response to activation of a mitogen-activated protein kinase pathway in yeast by mating pheromone. Interestingly, in 26% of those transcripts where the predicted protein synthesis rate changed by at least 3-fold, more than half of these changes resulted from altered translational efficiencies. These observations underscore that analysis of transcript level, albeit extremely important, is insufficient by itself to describe completely the phenotypes of cells under different conditions.

Computational Biology↗

MultiProtIdent: identifying proteins using database search and protein-protein interactions.

Protein identification is important in proteomics. Proteomic analyses based on mass spectra (MS) constitute innovative ways to identify the components of protein complexes. Instruments can obtain the mass spectrum to an accuracy of 0.01 Da or better, but identification errors are inevitable. This study shows a novel tool, MultiProtIdent, which can identify proteins using additional information about protein-protein interactions and protein functional associations. Both single and multiple Peptide Mass Fingerprints (PMFs) are input to MultiProtIdent, which matches the PMFs to a theoretical peptide mass database. The relationships or interactions among proteins are considered to reduce false positives in PMF matching. Experiments to identify protein complexes reveal that MultiProtIdent is highly promising. The website associated with this study is http://dbms104.csie.ncu.edu.tw/.

Algorithms↗

SPS' Digest: the Swiss Proteomics Society selection of proteomics articles.

Despite the consolidation of the specialized proteomics literature around a few established journals, such as Proteomics, Molecular and Cellular Proteomics, and the Journal of Proteome Research, a lot of information is still spread in many different publications from different fields, such as analytical sciences, MS, bioinformatics, etc. The purpose of SPS' Digest is to gather a selection of proteomics articles, to categorize them, and to make the list available on a periodic basis through a web page and email alerts.

Animals↗

The HUPO Plasma Proteome Project: a report from the Munich congress.

The Human Proteome Organization has several major collaborative research initiatives, including the Plasma Proteome Project. A major feature of the HUPO World Congress in Munich in August 2005 was the release of the special issue of PROTEOMICS with 28 articles from the pilot phase of the Plasma Proteome Project. An open Workshop and a presentation in the closing plenary session of the congress focused on next phases for the Plasma Proteome Project.

Antibodies↗