Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

A wiring of the human nucleolus.

Recent proteomic efforts have created an extensive inventory of the human nucleolar proteome. However, approximately 30% of the identified proteins lack functional annotation. We present an approach of assigning function to uncharacterized nucleolar proteins by data integration coupled to a machine-learning method. By assembling protein complexes, we present a first draft of the human ribosome biogenesis pathway encompassing 74 proteins and hereby assign function to 49 previously uncharacterized proteins. Moreover, the functional diversity of the nucleolus is underlined by the identification of a number of protein complexes with functions beyond ribosome biogenesis. Finally, we were able to obtain experimental evidence of nucleolar localization of 11 proteins, which were predicted by our platform to be associates of nucleolar complexes. We believe other biological organelles or systems could be "wired" in a similar fashion, integrating different types of data with high-throughput proteomics, followed by a detailed biological analysis and experimental validation.

Artificial Intelligence↗

BRIGEP--the BRIDGE-based genome-transcriptome-proteome browser.

The growing amount of information resulting from the increasing number of publicly available genomes and experimental results thereof necessitates the development of comprehensive systems for data processing and analysis. In this paper, we describe the current state and latest developments of our BRIGEP bioinformatics software system consisting of three web-based applications: GenDB, EMMA and ProDB. These applications facilitate the processing and analysis of bacterial genome, transcriptome and proteome data and are actively used by numerous international groups. We are currently in the process of extensively interconnecting these applications. BRIGEP was developed in the Bioinformatics Resource Facility of the Center for Biotechnology at Bielefeld University and is freely available. A demo project with sample data and access to all three tools is available at https://www.cebitec.uni-bielefeld.de/groups/brf/software/brigep/. Code bundles for these and other tools developed in our group are accessible on our FTP server at ftp.cebitec.uni-bielefeld.de/pub/software/.

Bacterial Proteins↗

BLASTing proteomes, yielding phylogenies.

We develop a procedure called RiPE (Retrieval-induced Phylogeny Environment) that automatically performs an evolutionary analysis of a protein (sub)family, (i) by retrieving the relevant sequences via a homology search, (ii) by using the search report to construct the alignment using only homologous subsequences (taking into account their neighborhood with a low chance of homology), (iii) by realigning, and (iv) by generating phylogenetic trees based on the alignment. In a first implementation of our scheme, we start with the available proteome data of model organisms, perform a PSI-BLAST search, use MView to convert hits into a multiple alignment, and perform realignment and tree building. As a test case, we have investigated the human ABC transporters of the subfamily G, starting with the five known human ABCG transporters. Our method retrieved homologous sequences not previously analyzed, generating a tree that is more plausible and better supported than previously published trees. The RiPE 0.1 prototype is available at the RiPE website, http://ifg-izkf.uni-muenster.de/fuellen/RiPE/ripe.html.

ATP Binding Cassette Transporter, Subfamily G, Mem↗

Proteomics analysis of human cerebrospinal fluid.

Cerebrospinal fluid (CSF) is secreted from several different central nervous system (CNS) structures, and any changes in the CSF composition will accurately reflect pathological processes. Proteomics offers a comprehensive bird's eye view to analyze CSF proteins at a systems level. This paper reviews the variety of analytical methods that have been used for proteomics analysis of CSF, including sample preparation, two-dimensional liquid and gel electrophoresis, mass spectrometry, bioinformatics, and non-gel methods. The differentially expressed CSF proteins that have been identified by proteomics methods are discussed.

Cerebrospinal Fluid Proteins↗

[Proteomic analysis: why and how ?].

The proteome, first formalized in 1995, designs all the proteins expressed by the genome of a cell, tissu, or organ at a defined time. Proteomic analysis leads to a description of the regulation of gene expression by the study of proteins and of their post-translational modifications. Proteomic analysis is based on three technologies: 1) Two-dimensional electrophoresis allowing the separation of thousands of proteins from a single mixture; 2) mass spectrometry allowing the characterization of picoquantities of polypeptides and providing data on post-translational modifications; 3) Bioinformatic which is required for the quantification of protein level and for the constitution of databases of protein expression profiles. Complementing the methods of the genomics, the use of proteomic analysis is widely spreading in the fields of fundamental biology, biomedicine and pharmacology for the identification of new biological markers and therapeutic targets.

Animals↗

Protein profiling of the medicinal leech salivary gland secretion by proteomic analytical methods.

Protein diversity of the high molecular weight fraction (molecular mass > 500 daltons) of salivary grand secretion of the medicinal leech Hirudo medicinalis has been demonstrated using methods of proteomic analysis. One-dimensional (1D) electrophoresis revealed the presence of more than 60 bands corresponding to molecular masses ranging from 11 to 483 kD. 2D-electrophoresis revealed more than 100 specific protein spots differing in molecular masses and pI values. SELDI-mass spectrometry analysis using the ProteinChip. System based on chromatography surfaces of strong anion or weak cation exchanger detected 45 individual compounds of molecular masses ranged from 1.964 to 66.5 kD. Comparison of SELDI-MS data with protein databases revealed eight known proteins from the medicinal leech. Other masses detected by proteomic analytical methods may be related to both modifications of known proteins and unknown biologically active components of leech saliva secretion.

Animals↗

Proteomics for biodefense applications: progress and opportunities.

The increasing threat of bioterrorism and continued emergence of new infectious diseases has driven a major resurgence in biomedical research efforts to develop improved treatments, diagnostics and vaccines, as well as increase the fundamental understanding of the host immune response to infectious agents. The availability of multiple mass spectrometry platforms combined with multidimensional separation technologies and microbial genomic databases provides an unprecedented opportunity to develop these much needed resources. An overview of current proteomic strategies applied to microbes and viruses considered potential bioterrorism agents is presented. The emerging area of immunoproteomics as applied to the development of new vaccine targets is also summarized. These powerful research approaches can generate a multitude of potential new protein targets; however, translating these proteomic discoveries to useful counter-bioterrorism products will require large collaborative research efforts across multiple basic science and clinical disciplines. A translational proteomic research paradigm illustrating this approach using influenza virus as an example is discussed.

Adult↗

Construction of a two-dimensional gel electrophoresis protein database for the Nicotiana tabacum cv. Bright Yellow-2 cell suspension culture.

Using two-dimensional gel electrophoresis (2-DE) and electrospray-tandem mass spectrometry (ESI-MS/MS), we have started the proteome analysis of the cell line Nicotiana tabacum cv. Bright Yellow-2 (tobacco BY-2). The BY-2 cell suspension culture is widely used as a model system to study the growth and development of plant cells. We present a protocol describing the sample preparation and 2-DE, enabling us to separate and display more than 1000 proteins from this cell culture. A reference gel was generated, using immobilized pH gradient isoelectric focusing in a linear gradient from pH 3 to 10 and 12% Sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). Although the tobacco genome is not sequenced yet, a range of protein spots from this reference map was identified by means of a semi-automated liquid chromatography-ESI-quadrupole time of flight-tandem MS (LC-ESI-QTOF-MS-MS) setup and cross-species matching. These data were integrated in a database, which can be accessed at http://tby2-www.uia.ac.be/tby2/. On the on-line reference map, the identified protein spots are hyperlinked to individual protein entries. Each protein entry contains all identification information, as well as links to relevant entries in other on-line databases. Comprehensive search functions are implemented. Especially for an unsequenced but widespread model organism like tobacco BY-2, such a reference database is a convenient source for protein information that brings protein identification within reach without the need for extensive MS. This publicly accessible database provides a solid basis for tobacco BY-2 proteomics in the future.

Databases as Topic↗

The HUPO proteomics standards initiative--overcoming the fragmentation of proteomics data.

Proteomics is a key field of modern biomolecular research, with many small and large scale efforts producing a wealth of proteomics data. However, the vast majority of this data is never exploited to its full potential. Even in publicly funded projects, often the raw data generated in a specific context is analysed, conclusions are drawn and published, but little attention is paid to systematic documentation, archiving, and public access to the data supporting the scientific results. It is often difficult to validate the results stated in a particular publication, and even simple global questions like "In which cellular contexts has my protein of interest been observed?" can currently not be answered with realistic effort, due to a lack of standardised reporting and collection of proteomics data. The Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organisation (HUPO), defines community standards for data representation in proteomics to facilitate systematic data capture, comparison, exchange and verification. In this article we provide an overview of PSI organisational structure, activities, and current results, as well as ways to get involved in the broad-based, open PSI process.

Databases, Protein↗

Profiling the malaria genome: a gene survey of three species of malaria parasite with comparison to other apicomplexan species.

We have undertaken the first comparative pilot gene discovery analysis of approximately 25,000 random genomic and expressed sequence tags (ESTs) from three species of Plasmodium, the infectious agent that causes malaria. A total of 5482 genome survey sequences (GSSs) and 5582 ESTs were generated from mung bean nuclease (MBN) and cDNA libraries, respectively, of the ANKA line of the rodent malaria parasite Plasmodium berghei, and 10,874 GSSs generated from MBN libraries of the Salvador I and Belem lines of Plasmodium vivax, the most geographically wide-spread human malaria pathogen. These tags, together with 2438 Plasmodium falciparum sequences present in GenBank, were used to perform first-pass assembly and transcript reconstruction, and non-redundant consensus sequence datasets created. The datasets were compared against public protein databases and more than 1000 putative new Plasmodium proteins identified based on sequence similarity. Homologs of previously characterized Plasmodium genes were also identified, increasing the number of P. vivax and P. berghei sequences in public databases at least 10-fold. Comparative studies with other species of Apicomplexa identified interesting homologs of possible therapeutic or diagnostic value. A gene prediction program, Phat, was used to predict probable open reading frames for proteins in all three datasets. Predicted and non-redundant BLAST-matched proteins were submitted to InterPro, an integrated database of protein domains, signatures and families, for functional classification. Thus a partial predicted proteome was created for each species. This first comparative analysis of Plasmodium protein coding sequences represents a valuable resource for further studies on the biology of this important pathogen.

Animals↗

The dynamic range of protein expression: a challenge for proteomic research.

Proteomic research, for its part, is benefiting enormously from the last decade of genomic research as we now have archived, annotated and audited sequence databases to correlate and query experimental data. While the two-dimensional electrophoresis (2-DE) gels are still a central part of proteomics, we reflect on the possibilities and realities of the current 2-DE technology with regard to displaying and analysing proteomes. Limitations of analysing whole cell/tissue lysates by 2-DE alone are discussed, and we investigate whether extremely narrow p/ranges (1 pH unit/25 cm) provide a solution to display comprehensive protein expression profiles. We are confronted with a challenging task: the dynamic range of protein expression. We believe that most of the existing technology is capable of displaying many more proteins than is currently achievable by integrating existing and new techniques to prefractionate samples prior to 2-DE display or analysis. The availability of a "proteomics toolbox", consisting of defined reagents, methods, and equipment, would assist a comprehensive analysis of defined biological systems.

Chemical Fractionation↗

Methods for on-chip protein analysis.

The unambiguous identification of peptides/proteins is crucial for the definition of the proteome. Using ProteinChip Array technology also known as surface-enhanced laser desorption/ionization-time of flight mass spectrometry (SELDI-TOF MS), we developed experimental protocols and probed test conditions required for the protein identification on ProteinChip surfaces. We were able to directly digest peptides/proteins on-chip surfaces by specific proteases, such as trypsin, and to obtain the peptide mass fingerprint of the sample under investigation by its direct analysis on a simple laser desorption/ionization mass spectrometer. Furthermore, tandem mass spectrometry was performed on several of the resulting tryptic peptides by using collision quadrupole time of flight (Qq-TOF) MS/MS via the ProteinChip interface, thus allowing the unambiguous identification of the protein(s) within the sample. In addition, we were able to identify the C-terminal sequence of peptides by their digestion with carboxypeptidase Y directly on ProteinChip surfaces coupled with SELDI-TOF MS analysis of the resulting peptide mass ladders employing the instrument's protein ladder sequence software. Moreover, the removal of up to nine amino acid residues from the C-terminal end of a peptide extends the functional range of Qq-TOF MS/MS sequence determination to over 3000 m/z. The utility of these procedures for the proteome exploration are discussed.

Amino Acid Sequence↗

Importance of databases in experimental and clinical allergology.

Information technology (IT) is leading us to reconsider some of the approaches we have been using in both basic research and clinical work in allergology. Resources mainly coming from the advent of the Internet are further amplified by the parallel development of other new tools, such as molecular biology and nanotechnology. These three powerful tools are now available and are cross-linked to a certain degree to express their power when applied to biomedical fields. Bioinformatics applied to allergy simplifies our way of handling an increasing wealth of knowledge. This review assesses the current status of allergen databases that are mainly dedicated to sequence homology collection for computational purposes. Whether or not they integrate features that are now typical of IT in other biomedical fields is analyzed as well. The results of these analyses are discussed with a view to the critical need of integrating biochemical data with clinical, epidemiological information and how this goal can be reached by the use of proteomic microarrays for IgE detection. Future directions for a more comprehensive use of allergen databases are proposed.

Allergens↗

Multidimensional protein identification technology: current status and future prospects.

Protein profiling using high-throughput tandem mass spectrometry has become a powerful method for analyzing changes in global protein expression patterns in cells and tissues as a function of developmental, physiologic and disease processes. This review summarizes the utility and practical application of multidimensional protein identification technology as a platform for comprehensive proteomic profiling of complex biologic samples. The strengths and potential problems and limitations associated with this powerful technology are discussed, with an emphasis placed on one of the biggest challenges currently facing large-scale expression profiling projects -- namely, data analysis. Complementary bioinformatic computational data mining strategies, such as clustering, functional annotation and statistical inference, are also discussed as these are increasingly necessary for interpreting the results of global proteomic profiling studies.

Animals↗

Rapid validation of protein identifications with the borderline statistical confidence via de novo sequencing and MS BLAST searches.

Protein identifications with the borderline statistical confidence are typically produced by matching a few marginal quality MS/MS spectra to database peptide sequences and represent a significant bottleneck in the reliable and reproducible characterization of proteomes. Here, we present a method for rapid validation of borderline hits that circumvents the need in, often biased, manual inspection of raw MS/MS spectra. The approach takes advantage of the independent interpretation of corresponding MS/MS spectra by PepNovo de novo sequencing software followed by mass spectrometry-driven BLAST (MS BLAST) sequence-similarity database searches that utilize all partially inaccurate, degenerate and redundant candidate peptide sequences. In a case study involving the identification of more than 180 Caenorhabditis elegans proteins by nanoLC-MS/MS analysis on a linear ion trap LTQ mass spectrometer, the approach enabled rapid assignment (confirmation or rejection) of more than 70% of Mascot hits of borderline statistical confidence.

Amino Acid Sequence↗

Charge state estimation for tandem mass spectrometry proteomics.

High-throughput protein analysis by tandem mass spectrometry produces anywhere from thousands to millions of spectra that are being used for peptide and protein identifications. Though each spectrum corresponds only to one charged peptide (ion) state, repetitive database searches of multiple charge states are typically conducted since the resolution of many common mass spectrometers is not sufficient to determine the charge state. The resulting database searches are both error-prone and time-consuming. We describe a straightforward, accurate approach on charge state estimation (CHASTE). CHASTE relies on fragment ion peak distributions, and by using reliable logistic regression models, combines different measurements to improve its accuracy. CHASTE's performance has been validated on data sets, comprised of known peptide dissociation spectra, obtained by replicate analyses of our earlier developed protein standard mixture using ion trap mass spectrometers at different laboratories. CHASTE was able to reduce number of needed database searches by at least 60% and the number of redundant searches by at least 90% virtually without any informational loss. This greatly alleviates one of the major bottlenecks in high throughput peptide and protein identifications. Thresholds and parameter estimates can be tailored to specific analysis situations, pipelines, and instrumentations. CHASTE was implemented in Java GUI-based and command-line-based interfaces.

Computer Graphics↗

An automated matrix-assisted laser desorption/ionization quadrupole Fourier transform ion cyclotron resonance mass spectrometer for "bottom-up" proteomics.

Here we describe a new quadrupole Fourier transform ion cyclotron resonance hybrid mass spectrometer equipped with an intermediate-pressure MALDI ion source and demonstrate its suitability for "bottom-up" proteomics. The integration of a high-speed MALDI sample stage, a quadrupole analyzer, and a FT-ICR mass spectrometer together with a novel software user interface allows this instrument to perform high-throughput proteomics experiments. A set of linearly encoded stages allows sub-second positioning of any location on a microtiter-sized target with up to 1536 samples with micrometer precision in the source focus of the ion optics. Such precise control enables internal calibration for high mass accuracy MS and MS/MS spectra using separate calibrant and analyte regions on the target plate, avoiding ion suppression effects that would result from the spiking of calibrants into the sample. An elongated open cylindrical analyzer cell with trap plates allows trapping of ions from 1000 to 5000 m/z without notable mass discrimination. The instrument is highly sensitive, detecting less than 50 amol of angiotensin II and neurotensin in a microLC MALDI MS run under standard experimental conditions. The automated tandem MS of a reversed-phase separated bovine serum albumin digest demonstrated a successful identification for 27 peptides covering 45% of the sequence. An automated tandem MS experiment of a reversed-phase separated yeast cytosolic protein digest resulted in 226 identified peptides corresponding to 111 different proteins from 799 MS/MS attempts. The benefits of accurate mass measurements for data validation for such experiments are discussed.

Calibration↗