Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

FCP: functional coverage of the proteome by structures.

MOTIVATION: Tools and resources for translating the remarkable growth witnessed in recent years in the number of protein structures determined experimentally into actual gain in the functional coverage of the proteome are becoming increasingly necessary. We introduce FCP, a publicly accessible web tool dedicated to analyzing the current state and trends of the population of structures within protein families. FCP offers both graphical and quantitative data on the degree of functional coverage of enzymes and nuclear receptors by existing structures, as well as on the bias observed in the distribution of structures along their respective functional classification schemes. AVAILABILITY: http://cgl.imim.es/fcp CONTACT: jmestres@imim.es.

Algorithms↗

Use of performic acid oxidation to expand the mass distribution of tryptic peptides.

Significant identification of proteins by mass fingerprinting and partial sequencing of tryptic peptides is central to proteomics. However, peptide masses cluster with distances of approximately 1 Da. Expanding these clusters will give more peptides of unique masses, thereby identifying proteins with a higher significance. The mass clusters can be expanded downward by including more oxygen atoms in the peptides. Classic performic acid oxidation modifies three residues, Cys to CysO(3), Met to MetO(2), and Trp to TrpO(2). In this study, we compare the mass distributions of tryptic peptides computed from the predicted proteomes of Bacillus subtilis, Drosophila melanogaster, Arabidopsis thaliana, and Homo sapiens modified by oxidation, reduction, and reduction followed by carboxymethylation, carboxamidomethylation, or pyridylethylation. Forty to 46% of the eukaryotic tryptic peptides contain Cys, Met, or Trp. Additionally, the importance of mass accuracy of differentially modified tryptic peptides for significant protein identification by database searches was analyzed. The results show that performic acid oxidation gives markedly extended mass distributions at mass accuracies from +/-0.002 to +/-0.25 Da for the eukaryotes. The effect of the expanded mass distribution on significant protein identification was illustrated by searching simulated mass peak lists against the databases containing oxidized and reduced tryptic peptides. The specificity of formic acid oxidation was tested experimentally, and no general adverse effects were detected. Tryptic peptides provided a 100% sequence coverage of oxidized barley grain peroxidase by LC-MS, and the sequence coverages of oxidized and carboxymethylated bovine serum albumin were similar by MALDI-TOF MS analyses.

Amino Acids↗

Changes of chondrocyte metabolism in vitro: an approach by proteomic analysis.

Changes in chondrocyte metabolism in vitro using different support systems and under different culture conditions were studied with a proteomic approach. Qualitative and quantitative modifications in the synthesis of chondrocyte proteins were investigated using two-dimensional (2D) gel electrophoresis. This technique provided a simple way to visualize the most abundant chondrocyte proteins. Proteins were identified after in-gel proteolysis with trypsin and matrix-assisted laser desorption ionization-time of flight mass spectrometry, using peptide mass fingerprinting. Tryptic peptide masses were measured and matched against a computer-generated list from the simulated trypsin proteolysis of a protein database (SwissProt).

Animals↗

Characterization of the Drosophila melanogaster mitochondrial proteome.

We have combined high-resolution two-dimensional (2-D) gel electrophoresis with mass spectrometry with the aim of identifying proteins represented in the 2-D gel database of Drosophila melanogaster mitochondria. First, we purified mitochondria from third instar Drosophila larvae and constructed a high-resolution 2-D gel database containing 231 silver-stained polypeptides. Next, we carried out preparative 2-D PAGE to isolate some of the polypeptides and characterize them by MALDI-TOF analysis. Using this strategy, we identified 66 mitochondrial spots in the database, and in each case confirmed their identity by MALDI-TOF/TOF analysis. In addition, we generated antibodies against two of the mitochondrial proteins as tools for characterizing the organelle.

Amino Acid Sequence↗

Proteomics strategies for protein identification.

The information from genome sequencing provides new approaches for systems-wide understanding of protein networks and cellular function. DNA microarray technologies have advanced to the point where nearly complete monitoring of gene expression is feasible in several organisms. An equally important goal is to comprehensive survey cellular proteomes and profile protein changes under different cellular states. This presents a complex analytical problem, due to the chemical variability between proteins and peptides. Here, we discuss strategies to improve accuracy and sensitivity of peptide identification, distinguish represented protein isoforms, and quantify relative changes in protein abundance.

Computational Biology↗

Proteome analysis using isoelectric focusing in immobilized pH gradient gels followed by mass spectrometry.

Over the past several years, a large effort has been focused on improvements of two-dimensional (2-D) gel electrophoresis-based proteomics technology, and on development of novel approaches for proteome analysis. Here, we describe the application of an alternative strategy for the analysis of complex proteomes. The strategy combines isoelectric focusing in immobilized pH gradient strips (in-gel IEF), mass spectrometry (MS), and bioinformatics. A protein mixture is separated by in-gel IEF, and the entire strip is cut into a set of gel sections. Proteins in each gel section are digested with trypsin, and the tryptic peptides are subjected to liquid chromatography-nanoelectrospray-quadrupole ion-trap tandem mass spectrometry (LC-ESI-MS/MS). The LC-ESI-MS/MS data are used to identify the proteins through searches of a protein sequence database. Using this in-gel IEF-LC-MS/MS strategy, we have identified 127 proteins from a human pituitary. This study demonstrates the potential of the in-gel IEF-LC-MS/MS approach for analyses of complex mammalian proteomes.

Computational Biology↗

Protein functions and biological contexts.

The availability of a rough draft of the predicted human proteome allows an evaluation of the extent to which the predicted and biochemical functions of proteins are in alignment, and the roles of different technologies and approaches to understanding human diseases and instantiating therapeutics. Microarray technologies at the transcriptomic and proteomic levels can be high throughput and excellent for diagnostic purposes, but their informational outputs are inferior in quality to those emerging from the co- and post-translational levels and from antibody-based molecular anatomy. It is now abundantly clear that data transfer between the transcriptome and proteome is not straightforward, and that increasing emphasis needs to be placed on pure proteomic approaches at the structural, quantitative, cell biological and phenomic levels, with special focus on embryogenic and foetal processes. Finally, the precision genetic engineering that is required to evaluate the functional significance of context-dependent protein interactions underpinned by post-translational modifications and proteolytic cleavage events, is still too time consuming and rudimentary to be implemented on a large scale in the mouse, and basic principles and first order networks will need to be sorted out in even simpler model systems such as Drosophila.

Databases, Protein↗

A Fourier transformation based method to mine peptide space for antimicrobial activity.

BACKGROUND: Naturally occurring antimicrobial peptides are currently being explored as potential candidate peptide drugs. Since antimicrobial peptides are part of the innate immune system of every living organism, it is possible to discover new candidate peptides using the available genomic and proteomic data. High throughput computational techniques could also be used to virtually scan the entire peptide space for discovering out new candidate antimicrobial peptides. RESULT: We have identified a unique indexing method based on biologically distinct characteristic features of known antimicrobial peptides. Analysis of the entries in the antimicrobial peptide databases, based on our indexing method, using Fourier transformation technique revealed a distinct peak in their power spectrum. We have developed a method to mine the genomic and proteomic data, for the presence of peptides with potential antimicrobial activity, by looking for this distinct peak. We also used the Euclidean metric to rank the potential antimicrobial peptides activity. We have parallelized our method so that virtually any given protein space could be data mined, in search of antimicrobial peptides. CONCLUSION: The results show that the Fourier transform based method with the property based coding strategy could be used to scan the peptide space for discovering new potential antimicrobial peptides.

Amino Acid Sequence↗

Informatics and data management in proteomics.

Proteomics has become dominated by large amounts of experimental data and interpreted results. This experimental data cannot be effectively used without understanding the fundamental structure of its information content and representing that information in such a way that knowledge can be extracted from it. This review explores the structure of this information with regard to three fundamental issues: the extraction of relevant information from raw data, the scale of the projects involved and the statistical significance of protein identification results.

Algorithms↗

A dataset of human liver proteins identified by protein profiling via isotope-coded affinity tag (ICAT) and tandem mass spectrometry.

Proteins from human liver carcinoma Huh7 cells, representing transformed liver cells, and cultured primary human fetal hepatocytes (HFH) and human HH4 hepatocytes, representing nontransformed liver cells, were extracted and processed for proteome analysis. Proteins from stimulated cells (interferon-alpha treatment for the Huh7 and HFH cells and induction of hepatitis C virus [HCV] proteins for the HH4 cells) and corresponding control cells were labeled with light and heavy cleavable ICAT reagents, respectively. The labeled samples were combined, trypsinized, and subject to cation-exchange and avidin-affinity chromatographies. The resulting cysteine-containing peptides were analyzed by microcapillary LC-MS/MS. The MS/MS spectra were initially analyzed by searching the human International Protein Index database using the SEQUEST software (1). Subsequently, new statistical algorithms were applied to the collective SEQUEST search results of each experiment. First, the PeptideProphet software (2) was applied to discriminate true assignments of MS/MS spectra to peptide sequences from false assignments, to assign a probability value for each identified peptide, and to compute the sensitivity and error rate for the assignment of spectra to sequences in each experiment. Second, the ProteinProphet software (3) was used to infer the protein identifications and to compute probabilities that a protein had been correctly identified, based on the available peptide sequence evidence. The resulting protein lists were filtered by a ProteinProphet probability score p > or = 0.5, which corresponded to an error rate of less than 5%. A total of 1,296, 1,430, and 1,476 proteins or related protein groups were identified in three subdatasets from the Huh7, HFH, and HH4 cells, respectively. In total, these subdatasets contained 2,486 unique protein identifications from human liver cells. An increase of the threshold to p > or = 0.9 (corresponding to an error rate of less than 1%) resulted in 2,159 unique protein identifications (1,146, 1,235, and 1,318 for the Huh7, HFH, and HH4 cells, respectively).

Algorithms↗

Informatics-assisted protein profiling in a transgenic mouse model of amyotrophic lateral sclerosis.

One of the causes of amyotrophic lateral sclerosis (ALS) is due to mutations in Cu,Zn-superoxide dismutase (SOD1). The mutant protein exhibits a toxic gain of function that adversely affects the function of neurons in the spinal cord, brain stem, and motor cortex. A proteomic analysis of protein expression in a widely used mouse model of ALS was undertaken to identify differences in protein expression in the spinal cords of mice expressing a mutant protein with the G93A mutation found in human ALS. Protein profiling was done on soluble and particulate fractions of spinal cord extracts using high throughput two-dimensional liquid chromatography coupled to tandem mass spectrometry. An integrated proteomics-informatics platform was used to identify relevant differences in protein expression based upon the abundance of peptides identified by database searching of mass spectrometry data. Changes in the expression of proteins associated with mitochondria were particularly prevalent in spinal cord proteins from both mutant G93A-SOD1 and wild-type SOD1 transgenic mice. G93A-SOD1 mouse spinal cord also exhibited differences in proteins associated with metabolism, protein kinase regulation, antioxidant activity, and lysosomes. Using gene ontology analysis, we found an overlap of changes in mRNA expression in presymptomatic mice (from microarray analysis) in three different gene categories. These included selected protein kinase signaling systems, ATP-driven ion transport, and neurotransmission. Therefore, alterations in selected cellular processes are detectable before symptomatic onset in ALS mouse models. However, in late stage disease, mRNA expression analysis did not reveal significant changes in mitochondrial gene expression but did reveal concordant changes in lipid metabolism, lysosomes, and the regulation of neurotransmission. Thus, concordance of proteomic and mRNA expression data within multiple categories validates the use of gene ontology analysis to compare different types of "omic" data.

Amyotrophic Lateral Sclerosis↗

NCBI GEO: mining millions of expression profiles--database and tools.

The Gene Expression Omnibus (GEO) at the National Center for Biotechnology Information (NCBI) is the largest fully public repository for high-throughput molecular abundance data, primarily gene expression data. The database has a flexible and open design that allows the submission, storage and retrieval of many data types. These data include microarray-based experiments measuring the abundance of mRNA, genomic DNA and protein molecules, as well as non-array-based technologies such as serial analysis of gene expression (SAGE) and mass spectrometry proteomic technology. GEO currently holds over 30,000 submissions representing approximately half a billion individual molecular abundance measurements, for over 100 organisms. Here, we describe recent database developments that facilitate effective mining and visualization of these data. Features are provided to examine data from both experiment- and gene-centric perspectives using user-friendly Web-based interfaces accessible to those without computational or microarray-related analytical expertise. The GEO database is publicly accessible through the World Wide Web at http://www.ncbi.nlm.nih.gov/geo.

Animals↗

Computational analyses of high-throughput protein-protein interaction data.

Protein-protein interactions play important roles in nearly all events that take place in a cell. High-throughput experimental techniques enable the study of protein-protein interactions at the proteome scale through systematic identification of physical interactions among all proteins in an organism. High-throughput protein-protein interaction data, with ever-increasing volume, are becoming the foundation for new biological discoveries. A great challenge to bioinformatics is to manage, analyze, and model these data. In this review, we describe several databases that store, query, and visualize protein-protein interaction data. Comparison between experimental techniques shows that each high-throughput technique such as yeast two-hybrid assay or protein complex identification through mass spectrometry has its limitations in detecting certain types of interactions and they are complementary to each other. In silico methods using protein/DNA sequences, domain and structure information to predict protein-protein interaction can expand the scope of experimental data and increase the confidence of certain protein-protein interaction pairs. Protein-protein interaction data correlate with other types of data, including protein function, subcellular location, and gene expression profile. Highly connected proteins are more likely to be essential based on the analyses of the global architecture of large-scale interaction network in yeast. Use of protein-protein interaction networks, preferably in conjunction with other types of data, allows assignment of cellular functions to novel proteins and derivation of new biological pathways. As demonstrated in our study on the yeast signal transduction pathway for amino acid transport, integration of high-throughput data with traditional biology resources can transform the protein-protein interaction data from noisy information into knowledge of cellular mechanisms.

Animals↗

Utilisation of proteomics datasets generated via multidimensional protein identification technology (MudPIT).

Technological developments in proteomics have had a dramatic impact on biology in recent years. One of these developments--named multidimensional protein identification technology (MudPIT)--couples two-dimensional chromatography of peptides in mass spectrometry-compatible solutions directly to tandem mass spectrometry, allowing for the identification of proteins from highly complex mixtures. Since the initial descriptions of MudPIT, this approach has been implemented in the analysis of whole proteomes, organelles and protein complexes. Key aspects of many of the analyses are the validation of MudPIT datasets with alternate strategies and the integration of MudPIT datasets with other biochemical, cell biology or molecular biology approaches. This paper presents strategies for validating MudPIT datasets and incorporating these datasets into biologically driven experimental design.

Automation↗

Systematic identification in silico of covalently bound cell wall proteins and analysis of protein-polysaccharide linkages of the human pathogen Candida glabrata.

Candida glabrata is an important cause of systemic candidiasis in humans. This paper reports a systematic analysis of the putative glycosylphosphatidylinositol-modified (GPI) proteins of C. glabrata, a large part of which are covalently bound to the cell wall glucan network and the remainder of which are retained in the plasma membrane, and of cell wall proteins (CWPs) which are covalently bound in a mild-alkali-sensitive manner. In silico genomic analysis revealed 106 putative GPI proteins. Fifty-one of these GPI proteins could be categorized as adhesive proteins, potentially implicated in fungus-host interactions or biofilm formation during the development of fungal infections. Eleven proteins belonged to well-known GPI protein families of glycoside hydrolases, probably involved in cell wall expansion and remodelling during growth. Other identified GPI proteins included phospholipases, aspartic proteases, homologues of ScEcm33p and ScKre1p, and structural CWPs. Interestingly, the GPI algorithm predicted three orthologues of an abundant CWP in S. cerevisiae, Cwp1p, which is absent in Candida albicans. To evaluate the in silico predictions, isolated cell walls were extracted using HF-pyridine, which specifically cleaves phosphodiester bonds, to release GPI-CWPs. Immunological analysis of the extract using one-dimensional SDS-PAGE and anti-ScCwp1p antiserum indicated the presence of a Cwp1p homologue in C. glabrata cell walls. Further analysis by two-dimensional gel electrophoresis and electrospray ionization tandem mass spectrometry (ESI-MS/MS) confirmed the presence of two of the predicted Cwp1p proteins, Cwp1.1p and Cwp1.2p. Crh1p, a putative 1,3-beta-glucan remodelling enzyme, was also identified. In silico genomic analysis further revealed five putative Pir proteins (Pir1-5p) and five members of the Bgl2 glycoside hydrolase family 17, belonging to a class of putative CWPs that can be extracted with NaOH. Immunological analysis of mild-alkali-extracted CWPs showed the presence of a ScPir2p homologue. Together, these experimental data and in silico predictions represent the first systematic analysis of the C. glabrata cell wall proteome.

Amino Acid Sequence↗

Further steps towards data standardisation: the Proteomic Standards Initiative HUPO 3(rd) annual congress, Beijing 25-27(th) October, 2004.

The increasing volume of proteomics data currently being generated by increasingly high-throughput methodologies has led to an increasing need for methods by which such data can be accurately described, stored and exchanged between experimental researchers and data repositories. Work by the Proteomics Standards Initiative of the Human Proteome Organisation has laid the foundation for the development of standards by which experimental design can be described and data exchange facilitated. The progress of these efforts, and the direct benefits already accruing from them, were described at a plenary session of the 3(rd) Annual HUPO congress. Parallel sessions allowed the three work groups to present their progress to interested parties and to collect feedback from groups already implementing the available formats.

China↗

[Application of proteomics in the research of discrepancy proteins in gastric cancer].

OBJECTIVE: To explore the discrepancy proteins in gastric cancer by proteome analysis. METHODS: Total proteins of gastric cancer tissues and matched normal gastric epithelial tissues were separated respectively by two-dimensional polyacrylamide gel electrophoresis (2-DE). Mass spectrometry was used to test the differentially expressed proteins. RESULTS: One thousand one hundred and forty-seven protein spots from gastric cancer tissue and 1079 spots from the normal tissue were gained. Out of 164 different protein spots, 41 were only expressed in gastric cancer tissue, 27 were unique in normal tissue, 39 were up-regulated and 57 were down-regulated in gastric cancer. Seven proteins, which were highly expressed in gastric cancer tissue, were identified. CONCLUSION: Different protein spots between gastric cancer tissues and normal gastric epithelial tissue were gained by proteomics. The 7 discrepancy proteins were further identified. It establishes the foundation of finding specific gastric cancer proteins, which act as biomarkers for the diagnosis and prognosis of gastric cancer.

Case-Control Studies↗