Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Proteomics and disease: opportunities and challenges.

Since the human genome was sequenced, there has been intense activity to understand the function of the 30,000 identified genes; attention has now turned to the products of genes--proteins. Proteomics is the large-scale study of the structure and function of proteins; it includes the rapidly evolving field of disease proteomics, which aims to identify proteins involved in human disease and to understand how their expression, structure and function cause illness. Proteomics has identified proteins that offer promise as diagnostic or prognostic markers, or as therapeutic targets in a range of illnesses, including cancer, immune rejection after transplantation, and infectious diseases such as tuberculosis and malaria; it has the potential to allow patient-tailored therapy. Some major challenges remain, both technical (eg, detecting "low-abundance" proteins, and maintaining sample stability) and in data management (eg, correlating changes in proteins with disease processes).

Databases, Factual↗

A computational approach for ordering signal transduction pathway components from genomics and proteomics Data.

BACKGROUND: Signal transduction is one of the most important biological processes by which cells convert an external signal into a response. Novel computational approaches to mapping proteins onto signaling pathways are needed to fully take advantage of the rapid accumulation of genomic and proteomics information. However, despite their importance, research on signaling pathways reconstruction utilizing large-scale genomics and proteomics information has been limited. RESULTS: We have developed an approach for predicting the order of signaling pathway components, assuming all the components on the pathways are known. Our method is built on a score function that integrates protein-protein interaction data and microarray gene expression data. Compared to the individual datasets, either protein interactions or gene transcript abundance measurements, the integrated approach leads to better identification of the order of the pathway components. CONCLUSIONS: As demonstrated in our study on the yeast MAPK signaling pathways, the integration analysis of high-throughput genomics and proteomics data can be a powerful means to infer the order of pathway components, enabling the transformation from molecular data into knowledge of cellular mechanisms.

Computational Biology↗

Location proteomics: a systems approach to subcellular location.

Systems Biology requires comprehensive systematic data on all aspects and levels of biological organization and function. In addition to information on the sequence, structure, activities and binding interactions of all biological macromolecules, the creation of accurate predictive models of cell behaviour will require detailed information on the distribution of those molecules within cells and the ways in which those distributions change over the cell cycle and in response to mutations or external stimuli. Current information on subcellular location in protein databases is limited to unstructured text descriptions or sets of terms assigned by human curators. These entries do not permit basic operations that are common to other biological databases, such as measurement of the degree of similarity between the distributions of two proteins, and they are not able to fully capture the complexity of protein patterns that can be observed. The field of location proteomics seeks to provide automated, objective high-resolution descriptions of protein location patterns within cells. Methods have been developed to group proteins into statistically indistinguishable location patterns using automated analysis of fluorescence microscope images. The resulting clusters, or location families, are analogous to clusters found for other domains, such as protein sequence families. Preliminary work suggests the feasibility of expressing each unique pattern as a generative model that can be incorporated into comprehensive models of cell behaviour.

Animals↗

Stem cell proteomes: a profile of human mesenchymal stem cells derived from umbilical cord blood.

Multipotent mesenchymal stem cells (MSCs) derived from human umbilical cord blood (UCB) represent promising candidates for the development of future strategies in cellular therapy. To create a comprehensive protein expression profile for UCB-MSCs, one UCB unit from a full-term delivery was isolated from the unborn placenta, transferred into culture, and their whole-cell protein fraction was subjected to two-dimensional electrophoresis (2-DE). Unambiguous protein identification was achieved with peptide mass fingerprinting matrix-assisted laser desorption/ionization - time of flight - mass spectrometry (MALDI-TOF-MS), peptide sequencing (MALDI LIFT-TOF/TOF MS), as well as gel-matching with previously identified databases. In overall five replicate 2-DE runs, a total of 2037 +/- 437 protein spots were detected of which 205 were identified representing 145 different proteins and 60 isoforms or post-translational modifications. The identified proteins could be grouped into several functional categories, such as metabolism, folding, cytoskeleton, transcription, signal transduction, protein degradation, detoxification, vesicle/protein transport, cell cycle regulation, apoptosis, and calcium homeostasis. The acquired proteome map of nondifferentiated UCB-MSCs is a useful inventory which facilitates the identification of the normal proteomic pattern as well as its changes due to activated or suppressed pathways of cytosolic signal transduction which occur during proliferation, differentiation, or other experimental conditions.

Databases, Protein↗

Large-gel two-dimensional electrophoresis-matrix assisted laser desorption/ionization-time of flight-mass spectrometry: an analytical challenge for studying complex protein mixtures.

The large-gel two-dimensional electrophoresis (2-DE) technique, developed by Klose and co-workers over the past 25 years, provides the resolving power necessary to separate crude proteome extracts of higher eukaryotes. Matrix assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS) provides the sample throughput necessary to identify thousands of different protein species in an adequate time period. Spot excision, in situ proteolysis, and extraction of the cleavage products from the gel matrix, peptide purification and concentration as well as the mass spectrometric sample preparation are the crucial steps that interface the two analytical techniques. Today, these routines and not the mass spectrometric instrumentation determine how many protein digests can be analyzed per day per instrument. The present paper focuses on this analytical interface and reports on an integrated protocol and technology developed in our laboratory. Automated identification of proteins in sequence databases by mass spectrometric peptide mapping requires a powerful search engine that makes full use of the information contained in the experimental data, and scores the search results accordingly. This challenge is heading a second part of the paper.

Animals↗

Proteomic study of the Arabidopsis thaliana chloroplastic envelope membrane utilizing alternatives to traditional two-dimensional electrophoresis.

With the completion of the sequencing of the Arabidopsis genome and with the significant increase in the amount of other plant genome and expressed sequence tags (ESTs) data, plant proteomics is rapidly becoming a very active field. We have pursued a high-throughput mass spectrometry-based proteomics approach to identify and characterize membrane proteins localized to the Arabidopsis thaliana chloroplastic envelope membrane. In this study, chloroplasts were prepared from plate- or soil-grown Arabidopsis plants using a novel isolation procedure, and "mixed" envelopes were subsequently isolated using sucrose step gradients. We applied two alternative methodologies, off-line multidimensional protein identification technology (Off-line MUDPIT) and one-dimensional (1D) gel electrophoresis followed by proteolytic digestion and liquid chromatography coupled with tandem mass spectrometry (Gel-C-MS/MS), to identify envelope membrane proteins. This proteomic study enabled us to identify 392 nonredundant proteins.

Amino Acid Sequence↗

[Pharmacogenomics and pharmainformatics].

Pharmacogenomics is defined to identify the genes which are involved in determining the responsiveness and to distinguish responders and non-responders to a given drug. Genome sequencing, transcriptome and proteome analysis are of particular significance in pharmacogenomics. Sequencing is used to locate polymorphisms, and monitoring of gene expression can provide clue about the genomic response to disease and treatment. The transcriptome analysis can be done by methods of random cDNA sequencing (expressed sequence tag project, body map project, serial analysis of gene expression, et al), mRNA display (differential display, fluorescent differential display, RNA arbitraly primed PCR, molecular indexing, gene expression fingerprinting, et al) and differential hybridization(cDNA high density filter, cDNA microarray, oligomicrochip, et al). We used transcriptome analysis to identify therapeutic target genes by studying change of gene expression in animal models of cerebral vasospasm (1) and of hypoxia/ischemia and found novel drug target candidates through this pharmacogenomic strategy (2). We found remarkable up-regulation of heme oxygenase-1(HO-1) mRNA in the basilar artery and it might be closely related to the occurrence of delayed vasospasm after subarachnoid hemorrhage. In this report, we clearly demonstrate that intrathecal administration of antisense HO-1 oligodeoxynucleotide aggravates vasospasm, suggesting HO-1 gene induction has spasmolytic effects. Furthermore, we found the protective effects of HO-1 gene induction by endogenous or clinical compounds in cerebral vasospasm. Therapeutic gene induction of HO-1 could be a novel strategy for the prevention and treatment of Hb-induced pathologic conditions including delayed cerebral vasospasm. Our results suggest that the pharmacogenomic transcriptome analysis and pharmainformatics has the potential for strategy to define novel drug targets in various diseases (3). (1) J Clin Invest 104: 59-66, 1999. (2) J Biol Chem 276: 19921-19928, 2001. (3) J Cardiovasc Pharm 36: S1-S4, 2000.

Animals↗

A novel MALDI LIFT-TOF/TOF mass spectrometer for proteomics.

A new matrix-assisted laser-desorption/ionization time-of-flight/time-of-flight mass spectrometer with the novel "LIFT" technique (MALDI LIFT-TOF/TOF MS) is described. This instrument provides high sensitivity (attomole range) for peptide mass fingerprints (PMF). It is also possible to analyze fragment ions generated by any one of three different modes of dissociation: laser-induced dissociation (LID) and high-energy collision-induced dissociation (CID) as real MS/MS techniques and in-source decay in the reflector mode of the mass analyzer (reISD) as a pseudo-MS/MS technique. Fully automated operation including spot picking from 2D gels, in-gel digestion, sample preparation on MALDI plates with hydrophilic/hydrophobic spot profiles and spectrum acquisition/processing lead to an identification rate of 66% after the PMF was obtained. The workflow control software subsequently triggered automated acquisition of multiple MS/MS spectra. This information, combined with the PMF increased the identification rate to 77%, thus providing data that allowed protein modifications and sequence errors in the protein sequence database to be detected. The quality of the MS/MS data allowed for automated de novo sequencing and protein identification based on homology searching.

Amino Acid Sequence↗

The holm oak leaf proteome: analytical and biological variability in the protein expression level assessed by 2-DE and protein identification tandem mass spectrometry de novo sequencing and sequence similarity searching.

As a first approach in establishing the holm oak leaf proteome, we have optimised a protocol for this plant and tissue which includes the following steps: trichloroacetic acid-acetone extraction, two-dimensional gel electrophoresis (2-DE) on pH 5 to 8 linear gradient immobilised pH gradient strips as the first dimension, and sodium dodecyl sulfate-polyacrylamide gel electrophoresis on 13% polyacrylamide gels as the second one. Proteins were detected by Coomassie staining. Gel images were recorded and digitalized, and the protein spots quantified by using a linear regression equation of protein quantity on spot volume obtained against standard proteins. Analytical variance was calculated for one-hundred protein spots from three replicate 2-DE gels of the same protein extract. Biological variance was determined for the same protein spots from independent tissue extracts corresponding to leaves from different trees, or the same tree at different orientations or sampling times during a day. Values of 26% for the analytical variance and 58.6% for the biological variance among independent trees were obtained. These values provide a quantified and statistical basis for the evaluation of protein expression changes in comparative proteomic investigations with this species. A representative set of the major proteins, covering the isoelectric point range of 5 to 8 and the relative molecular mass(r) range of 14 to 78 kDa, were subjected to liquid chromatography-tandem mass spectrometry analysis. Due to the absence of Quercus DNA or protein sequence databases, a method based on the procedure reported by Liska and Shevchenko including de novo sequencing and BLAST similarity searching against other plant species databases was used for protein identification. Out of 43 analysed spots, 35 were positively identified. The identified proteins mainly corresponded to enzymes involved in photosynthesis and energetic metabolism, with a significant number corresponding to RubisCO.

Amino Acid Sequence↗

Protein identification from product ion spectra of peptides validated by correlation between measured and predicted elution times in liquid chromatography/mass spectrometry.

Reversed-phase liquid chromatography (LC) directly coupled with electrospray-tandem mass spectrometry (MS/MS) is a successful choice to obtain a large number of product ion spectra from a complex peptide mixture. We describe a search validation program, ScoreRidge, developed for analysis of LC-MS/MS data. The program validates peptide assignments to product ion spectra resulting from usual probability-based searches against primary structure databases. The validation is based only on correlation between the measured LC elution time of each peptide and the deduced elution time from the amino acid sequence assigned to product ion spectra obtained from the MS/MS analysis of the peptide. Sufficient numbers of probable assignments gave a highly correlative curve. Any peptide assignments within a certain tolerance from the correlation curve were accepted for the following arrangement step to list identified proteins. Using this data validation program, host protein candidates responsible for interaction with human hepatitis B virus core protein were identified from a partially purified protein mixture. The present simple and practical program complements protein identification from usual product ion search algorithms and reduces manual interpretation of the search result data. It will lead to more explicit protein identification from complex peptide mixtures such as whole proteome digests from tissue samples.

Chromatography, Liquid↗

Identification of protein vaccine candidates from Helicobacter pylori using a preparative two-dimensional electrophoretic procedure and mass spectrometry.

Helicobacter pylori is an important human gastric pathogen for which the entire genome sequence is known. This microorganism displays a uniquely complex pattern of binding to complex carbohydrates presented on host mucosal surfaces and other tissues, through adhesion molecules (adhesins) on the microbial cell surface. Adhesins and other membrane-associated proteins are important targets for vaccine development. The identification and characterization of cell-surface proteins expressed by H. pylori is a prerequisite for the development of vaccines designed to interfere with bacterial colonization of host tissues. However, identification of membrane proteins is difficult using a traditional proteomics approach employing 2D-PAGE. We have used a novel approach in the identification of microbial proteins that employs a rapid preparative two-dimensional electrophoretic separation followed by mass spectrometry and database searches. No pre-enrichment of bacterial membranes was required. The entire process, from sample preparation to protein identification, can be completed in less than 18 hours, and the presence of proteins can be monitored after both the first- and second-dimensional separations using mass spectrometry. We were able to identify 40 proteins from a detergent-solubilized H. pylori preparation; over one-third of these were membrane or membrane-associated proteins. A functionally characterized low-abundance membrane protein, the Leb-binding adhesin, was found in this group. The use of this rapid 2D electrophoretic separation in proteomic studies of H. pylori is expected to speed up the identification of expressed virulence proteins and vaccine targets in this and other microbial pathogens.

Amino Acid Sequence↗

Megavariate data analysis of mass spectrometric proteomics data using latent variable projection method.

There are many data mining techniques for processing and general learning of multivariate data. However, we believe the wavelet transformation and latent variable projection method are particularly useful for spectroscopic and chromatographic data. Projection based methods are designed to handle hugely multivariate nature of such data effectively. For the actual analysis of the data we have used latent variable projection methods such as principal component analysis (PCA) and partial least squares projection to latent structures based discriminant analysis (PLS-DA) to analyze the raw data presented to the participants of the First Duke Proteomics Data Mining Conference. PCA was used to solve problem #1 (clustering problem) and the PLS-DA was used to solve problem #2 (classification problem). The idea of internal and external cross-validation was used to validate the model obtained from the classification analysis. The simple two-component PLS-DA model obtained from the analysis performed well. The model has completely separated the two groups from all the data. The same model applied on two-thirds of the data showed good performance by external validation with independent test set of remaining 13 specimens obtained by setting aside the spectra of every third specimen (accuracy of 85%).

Artificial Intelligence↗

Proteomic determination of metabolic enzymes of the amnion cell: basis for a possible diagnostic tool?

Amniocentesis is a valuable and standard procedure for prenatal diagnosis of genetic or inborn errors of metabolism. Amnion cells are cultivated and chromosomes or proteins can be examined to provide molecular diagnosis. Mainly individual proteins are searched for based upon pedigrees and/or anamnesis. As inborn errors of metabolism involve a vast diversity of metabolic enzymes, we aimed to find a screening method for a large series of metabolic enzymes. Amnion cells were obtained from amniocentesis and subjected to proteomic analysis. We used two-dimensional gel electrophoresis with in-gel digestion followed by matrix-assisted laser desorption/ionization-time of flight analysis, to identify metabolic enzymes. Furthermore, we compared metabolic proteins in amnion cells from controls with those from Down Syndrome (DS). Enzymes involved in carbohydrate handling, amino acid handling, -purine metabolism and intermediary metabolism as well as miscellaneous metabolic pathways were detected. Protein levels of several enzymes were significantly deranged in samples obtained from patients with DS. This approach, with the advantage of the concomitant determination of many enzyme proteins, may form the basis for future metabolic screens when amniocentesis is carried out.

Amniocentesis↗

Prefractionation of protein samples for proteome analysis using reversed-phase high-performance liquid chromatography.

We describe an approach for fractionating complex protein samples prior to two-dimensional gel electrophoresis using reversed-phase high-performance liquid chromatography. Whole lysates of cells and tissue were prefractionated by reversed-phase chromatography and elution with a five-step gradient of increasing acetonitrile concentrations. The proteins obtained at each step were subsequently separated by high-resolution two-dimensional gel electrophoresis (2-DE). The reproducibility of this prefractionation technique proved to be optimal for comparing 2-DE gels from two different cell states. In addition, this method is suitable for enriching low-abundance proteins barely detectable by silver staining to amounts that can be detected by Coomassie blue and further analyzed by mass spectrometry.

Animals↗

Predicting subcellular localization of proteins using machine-learned classifiers.

MOTIVATION: Identifying the destination or localization of proteins is key to understanding their function and facilitating their purification. A number of existing computational prediction methods are based on sequence analysis. However, these methods are limited in scope, accuracy and most particularly breadth of coverage. Rather than using sequence information alone, we have explored the use of database text annotations from homologs and machine learning to substantially improve the prediction of subcellular location. RESULTS: We have constructed five machine-learning classifiers for predicting subcellular localization of proteins from animals, plants, fungi, Gram-negative bacteria and Gram-positive bacteria, which are 81% accurate for fungi and 92-94% accurate for the other four categories. These are the most accurate subcellular predictors across the widest set of organisms ever published. Our predictors are part of the Proteome Analyst web-service.

Algorithms↗

The human plasma proteome: a nonredundant list developed by combination of four separate sources.

We have merged four different views of the human plasma proteome, based on different methodologies, into a single nonredundant list of 1175 distinct gene products. The methodologies used were 1) literature search for proteins reported to occur in plasma or serum; 2) multidimensional chromatography of proteins followed by two-dimensional electrophoresis and mass spectroscopy (MS) identification of resolved proteins; 3) tryptic digestion and multidimensional chromatography of peptides followed by MS identification; and 4) tryptic digestion and multidimensional chromatography of peptides from low-molecular-mass plasma components followed by MS identification. Of 1,175 nonredundant gene products, 195 were included in more than one of the four input datasets. Only 46 appeared in all four. Predictions of signal sequence and transmembrane domain occurrence, as well as Genome Ontology annotation assignments, allowed characterization of the nonredundant list and comparison of the data sources. The "nonproteomic" literature (468 input proteins) is strongly biased toward signal sequence-containing extracellular proteins, while the three proteomics methods showed a much higher representation of cellular proteins, including nuclear, cytoplasmic, and kinesin complex proteins. Cytokines and protein hormones were almost completely absent from the proteomics data (presumably due to low abundance), while categories like DNA-binding proteins were almost entirely absent from the literature data (perhaps unexpected and therefore not sought). Most major categories of proteins in the human proteome are represented in plasma, with the distribution at successively deeper layers shifting from mostly extracellular to a distribution more like the whole (primarily cellular) proteome. The resulting nonredundant list confirms the presence of a number of interesting candidate marker proteins in plasma and serum.

Biomarkers, Tumor↗

Exploring candidate genes for human brain diseases from a brain-specific gene network.

It is believed that large numbers of genes are involved in common human brain diseases. Here, we propose a novel computational strategy for simultaneously identifying multiple candidate genes for genetic human brain diseases from a brain-specific gene network-level perspective. By integrating diverse genomic and proteomic datasets based on Bayesian statistical model, we built a large-scale human brain-specific gene network. Based on this network and minor prior knowledge of a specific brain disease, we can effectively identify multiple candidate genes for this disease. When four known Alzheimer's disease genes were used as the prior knowledge, among the top 46 high-scoring genes that we have found, 37 were previously reported to be associated with Alzheimer's disease. And the higher score a gene has, the more likely this gene is a disease-related one. The results suggest that the proposed method is effective, convenient, and applicable in the future genetic studies.

Alzheimer Disease↗

Proteomics: advanced technology for the analysis of cellular function.

Proteomics developed initially from the decade-long study of comprehensive protein visualization on two-dimensional electrophoresis gels has been expanded by mass spectrometry and the growth in searchable sequence databases. Currently, by use of more sophisticated technology such as a combination of multidimensional chromatography and mass spectrometry, thousands of proteins can automatically be identified in a day along with semiquantitative information on differential-protein expression. As with differential gene expression by cDNA-chips, the differential-protein analysis is useful for monitoring and identifying proteins involved in various physiological changes in cells or organisms, although the analysis alone does not necessarily provide information regarding the cause of the change or the function of the proteins. However, proteomics also provides the tools to expand into more sophisticated biochemical approaches, such as the study of protein interactions that can be determined directly by performing a pull-down assay with a bait protein followed by mass spectrometric identification of the bound proteins. Proteomics, thus, is useful for both large-scale surveys of proteins and detailed studies of the functional relationships among the proteins of interest. Certainly this approach can be applicable to the assessment of amino acid adequacy and safety.

Animals↗