Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

[Identification of proteome molecules by proteomics using two-dimensional gel electrophoresis and MALDI-TOF MS].

Genomic technologies have enabled rapid accumulation of information from complex biological systems over the last two decades. The complete DNA sequence is now known for many organisms and the informational database obtained from genome sequencing projects has provided the base for the specification of proteome - the protein complement of genome. Genomic functions can be inferred from the analysis of gene structure and gene expression profiles because proteins are the functional molecules of an organism. Integrated technologies including protein separation, identification, characterization and information manage system are essential to analyze the proteins in complex cellular matrix. This study is focusing on the strategies of proteome analysis using sample preparation, 2-dimensional gel electrophoresis, processing of protein spots and identification of proteins, protein-protein interaction and posttranslational modification using MALDI-TOF-MS. 2-D gel electrophoresis is currently the most powerful protein separation technique and MALDI-TOF MS is powerful identification technique for protein and peptides as a sensitive, rapid, and high resolution analytical method. The developed integrated proteome technologies are very useful to understand the biological phenomena at molecular level by identifying the new molecules and their modifications in various cellular processes, and can be applied for biotechnology including medical science.

Databases, Protein↗

The importance of intrinsic disorder for protein phosphorylation.

Reversible protein phosphorylation provides a major regulatory mechanism in eukaryotic cells. Due to the high variability of amino acid residues flanking a relatively limited number of experimentally identified phosphorylation sites, reliable prediction of such sites still remains an important issue. Here we report the development of a new web-based tool for the prediction of protein phosphorylation sites, DISPHOS (DISorder-enhanced PHOSphorylation predictor, http://www.ist.temple. edu/DISPHOS). We observed that amino acid compositions, sequence complexity, hydrophobicity, charge and other sequence attributes of regions adjacent to phosphorylation sites are very similar to those of intrinsically disordered protein regions. Thus, DISPHOS uses position-specific amino acid frequencies and disorder information to improve the discrimination between phosphorylation and non-phosphorylation sites. Based on the estimates of phosphorylation rates in various protein categories, the outputs of DISPHOS are adjusted in order to reduce the total number of misclassified residues. When tested on an equal number of phosphorylated and non-phosphorylated residues, the accuracy of DISPHOS reaches 76% for serine, 81% for threonine and 83% for tyrosine. The significant enrichment in disorder-promoting residues surrounding phosphorylation sites together with the results obtained by applying DISPHOS to various protein functional classes and proteomes, provide strong support for the hypothesis that protein phosphorylation predominantly occurs within intrinsically disordered protein regions.

Amino Acids↗

DeNovoID: a web-based tool for identifying peptides from sequence and mass tags deduced from de novo peptide sequencing by mass spectroscopy.

One of the core activities of high-throughput proteomics is the identification of peptides from mass spectra. Some peptides can be identified using spectral matching programs like Sequest or Mascot, but many spectra do not produce high quality database matches. De novo peptide sequencing is an approach to determine partial peptide sequences for some of the unidentified spectra. A drawback of de novo peptide sequencing is that it produces a series of ordered and disordered sequence tags and mass tags rather than a complete, non-degenerate peptide amino acid sequence. This incomplete data is difficult to use in conventional search programs such as BLAST or FASTA. DeNovoID is a program that has been specifically designed to use degenerate amino acid sequence and mass data derived from MS experiments to search a peptide database. Since the algorithm employed depends on the amino acid composition of the peptide and not its sequence, DeNovoID does not have to consider all possible sequences, but rather a smaller number of compositions consistent with a spectrum. DeNovoID also uses a geometric indexing scheme that reduces the number of calculations required to determine the best peptide match in the database. DeNovoID is available at http://proteomics.mcw.edu/denovoid.

Algorithms↗

New perspectives on host-parasite interplay by comparative transcriptomic and proteomic analyses of Schistosoma japonicum.

Schistosomiasis remains a serious public health problem with an estimated 200 million people infected in 76 countries. Here we isolated ~ 8,400 potential protein-encoding cDNA contigs from Schistosoma japonicum after sequencing circa 84,000 expressed sequence tags. In tandem, we undertook a high-throughput proteomics approach to characterize the protein expression profiles of a number of developmental stages (cercariae, hepatic schistosomula, female and male adults, eggs, and miracidia) and tissues at the host-parasite interface (eggshell and tegument) by interrogating the protein database deduced from the contigs. Comparative analysis of these transcriptomic and proteomic data, the latter including 3,260 proteins with putative identities, revealed differential expression of genes among the various developmental stages and sexes of S. japonicum and localization of putative secretory and membrane antigens, enzymes, and other gene products on the adult tegument and eggshell, many of which displayed genetic polymorphisms. Numerous S. japonicum genes exhibited high levels of identity with those of their mammalian hosts, whereas many others appeared to be conserved only across the genus Schistosoma or Phylum Platyhelminthes. These findings are expected to provide new insights into the pathophysiology of schistosomiasis and for the development of improved interventions for disease control and will facilitate a more fundamental understanding of schistosome biology, evolution, and the host-parasite interplay.

Amino Acid Sequence↗

Structure-based functional discovery of proteins: structural proteomics.

The discovery of biochemical and cellular functions of unannotated gene products begins with a database search of proteins with structure/sequence homologues based on known genes. Very recently, a number of frontier groups in structural biology proposed a new paradigm to predict biological functions of an unknown protein on the basis of its three-dimensional structure on a genomic scale. Structural proteomics (genomics), a research area for structure-based functional discovery, aims to complete the protein-folding universe of all gene products in a cell. It would lead us to a complete understanding of a living organism from protein structure. Two major complementary experimental techniques, X-ray crystallography and NMR spectroscopy, combined with recently developed high throughput methods have played a central role in structural proteomics research; however, an integration of these methodologies together with comparative modeling and electron microscopy would speed up the goal for completing a full dictionary of protein folding space in the near future.

Animals↗

How many nuclear hormone receptors are there in the human genome?

The sequence of the human genome now allows the definition of the complete set of genes for specific protein families in humans. Because of their involvement in many physiological and pathological processes, the nuclear hormone receptors are a superfamily of crucial medical significance. Although 48 human nuclear receptor genes were identified previously, their total number is unclear from early human genome reports. Here, we report the identification and classification of all nuclear receptor genes in the human genome, and we discuss corresponding transcriptome and proteome diversity.

Alternative Splicing↗

Genome-based proteomics.

Protein-protein interactions play crucial roles in various biological pathways and functions. Therefore, the characterization of protein levels and also the network of interactions within an organism would contribute considerably to the understanding of life. The availability of the human genome sequence has created a range of new possibilities for biomedical research. A crucial challenge is to utilize the genetic information for better understanding of protein distribution and function in normal as well as in pathological biological processes. In this review, we have focused on different platforms used for systematic genome-based proteome analyses. These technologies are in many ways complementary and should be seen as various ways to elucidate different functions of the proteome.

Chromatography, Liquid↗

Human Proteome Organisation Proteomics Standards Initiative. Pre-Congress Initiative.

The plenary session of the Proteomics Standards Initiative of the Human Proteome Organisation discussed the current status of the ongoing work in the fields of molecular interactions, mass spectrometry and the description of protein modifications. In addition, new areas are being opened up, in particular developing standards for the description and exchange of data from gel electrophoresis experiments. The General Proteomics Standards group is now working closely with the Functional Genomics Experiment efforts to define a general standard in which to encode data that will enable a systems biology approach to data analysis.

Databases, Genetic↗

Probability-based evaluation of peptide and protein identifications from tandem mass spectrometry and SEQUEST analysis: the human proteome.

Large-scale protein identifications from highly complex protein mixtures have recently been achieved using multidimensional liquid chromatography coupled with tandem mass spectrometry (LC/LC-MS/MS) and subsequent database searching with algorithms such as SEQUEST. Here, we describe a probability-based evaluation of false positive rates associated with peptide identifications from three different human proteome samples. Peptides from human plasma, human mammary epithelial cell (HMEC) lysate, and human hepatocyte (Huh)-7.5 cell lysate were separated by strong cation exchange (SCX) chromatography coupled offline with reversed-phase capillary LC-MS/MS analyses. The MS/MS spectra were first analyzed by SEQUEST, searching independently against both normal and sequence-reversed human protein databases, and the false positive rates of peptide identifications for the three proteome samples were then analyzed and compared. The observed false positive rates of peptide identifications for human plasma were significantly higher than those for the human cell lines when identical filtering criteria were used, suggesting that the false positive rates are significantly dependent on sample characteristics, particularly the number of proteins found within the detectable dynamic range. Two new sets of filtering criteria are proposed for human plasma and human cell lines, respectively, to provide an overall confidence of >95% for peptide identifications. The new criteria were compared, using a normalized elution time (NET) criterion (Petritis et al. Anal. Chem. 2003, 75, 1039-1048), with previously published criteria (Washburn et al. Nat. Biotechnol. 2001, 19, 242-247). The results demonstrate that the present criteria provide significantly higher levels of confidence for peptide identifications from mammalian proteomes without greatly decreasing the number of identifications.

Blood Proteins↗

Integrating a functional proteomic approach into the target discovery process.

Functional proteomics is a promising technique for the rational identification of novel therapeutic targets by elucidation of the function of newly identified proteins in disease-relevant cellular pathways. Of the recently described high-throughput approaches for analyzing protein-protein interactions, the yeast two-hybrid (Y2H) system has turned out to be one of the most suitable for genome-wide analysis. However, this system presents a challenging technical problem: the high prevalence of false positives and false negatives in datasets due to intrinsic limitations of the technology and the use of a high-throughput, genetic assay. We discuss here the different experimental strategies applied to Y2H assays, their general limitations and advantages. We also address the issue of the contribution of protein interaction mapping to functional biology, especially when combined with complementary genomic and proteomic analyses. Finally, we illustrate how the combination of protein interaction maps with relevant functional assays can provide biological support to large-scale protein interaction datasets and contribute to the identification and validation of potential therapeutic targets.

Animals↗

Large-scale database searching using tandem mass spectra: looking up the answer in the back of the book.

Database searching is an essential element of large-scale proteomics. Because these methods are widely used, it is important to understand the rationale of the algorithms. Most algorithms are based on concepts first developed in SEQUEST and PeptideSearch. Four basic approaches are used to determine a match between a spectrum and sequence: descriptive, interpretative, stochastic and probability-based matching. We review the basic concepts used by most search algorithms, the computational modeling of peptide identification and current challenges and limitations of this approach for protein identification.

Algorithms↗

The UCSC Proteome Browser.

The University of California Santa Cruz (UCSC) Proteome Browser provides a wealth of protein information presented in graphical images and with links to other protein-related Internet sites. The Proteome Browser is tightly integrated with the UCSC Genome Browser. For the first time, Genome Browser users have both the genome and proteome worlds at their fingertips simultaneously. The Proteome Browser displays tracks of protein and genomic sequences, exon structure, polarity, hydrophobicity, locations of cysteine and glycosylation potential, Superfamily domains and amino acids that deviate from normal abundance. Histograms show genome-wide distribution of protein properties, including isoelectric point, molecular weight, number of exons, InterPro domains and cysteine locations, together with specific property values of the selected protein. The Proteome Browser also provides links to gene annotations in the Genome Browser, the Known Genes details page and the Gene Sorter; domain information from Superfamily, InterPro and Pfam; three-dimensional structures at the Protein Data Bank and ModBase; and pathway data at KEGG, BioCarta/CGAP and BioCyc. As of August 2004, the Proteome Browser is available for human, mouse and rat proteomes. The browser may be accessed from any Known Genes details page of the Genome Browser at http://genome.ucsc.edu. A user's guide is also available on this website.

California↗

Identification of novel human genes evolutionarily conserved in Caenorhabditis elegans by comparative proteomics.

Modern biomedical research greatly benefits from large-scale genome-sequencing projects ranging from studies of viruses, bacteria, and yeast to multicellular organisms, like Caenorhabditis elegans. Comparative genomic studies offer a vast array of prospects for identification and functional annotation of human ortholog genes. We presented a novel comparative proteomic approach for assembling human gene contigs and assisting gene discovery. The C. elegans proteome was used as an alignment template to assist in novel human gene identification from human EST nucleotide databases. Among the available 18,452 C. elegans protein sequences, our results indicate that at least 83% (15,344 sequences) of C. elegans proteome has human homologous genes, with 7,954 records of C. elegans proteins matching known human gene transcripts. Only 11% or less of C. elegans proteome contains nematode-specific genes. We found that the remaining 7,390 sequences might lead to discoveries of novel human genes, and over 150 putative full-length human gene transcripts were assembled upon further database analyses. [The sequence data described in this paper have been submitted to the

Amino Acid Sequence↗

Proteomics: state of the art and its application in cardiovascular research.

The cellular and molecular mechanisms underlying cardiovascular dysfunctions are widely unknown. Basically, pathological changes in the cardiovascular system arise from protein alterations. Proteomics comprises a set of tools for the large-scale study of gene expression at the protein level thereby allowing for the identification of protein alterations responsible for the development and the pathological outcome of diseases including those of the cardiovascular system. In principle these alterations include those of suitable candidates for drug targets and disease biomarkers as well as therapeutic proteins/peptides. Since gene therapy depends on the function of a therapeutic protein encoded by a "therapeutic" gene proteomic analyses also provide the basis for the design and application of gene therapies. Proteomic technologies allow to identify not only proteins but also the nature of their posttranslational modifications thus enabling the elucidation of signal transduction pathways and their deregulation under pathological conditions. The linkage of information about proteome changes with functional consequences lead to the development of functional proteomic studies. Functional proteomic analyses will particularly help to better understand the relations between proteome changes and cardiovascular dysfunctions. The storage and administration of experimental data obtained by the application of proteomic analyses is supported by species- and tissue-specific protein databases and specific software. Publications in this field are reviewed in this paper.

Animals↗

Predicting protease types by hybridizing gene ontology and pseudo amino acid composition.

Proteases play a vitally important role in regulating most physiological processes. Different types of proteases perform different functions with different biological processes. Therefore, it is highly desired to develop a fast and reliable means to identify the types of proteases according to their sequences, or even just identify whether they are proteases or nonproteases. The avalanche of protein sequences generated in the postgenomic era has made such a challenge become even more critical and urgent. By hybridizing the gene ontology approach and pseudo amino acid composition approach, a powerful predictor called GO-PseAA predictor was introduced to address the problems. To avoid redundancy and bias, demonstrations were performed on a dataset where none of proteins has >/= 25% sequence identity to any other. The overall success rates thus obtained by the jackknife cross-validation test in identifying protease and nonprotease was 91.82%, and that in identifying the protease type was 85.49% among the following five types: (1) aspartic, (2) cysteine, (3) metallo, (4) serine, and (5) threonine. The high jackknife success rates yielded for such a stringent dataset indicate the GO-PseAA predictor is very powerful and might become a useful tool in bioinformatics and proteomics.

Amino Acids↗

A new algorithm for the evaluation of shotgun peptide sequencing in proteomics: support vector machine classification of peptide MS/MS spectra and SEQUEST scores.

Shotgun tandem mass spectrometry-based peptide sequencing using programs such as SEQUEST allows high-throughput identification of peptides, which in turn allows the identification of corresponding proteins. We have applied a machine learning algorithm, called the support vector machine, to discriminate between correctly and incorrectly identified peptides using SEQUEST output. Each peptide was characterized by SEQUEST-calculated features such as delta Cn and Xcorr, measurements such as precursor ion current and mass, and additional calculated parameters such as the fraction of matched MS/MS peaks. The trained SVM classifier performed significantly better than previous cutoff-based methods at separating positive from negative peptides. Positive and negative peptides were more readily distinguished in training set data acquired on a QTOF, compared to an ion trap mass spectrometer. The use of 13 features, including four new parameters, significantly improved the separation between positive and negative peptides. Use of the support vector machine and these additional parameters resulted in a more accurate interpretation of peptide MS/MS spectra and is an important step toward automated interpretation of peptide tandem mass spectrometry data in proteomics.

Algorithms↗

ProteomeCommons.org JAF: reference information and tools for proteomics.

SUMMARY: Analysis of proteomics data, specifically mass spectrometry data, commonly relies on libraries of known information such as atomic masses, known stable isotopes, atomic compositions of amino acids, observed modifications of known amino acids and ion masses that directly correspond to known amino acid sequences. The Java Analysis Framework (JAF) for proteomics provides a freely usable, open-source library of Java code that abstracts all of the aforementioned data, enabling more rapid development of proteomics tools. The JAF also includes several user tools that can be run directly from a web browser. AVAILABILITY: The current version and an archive of all older versions of the Java Analysis Framework for Proteomics is freely available, including complete source-code, at http://www.proteomecommons.org/current/511/.

Database Management Systems↗