Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.

BACKGROUND: The number of protein sequences deriving from genome sequencing projects is outpacing our knowledge about the function of these proteins. With the gap between experimentally characterized and uncharacterized proteins continuing to widen, it is necessary to develop new computational methods and tools for functional prediction. Knowledge of catalytic sites provides a valuable insight into protein function. Although many computational methods have been developed to predict catalytic residues and active sites, their accuracy remains low, with a significant number of false positives. In this paper, we present a novel method for the prediction of catalytic sites, using a carefully selected, supervised machine learning algorithm coupled with an optimal discriminative set of protein sequence conservation and structural properties. RESULTS: To determine the best machine learning algorithm, 26 classifiers in the WEKA software package were compared using a benchmarking dataset of 79 enzymes with 254 catalytic residues in a 10-fold cross-validation analysis. Each residue of the dataset was represented by a set of 24 residue properties previously shown to be of functional relevance, as well as a label {+1/-1} to indicate catalytic/non-catalytic residue. The best-performing algorithm was the Sequential Minimal Optimization (SMO) algorithm, which is a Support Vector Machine (SVM). The Wrapper Subset Selection algorithm further selected seven of the 24 attributes as an optimal subset of residue properties, with sequence conservation, catalytic propensities of amino acids, and relative position on protein surface being the most important features. CONCLUSION: The SMO algorithm with 7 selected attributes correctly predicted 228 of the 254 catalytic residues, with an overall predictive accuracy of more than 86%. Missing only 10.2% of the catalytic residues, the method captures the fundamental features of catalytic residues and can be used as a "catalytic residue filter" to facilitate experimental identification of catalytic residues for proteins with known structure but unknown function.

Algorithms↗

Bcipep: a database of B-cell epitopes.

BACKGROUND: Bcipep is a database of experimentally determined linear B-cell epitopes of varying immunogenicity collected from literature and other publicly available databases. RESULTS: The current version of Bcipep database contains 3031 entries that include 763 immunodominant, 1797 immunogenic and 471 null-immunogenic epitopes. It covers a wide range of pathogenic organisms like viruses, bacteria, protozoa, and fungi. The database provides a set of tools for the analysis and extraction of data that includes keyword search, peptide mapping and BLAST search. It also provides hyperlinks to various databases such as GenBank, PDB, SWISS-PROT and MHCBN. CONCLUSION: A comprehensive database of B-cell epitopes called Bcipep has been developed that covers information on epitopes from a wide range of pathogens. The Bcipep will be source of information for investigators involved in peptide-based vaccine design, disease diagnosis and research in allergy. It should also be a promising data source for the development and evaluation of methods for prediction of B-cell epitopes. The database is available at http://www.imtech.res.in/raghava/bcipep.

B-Lymphocytes↗

Bioinformatics in neurosurgery.

WITH THE COMPLETION of the Human Genome Project, the amount of molecular biological sequence data available in public databases has reached staggering proportions. Data continue to accumulate at an exponential rate in the postgenomic era. Compilation, storage, searching, sharing, studying, and transmitting of all these data present formidable challenges. To keep pace with this extant database, the science of bioinformatics (sometimes called computational biology) has evolved. Bioinformatics is the combination of biology and computers and usually involves the storage or analysis of molecular biological sequence data at either the deoxyribonucleic acid, ribonucleic acid, or protein (amino acid) level. Most bioinformatics tools are freely available on the Internet for use by investigators around the globe. The collective wisdom from bioinformatics databases worldwide will continue to spawn advances in the neurological sciences for generations to come. Neurosurgeons must be aware of the power and potential applications of bioinformatics for the analysis of neurosurgical diseases.

Animals↗

Case study: data management strategies in an integrated pathway tool.

This paper describes the development strategies for an integrated tool to support scientists in the creative exploration of data relating to biochemical pathways. The multiple user groups, diverse functionalities, and many types and sources of data demanded a flexible yet coherent approach. This paper summarises the software requirements and the implied modules and functions, and focuses on the design decisions relevant to the representation, management and flow of data. Finally, several case studies in the use of the software are described and evaluated, and recommendations are made for future work.

Animals↗

Proteome analysis of resting human neutrophils.

Neutrophils constitute the first line of host defense against pathogens. In the present study 2-D gel electrophoresis-mass spectrometry technology was employed to analyze the human resting neutrophils proteome. One hundred and two conserved spots were subjected to peptide mass fingerprinting, yielding 22 identifications. Among the identified proteins, nine are related to the inflammatory process, two polypeptides are assigned to metabolic functions and five are classified as structural.

Adult↗

Genomic pathways to antifungal discovery.

The limitations of the therapeutic antifungals are becoming increasingly apparent in the clinic due to their modest efficacy against life-threatening systemic fungal infections. These antifungals belong to only a few structural classes that affect a small range of targets, some are quite toxic in humans while the use of others, particularly the azole drugs, has encouraged the emergence of resistant clinical isolates and the selection of innately resistant fungal pathogens. Only a few new drugs based on novel targets are in clinical development, and these may be insufficient to overcome the changing tide of fungal disease. In parallel with the successful completion of the Saccharomyces cerevisiae and human genome sequencing projects, an increasing number of genome sequencing projects are being initiated and completed for significant fungal pathogens. The growing repository of genomic information, which is complemented by decades of genetic and biochemical study, is now available for genome-wide analysis of gene function and for incisive inter-genomic comparison, with the S. cerevisiae and human genomes providing key points of reference. Functional genomic and comparative genomic techniques, many of which were developed with S. cerevisiae, are being applied to fungal pathogens with the aim of obtaining an integrated view of fungal biology and to extract targets suitable for drug discovery. This review describes some of these techniques, their limitations and their increasing contribution to the antifungal discovery process through effective gene annotation, target identification and prioritization, and in the optimization of antifungal leads.

Antifungal Agents↗

Multimeric threading-based prediction of protein-protein interactions on a genomic scale: application to the Saccharomyces cerevisiae proteome.

MULTIPROSPECTOR, a multimeric threading algorithm for the prediction of protein-protein interactions, is applied to the genome of Saccharomyces cerevisiae. Each possible pairwise interaction among more than 6000 encoded proteins is evaluated against a dimer database of 768 complex structures by using a confidence estimate of the fold assignment and the magnitude of the statistical interfacial potentials. In total, 7321 interactions between pairs of different proteins are predicted, based on 304 complex structures. Quality estimation based on the coincidence of subcellular localizations and biological functions of the predicted interactors shows that our approach ranks third when compared with all other large-scale methods. Unlike other in silico methods, MULTIPROSPECTOR is able to identify the residues that participate directly in the interaction. Three hundred seventy-four of our predictions can be found by at least one of the other studies, which is compatible with the overlap between two different other methods. From the analysis of the mRNA abundance data, our method does not bias towards proteins with high abundance. Finally, several relevant predictions involved in various functions are presented. In summary, we provide a novel approach to predict protein-protein interactions on a genomic scale that is a useful complement to experimental methods.

DNA, Fungal↗

Towards the proteome of Brassica napus phloem sap.

The soluble proteins in sieve tube exudate from Brassica napus plants were systematically analyzed by 1-DE and high-resolution 2-DE, partial amino acid sequence determination by MS/MS, followed by database searches. 140 proteins could be identified by their high similarity to database sequences (135 from 2-DE, 5 additional from 1-DE). Most analyzed spots led to successful protein identifications, demonstrating that Brassica napus, a close relative of Arabidopsis thaliana, is a highly suitable model plant for phloem research. None of the identified proteins was formerly known to be present in Brassica napus phloem, but several proteins have been described in phloem sap of other species. The data, which is discussed with respect to possible physiological importance of the proteins in the phloem, further confirms and substantially extends earlier findings and uncovers the presence of new protein functions in the vascular system. For example, we found several formerly unknown phloem proteins that are potentially involved in signal generation and transport, e.g., proteins mediating calcium and G-protein signaling, a set of RNA-binding proteins, and FLOWERING LOCUS T (FT) and its twin sister that might be key components for the regulation of flowering time.

Brassica napus↗

Coverage of protein sequence space by current structural genomics targets.

By its purest definition the ultimate goal of structural genomics (SG) is the determination of the structures of all proteins encoded by genomes. Most of these will be obtained by homology modeling using the structures of a set of target proteins for experimental determination. Thanks to the open exchange of SG target information, we are able to analyze the sequences of the current target list to evaluate the extent of its coverage of protein sequence space. The presence of homologous sequences currently either in the Protein Data Bank (PDB) or among SG targets has been determined for each of the protein sequences in several organisms. In this way we are able to evaluate the coverage by existing or targeted structural data for the non-membranous parts of entire proteomes. For small bacterial proteomes such as that of H. influenzae almost all proteins have homologous sequences among SG targets or in the PDB. There is significantly lower coverage for more complex organisms, such as C. elegans. We have mapped the SG target list onto the ProtoMap clustering of protein sequences. Clusters occupied by SG targets represent over 150,000 protein sequences, which is approximately 44% of the total protein sequences classified by ProtoMap. The mapping of SG targets also enables an evaluation of the degree of overlap within the target list. An SG target typically occupies a ProtoMap cluster with more than six other homologous targets.

Algorithms↗

Using proteomics to mine genome sequences.

We present a method for mining unannotated or annotated genome sequences with proteomic data to identify open reading frames. The region of a genome coding for a protein sequence is identified by using information from the analysis of proteins and peptides with MALDI-TOF mass spectrometry. The raw genome sequence or any unassembled contigs of an organism are theoretically cleaved into a number of equal sized but overlapping fragments, and these are then translated in all six frames into a series of virtual proteins. Each virtual protein is then subjected to a theoretical enzymatic digestion. Standard proteomic sample preparation methods are used to separate, array, and digest the proteins of interest to peptides. The masses of the resulting peptides are measured using mass spectrometry and compared to the theoretical peptide masses of the virtual proteins. The region of the genome responsible for coding for a particular protein can then be identified when there are a large number of hits between peptides from the protein and peptides from the virtual protein. The method makes no assumptions about the location of a protein in a particular gene sequence or the positions or types of start and stop codons. To illustrate this approach, all 773 proteins of Pseudomonas aeruginosa contained in SWISS-PROT were used to theoretically test the method and optimize parameters. Increasing the size of the virtual proteins results in an overall improvement in the ability to detect the coding region, at the cost of decreasing the sensitivity of the method for smaller proteins. Increasing the minimum number of matching peptides, lowering the mass error tolerance, or increasing the signal-to-noise ratio of the simulated mass spectrum, improves the ability to detect coding regions. The method is further demonstrated on experimental data from Mycobacterium tuberculosis and is also shown to work with eukaryotic organisms (e.g., Homo sapiens).

Amino Acid Sequence↗

Proteomic characterization of human normal articular chondrocytes: a novel tool for the study of osteoarthritis and other rheumatic diseases.

Articular cartilage is composed of cells and an extracellular matrix. The chondrocyte is the only cell type present in mature cartilage, and it is important in the control of cartilage integrity. There is currently a great lack of knowledge about the chondrocyte proteome. To solve this deficiency, we have obtained the first reference map of the human normal articular chondrocyte. Cells were isolated from cartilages obtained from autopsies without history of joint disease. Cultured cells were used to obtain protein extracts which were resolved by 2-DE and visualized by silver nitrate or CBB staining. Almost 200 spots were excised from the gels and analyzed using MALDI-TOF or MALDI-TOF/TOF MS. The analysis leads to the identification of 136 spots that represent 93 different proteins. A significant proportion of proteins are involved in cell organization (26%), energy (16%), protein fate (14%), metabolism (12%), and cell stress (12%). From all the identified proteins, annexins, vimentin, transgelin, destrin, cathepsin D, heat shock protein 47, and mitochondrial superoxide dismutase were more abundant in chondrocytes than in other types of mesenchymal cells such as Jurkat-T cells. As metabolic program of chondrocytes is altered in osteoarthritis and other rheumatic diseases, this proteomic map is an important tool for future studies on these pathologies.

Adolescent↗

A suffix tree approach to the interpretation of tandem mass spectra: applications to peptides of non-specific digestion and post-translational modifications.

MOTIVATION: Tandem mass spectrometry combined with sequence database searching is one of the most powerful tools for protein identification. As thousands of spectra are generated by a mass spectrometer in one hour, the speed of database searching is critical, especially when searching against a large sequence database, or when the peptide is generated by some unknown or non-specific enzyme, even or when the target peptides have post-translational modifications (PTM). In practice, about 70-90% of the spectra have no match in the database. Many believe that a significant portion of them are due to peptides of non-specific digestions by unknown enzymes or amino acid modifications. In another case, scientists may choose to use some non-specific enzymes such as pepsin or thermolysin for proteolysis in proteomic study, in that not all proteins are amenable to be digested by some site-specific enzymes, and furthermore many digested peptides may not fall within the rang of molecular weight suitable for mass spectrometry analysis. Interpreting mass spectra of these kinds will cost a lot of computational time of database search engines. OVERVIEW: The present study was designed to speed up the database searching process for both cases. More specifically speaking, we employed an approach combining suffix tree data structure and spectrum graph. The suffix tree is used to preprocess the protein sequence database, while the spectrum graph is used to preprocess the tandem mass spectrum. We then search the suffix tree against the spectrum graph for candidate peptides. We design an efficient algorithm to compute a matching threshold with some statistical significance level, e.g. p = 0.01, for each spectrum, and use it to select candidate peptides. Then we rank these peptides using a SEQUEST-like scoring function. The algorithms were implemented and tested on experimental data. For post-translational modifications, we allow arbitrary number of any modification to a protein. AVAILABILITY: The executable program and other supplementary materials are available online at: http://hto-c.usc.edu:8000/msms/suffix/.

Algorithms↗

Pseudomonas aeruginosa and a proteomic approach to bacterial pathogenesis.

Pseudomonas aeruginosa is a Gram-negative bacterium that is ubiquitous in the environment and can cause a variety of diseases in compromised patients. The genome of P. aeruginosa strain PAO1 has been reported to contain 5570 potential proteins. The value of this genomic database is that new proteins can be recognized to use as diagnostic markers, novel drug targets, and to better understand the physiology of this organism. However, similar to what has been observed in other sequenced bacterial genomes, approximately one third of the potential proteins have no known function. This is somewhat surprising given the long-standing interest in P. aeruginosa as an opportunistic pathogen. Obviously new tools, in addition to sequence similarity analysis, are needed to determine the role of these proteins. Proteomics using two-dimensional gel electrophoresis followed by mass spectrometry to detect and identify P. aeruginosa proteins represents a novel approach to address this gap.

Animals↗

Development of gene ontology tool for biological interpretation of genomic and proteomic data.

We have designed and developed a Gene Ontology based navigation tool, GoMiner, which organizes lists of interesting genes from a microarray or a protein array experiment for biological interpretation. It provides quantitative and statistical output files and useful visualization (e.g., a tree-like structure) to map the list of genes to its biological functional categories. It also provides links to other resources such as pubmed, locuslink, and biological molecular interaction map and signaling pathway packages.

Computational Biology↗

Bioinformatics: use in bacterial vaccine discovery.

Bioinformatics has now become a common laboratory name for groups studying genomic sequences. It is composed of many different, yet interrelated scientific fields such as genomics, proteomics, and transcriptional profiling. The availability of complete genomic sequences, especially prokaryotic organisms, allows one to rapidly identify, analyze, and clone genes of interest. For bacterial vaccine discovery, one can "mine" the genomic sequence for potential surface targets using various algorithms, characterize these gene targets, and produce primers for cloning, all before one enters the wet laboratory. This review will focus on various genomic mining tools/algorithms available for predicting open reading frames and their associated annotation (if known), physical and functional characterization, and cellular localization. Finally, examples are given of how all of this is being used for the identification of potential bacterial vaccine candidates.

Animals↗

Using the global proteome machine for protein identification.

This chapter describes the use of an open-source, freely available informatics system for the identification of proteins using tandem mass spectra of peptides derived from an enzymatic digest of a mixture of mature proteins. The chapter describes the use of features of the Global Proteome Machine (GPM) interface that assist in making comprehensive assignments between spectra and sequences, including the detection of point mutations, posttranslational modifications, and experimental artifacts. The use of this interface to validate results using the GPM Database is also described. This data repository allows analysts to compare their own results to those obtained by other scientists to determine the degree to which their data are consistent with previous measurements.

Amino Acid Sequence↗

Proteome approaches to characterize seed storage proteins related to ditelocentric chromosomes in common wheat (Triticum aestivum L.).

Changes in protein composition of wheat endosperm proteome were investigated in 39 ditelocentric chromosome lines of common wheat (Triticum aestivum L.) cv. Chinese Spring. Two-dimensional gel electrophoresis followed by Coomassie Brilliant Blue staining has resolved a total of 105 protein spots in a gel. Quantitative image analysis of protein spots was performed by PDQuest. Variations in protein spots between the euploid and the 39 ditelocentric lines were evaluated by spot number, appearance, disappearance and intensity. A specific spot present in all gels was taken as an internal standard, and the intensity of all other spots was calculated as the ratio of the internal standard. Out of the 1755 major spots detected in 39 ditelocentric lines, 1372 (78%) spots were found variable in different spot parameters: 147 (11%) disappeared, 978 (71%) up-regulated and 247 (18%) down-regulated. Correlation studies in changes in protein intensities among 24 protein spots across the ditelocentric lines were performed. High correlations in changes of protein intensities were observed among the proteins encoded by genes located in the homoeologous arms. Locations of structural genes controlling 26 spots were identified in 10 chromosomal arms. Multiple regulators of the same protein located at various chromosomal arms were also noticed. Identification of structural genes for most of the proteins was found difficult due to multiple regulators encoding the same protein. Two novel subunits (1B(Z,) 1BDz), the structure of which are very similar to the high molecular weight glutenin subunit 12, were identified, and the chromosome arm locations of these subunits were assigned.

Chromosomes↗