Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Characterisation of organellar proteomes: a guide to subcellular proteomic fractionation and analysis.

Subcellular fractionation is being widely used to increase our understanding of the proteome. Fractionation is often coupled with 2-DE, thus allowing the visualisation of proteins and their subsequent identification and characterisation by MS. Whilst this strategy should be effective, to date, there has been little or no consideration given to differences in the mass, pI, hydropathy or abundance of proteins in the organelles and how analytical strategies can be tailored to match the idiosyncrasies of proteins in each particular compartment. To address this, we analysed 3962 Saccharomyces cerevisiae proteins, previously localised to one or more of 22 subcellular compartments. Different compartments showed significantly different distributions of protein pI and hydropathy. Mitochondrial and ER proteins showed the most dramatic differences to other organelles, in their protein pIs and hydropathy, respectively. We show that organelles can be clustered by similarities in these physicochemical protein characteristics. Interestingly, the distribution of protein abundance was also significantly different between many organelles. Our results show that to fully explore subcellular fractions of the proteome, specific analytical strategies should be employed. We outline strategies for all 22 subcellular compartments.

Cell Compartmentation↗

Two-dimensional nano-liquid chromatography-mass spectrometry system for applications in proteomics.

This work demonstrates the development of a method for the analysis of complex proteome samples by two-dimensional nano-liquid chromatography-mass spectrometry. This approach includes strong cation-exchange, sample enrichment, reversed-phase chromatography and nanospray ion trap mass spectroscopy with data dependent tandem mass spectrometry spectra acquisition, and subsequent database search. The new methodology was first evaluated using standard protein digest samples. Finally, data for the analysis of a total Escherichia coli proteome are provided.

Amino Acid Sequence↗

Global analysis of the membrane subproteome of Pseudomonas aeruginosa using liquid chromatography-tandem mass spectrometry.

Pseudomonas aeruginosa is one of the most significant opportunistic bacterial pathogens in humans causing infections and premature death in patients with cystic fibrosis, AIDS, severe burns, organ transplants, or cancer. Liquid chromatography coupled online with tandem mass spectrometry was used for the large-scale proteomic analysis of the P. aeruginosa membrane subproteome. Concomitantly, an affinity labeling technique, using iodoacetyl-PEO biotin to tag cysteinyl-containing proteins, permitted the enrichment and detection of lower abundance membrane proteins. The application of these approaches resulted in the identification of 786 proteins. A total of 333 proteins (42%) had a minimum of one transmembrane domain (ranging from 1 to14) and 195 proteins were classified as hydrophobic based on their positive GRAVY values (ranging from 0.01 to 1.32). Key integral inner and outer membrane proteins involved in adaptation and antibiotic resistance were conclusively identified, including the detection of 53% of all predicted opr-type porins (outer integral membrane proteins) and all the components of the mexA-mexB-oprM transmembrane protein complex. This work represents one of the most comprehensive proteomic analyses of the membrane subproteome of P. aeruginosa and for prokaryotes in general.

Affinity Labels↗

Using GO-PseAA predictor to identify membrane proteins and their types.

Cell membranes are crucial to the life of a cell. Although the basic structure of biological membrane is provided by the lipid bilayer, most of the specific functions are carried out by membrane proteins. Knowledge of membrane protein type often offers important clues toward determining the function of an uncharacterized protein. Therefore, predicting the type of a membrane protein from its primary sequence, or even just identifying whether the uncharacterized protein belongs to a membrane protein or not, is an important and challenging problem in bioinformatics and proteomics. To deal with these problems, the GO-PseAA predictor is introduced that is operated in a hybridization space by combining the gene ontology and pseudo amino acid composition. Meanwhile, to test the prediction quality, a dataset was constructed that contains 6476 non-membrane proteins and 5122 membrane proteins classified into five different types. To avoid redundancy and bias, none of the proteins included has > or = 40% sequence identity to any other. It has been observed that the overall success rate by the jackknife cross-validation test in identifying non-membrane proteins and membrane proteins was 94.76%, and that in identifying the five membrane protein types was 95.84%. The high success rates suggest that the GO-PseAA predictor can catch the core feature of the statistical samples concerned and may become an automated high throughput toll in molecular and cell biology.

Algorithms↗

Approach to systematic analysis of serine/threonine phosphoproteome using Beta elimination and subsequent side effects: intramolecular linkage and/or racemisation.

Complete analysis of the phosphorylation of serine and threonine residues directly from biological extracts is still at an early stage and will remain a challenging goal for many years. Analysis of phosphorylated proteins and identification of the phosphorylated sites in a crude biological extract is a major topic in proteomics, since phosphorylation plays a dominant role in post-translational protein modification. Beta elimination of the serine/threonine-bound phosphate by alkali action generates (methyl)dehydroalanine. The reactivity of this group susceptible of nucleophilic attacks might be used as a tool for phosphoproteome analysis. Most of the known serine/threonine kinases recognize motifs in protein targets that are rich in lysine(s) and/or arginine(s). The (methyl)dehydroalanine resulting from beta elimination of the serine/threonine-bound phosphate by alkali action is likely to react with the amino groups of these neighboring amino acids. Furthermore, the addition reaction of dehydroalanine-peptides with a nucleophilic group more likely generates diastereoisomers derivatives. The internal cyclic bonds and/or the stereoisomer peptide derivatives thus generated confer resistance to trypsin cleavage and/or constitute stop signals for exopeptidases such as carboxypeptidase. This might form the basis of a method to facilitate the systematic identification of phosphorylated peptides.

Amino Acid Sequence↗

Chromatographic separations as a prelude to two-dimensional electrophoresis in proteomics analysis.

Current methods of proteome analysis rely almost solely on two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) followed by the excision of individual spots and protein identification using mass spectrometry (MS) and database searching. 2-D PAGE is denaturing in both dimensions and, thus, cannot indicate functional associations between individual proteins. Moreover, less abundant proteins are difficult to identify. To simplify the proteome, and explore functional associations, nondenaturing anion exchange column chromatography was used to separate a soluble protein extract from Escherichia coli. Successive fractions were then analysed using 2-D PAGE and selected spots from both the gels for the start material and the fractionated material were quantified and identified by peptide mass fingerprinting using a MALDI-TOF mass spectrometer. Enrichments of up to 13-fold were attained for individual protein spots and peptide mass fingerprints were of significantly higher quality after chromatographic separation. The marked anomalies between predicted p/and column elution position contrasted with the almost perfect correlation with migration distance on isoelectric focusing (IEF) and were explored further for basic proteins.

Chromatography, Ion Exchange↗

Mass spectrometry allows direct identification of proteins in large genomes.

Proteome projects seek to provide systematic functional analysis of the genes uncovered by genome sequencing initiatives. Mass spectrometric protein identification is a key requirement in these studies but to date, database searching tools rely on the availability of protein sequences derived from full length cDNA, expressed sequence tags or predicted open reading frames (ORFs) from genomic sequences. We demonstrate here that proteins can be identified directly in large genomic databases using peptide sequence tags obtained by tandem mass spectrometry. On the background of vast amounts of noncoding DNA sequence, identified peptides localize coding sequences (exons) in a confined region of the genome, which contains the cognate gene. The approach does not require prior information about putative ORFs as predicted by computerized gene finding algorithms. The method scales to the complete human genome and allows identification, mapping, cloning and assistance in gene prediction of any protein for which minimal mass spectrometric information can be obtained. Several novel proteins from Arabidopsis thaliana and human have been discovered in this way.

Amino Acid Sequence↗

Using functional domain composition to predict enzyme family classes.

According to their main EC (Enzyme Commission) numbers, enzymes are classified into the following 6 main classes: oxidoreductases, transferases, hydrolases, lyases, isomerases, and ligases. A new method has been developed to predict the enzymatic attribute of proteins by introducing the functional domain composition to formulate a given protein sequence. The advantage by doing so is that both the sequence-order-related features and the function-related features are naturally incorporated in the predictor. As a demonstration, the jackknife cross-validation test was performed on a dataset that consists of proteins with only less than 20% sequence identity to each other in order to get rid of any homologous bias. The overall success rate thus obtained was 85% in identifying the enzyme family classes (including the identification of nonenzyme protein sequences as well). The success rate is significantly higher than those obtained by the other methods on such a stringent dataset. This indicates that using the functional domain composition to represent protein samples for statistical prediction is indeed very promising, and will become a powerful tool in bioinformatics and proteomics.

Algorithms↗

Automatic quality assessment of peptide tandem mass spectra.

MOTIVATION: A powerful proteomics methodology couples high-performance liquid chromatography (HPLC) with tandem mass spectrometry and database-search software, such as SEQUEST. Such a set-up, however, produces a large number of spectra, many of which are of too poor quality to be useful. Hence a filter that eliminates poor spectra before the database search can significantly improve throughput and robustness. Moreover, spectra judged to be of high quality, but that cannot be identified by database search, are prime candidates for still more computationally intensive methods, such as de novo sequencing or wider database searches including post-translational modifications. RESULTS: We report on two different approaches to assessing spectral quality prior to identification: binary classification, which predicts whether or not SEQUEST will be able to make an identification, and statistical regression, which predicts a more universal quality metric involving the number of b- and y-ion peaks. The best of our binary classifiers can eliminate over 75% of the unidentifiable spectra while losing only 10% of the identifiable spectra. Statistical regression can pick out spectra of modified peptides that can be identified by a de novo program but not by SEQUEST. In a section of independent interest, we discuss intensity normalization of mass spectra.

Algorithms↗

A hybrid genetic-neural system for predicting protein secondary structure.

BACKGROUND: Due to the strict relation between protein function and structure, the prediction of protein 3D-structure has become one of the most important tasks in bioinformatics and proteomics. In fact, notwithstanding the increase of experimental data on protein structures available in public databases, the gap between known sequences and known tertiary structures is constantly increasing. The need for automatic methods has brought the development of several prediction and modelling tools, but a general methodology able to solve the problem has not yet been devised, and most methodologies concentrate on the simplified task of predicting secondary structure. RESULTS: In this paper we concentrate on the problem of predicting secondary structures by adopting a technology based on multiple experts. The system performs an overall processing based on two main steps: first, a "sequence-to-structure" prediction is enforced by resorting to a population of hybrid (genetic-neural) experts, and then a "structure-to-structure" prediction is performed by resorting to an artificial neural network. Experiments, performed on sequences taken from well-known protein databases, allowed to reach an accuracy of about 76%, which is comparable to those obtained by state-of-the-art predictors. CONCLUSION: The adoption of a hybrid technique, which encompasses genetic and neural technologies, has demonstrated to be a promising approach in the task of protein secondary structure prediction.

Algorithms↗

Immunome research.

Immunology research has been transformed in the post-genomics era, with high throughput molecular biology and information technologies taking an increasingly central role. This has led to the development of a new area of science termed "Immunomics", that encompasses genomic, high throughput and bioinformatic approaches to immunology. In recognition of the increasing importance of this field, Immunome Research is a new Open Access, online journal, that will publish cutting edge research across the field of Immunomics. Immunome Research will publish a wide range of article types including specialty immunology databases, immunology database tools, immunome epitope research, epitope analysis tools, high-throughput technologies (gene sequencing, microarrays, proteomics), white papers, mathematical and theoretical models, and prediction tools. Immunome Research is the official journal of the International Immunomics Society (IIMMS).

Editorial↗

Computational prediction of human metabolic pathways from the complete human genome.

BACKGROUND: We present a computational pathway analysis of the human genome that assigns enzymes encoded therein to predicted metabolic pathways. Pathway assignments place genes in their larger biological context, and are a necessary first step toward quantitative modeling of metabolism. RESULTS: Our analysis assigns 2,709 human enzymes to 896 bioreactions; 622 of the enzymes are assigned roles in 135 predicted metabolic pathways. The predicted pathways closely match the known nutritional requirements of humans. This analysis identifies probable omissions in the human genome annotation in the form of 203 pathway holes (missing enzymes within the predicted pathways). We have identified putative genes to fill 25 of these holes. The predicted human metabolic map is described by a Pathway/Genome Database called HumanCyc, which is available at http://HumanCyc.org/. We describe the generation of HumanCyc, and present an analysis of the human metabolic map. For example, we compare the predicted human metabolic pathway complement to the pathways of Escherichia coli and Arabidopsis thaliana and identify 35 pathways that are shared among all three organisms. CONCLUSIONS: Our analysis elucidates a significant portion of the human metabolic map, and also indicates probable unidentified genes in the genome. HumanCyc provides a genome-based view of human nutrition that associates the essential dietary requirements of humans with a set of metabolic pathways whose existence is supported by the human genome. The database places many human genes in a pathway context, thereby facilitating analysis of gene expression, proteomics, and metabolomics datasets through a publicly available online tool called the Omics Viewer.

Arabidopsis↗

EYE on bioinformatics: dissecting complex disease traits in silico.

Bioinformatics has provided an unprecedented power and resource for us to decipher the enigma of complex diseases. It can reveal otherwise promiscuous information from the tremendous amount of data generated by the new, powerful and high-throughput technologies of genomics and proteomics. In this paper, we review the cutting edge developments in complex disease trait mapping, databases, computational gene recognition, gene function prediction, pathway reconstruction and disease classification by expression profiling, and computational modelling of living systems. Integration of all this knowledge and the different technologies, alongside cooperation between experts from different fields, will enhance our understanding of the molecular and mechanistic abnormalities in disease state, and greatly assist the rational development of effective therapies.

Chromosome Mapping↗

Added value for tandem mass spectrometry shotgun proteomics data validation through isoelectric focusing of peptides.

A very popular approach in proteomics is the so-called "shotgun LC-MS/MS" strategy. In its mostly used form, a total protein digest is separated by ion exchange fractionation in the first dimension followed by off- or on-line RP LC-MS/MS. We replaced the first dimension by isoelectric focusing in the liquid phase using the Off-Gel device producing 15 fractions. As peptides are separated by their isoelectric point in the first dimension and hydrophobicity in the second, those experimentally derived parameters (pI and R(T)) can be used for the validation of potentially identified peptides. We applied this strategy to a cellular extract of Drosophila Kc167 cells and identified peptides with two different database search engines, namely PHENYX and SEQUEST, with PeptideProphet validation of the SEQUEST results. PHENYX returned 7582 potential peptide identifications and SEQUEST 7629. The SEQUEST results were reduced to 2006 identifications by validation with PeptideProphet. Validation of the PeptideProphet, SEQUEST and PHENYX results by pI and R(T) parameters confirmed 1837 PeptideProphet identifications while in the remainder of the SEQUEST results another 1130 peptides were found to be likely hits. The validation on PHENYX resulted in the fixation of a solid p-value threshold of <1 x 10(-04) that sets by itself the correct identification confidence to >95%, and a final count of 2034 highly confident peptide identifications was achieved after pI and R(T) validation. Although the PeptideProphet and PHENYX datasets have a very high confidence the overlap of common identifications was only at 79.4%, to be explained by the fact that data interpretation was done searching different protein databases with two search engines of different algorithms. The approach used in this study allowed for an automated and improved data validation process for shotgun proteomics projects producing MS/MS peptide identification results of very high confidence.

Algorithms↗

A proteomic-based approach for the identification of Candida albicans protein components present in a subunit vaccine that protects against disseminated candidiasis.

Candidiasis has become a prevalent infection in different types of immunocompromised patients. The cell wall of Candida albicans plays important functions during the host-fungus interactions. Cell wall (surface) proteins of C. albicans are major elicitors of host immune responses during candidiasis, and represent candidates for vaccine development. Groups of mice were vaccinated subcutaneously with a beta-mercaptoethanol (beta-ME) extract from C. albicans containing cell wall proteins. Vaccinated mice were then infected with a lethal dose of C. albicans. Increased survival and decreased fungal burden were observed in vaccinated mice as compared to a control group, and 75% of vaccinated mice with the beta-ME extract survived this otherwise lethal infection. We used a proteomic approach (2-DE followed by immunoblotting) to demonstrate a complex polypeptidic pattern associated with the beta-ME extract used in the vaccine formulation and to detect immunogenic components recognized by antibodies in immune sera from vaccinated animals. Reactive protein spots were identified by MALDI-TOF-MS and searches in genomic databases. As a conclusion, vaccination strategies using C. albicans cell wall proteins induce protective responses. These antigens can be identified by proteomic approaches and may be used as components of subcellular vaccines against candidiasis.

Animals↗

PIMWalker: visualising protein interaction networks using the HUPO PSI molecular interaction format.

UNLABELLED: This article reports on PIMWalker, a free and interactive tool for visualising protein interaction networks. PIMWalker handles the unified molecular interaction (MI) format defined by members of the Proteomics Standards Initiative (the PSI MI format), and it is thus directly and easily usable by bench biologists. PIMWalker also comes with a documented, open-source Javatrade mark application programming interface allowing the bioinformatic programmer to easily extend the functions. AVAILABILITY: PIMWalker is available under a free license from http://pim.hybrigenics.com/pimwalker.

Computer Graphics↗

Performance of a genetic algorithm for mass spectrometry proteomics.

BACKGROUND: Recently, mass spectrometry data have been mined using a genetic algorithm to produce discriminatory models that distinguish healthy individuals from those with cancer. This algorithm is the basis for claims of 100% sensitivity and specificity in two related publicly available datasets. To date, no detailed attempts have been made to explore the properties of this genetic algorithm within proteomic applications. Here the algorithm's performance on these datasets is evaluated relative to other methods. RESULTS: In reproducing the method, some modifications of the algorithm as it is described are necessary to get good performance. After modification, a cross-validation approach to model selection is used. The overall classification accuracy is comparable though not superior to other approaches considered. Also, some aspects of the process rely upon random sampling and thus for a fixed dataset the algorithm can produce many different models. This raises questions about how to choose among competing models. How this choice is made is important for interpreting sensitivity and specificity results as merely choosing the model with lowest test set error rate leads to overestimates of model performance. CONCLUSIONS: The algorithm needs to be modified to reduce variability and care must be taken in how to choose among competing models. Results derived from this algorithm must be accompanied by a full description of model selection procedures to give confidence that the reported accuracy is not overstated.

Algorithms↗

SPINE workshop on automated X-ray analysis: a progress report.

The Structural Proteomics In Europe (SPINE) consortium contained a workpackage to address the automated X-ray analysis of macromolecules. The aim of this workpackage was to increase the throughput of three-dimensional structures while maintaining the high quality of conventional analyses. SPINE was able to bring together developers of software with users from the partner laboratories. Here, the results of a workshop organized by the consortium to evaluate software developed in the member laboratories against a set of bacterial targets are described. The major emphasis was on molecular-replacement suites, where automation was most advanced. Data processing and analysis, use of experimental phases and model construction were also addressed, albeit at a lower level.

Algorithms↗