Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Functional proteomic analysis of human nucleolus.

The notion of a "plurifunctional" nucleolus is now well established. However, molecular mechanisms underlying the biological processes occurring within this nuclear domain remain only partially understood. As a first step in elucidating these mechanisms we have carried out a proteomic analysis to draw up a list of proteins present within nucleoli of HeLa cells. This analysis allowed the identification of 213 different nucleolar proteins. This catalog complements that of the 271 proteins obtained recently by others, giving a total of approximately 350 different nucleolar proteins. Functional classification of these proteins allowed outlining several biological processes taking place within nucleoli. Bioinformatic analyses permitted the assignment of hypothetical functions for 43 proteins for which no functional information is available. Notably, a role in ribosome biogenesis was proposed for 31 proteins. More generally, this functional classification reinforces the plurifunctional nature of nucleoli and provides convincing evidence that nucleoli may play a central role in the control of gene expression. Finally, this analysis supports the recent demonstration of a coupling of transcription and translation in higher eukaryotes.

Cell Nucleolus↗

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational↗

High-throughput functional affinity purification of mannose binding proteins from Oryza sativa.

We have used affinity chromatography in combination with mass spectrometry to isolate, identify, and assign a preliminary functional annotation to a large number of both known and novel proteins from rice. Rice (Oryza sativa) leaf, root, and seed tissue extracts were fractionated by column affinity chromatography using alpha-D-mannose as the ligand. Bound fractions were eluted and subjected to one-dimensional electrophoresis, followed by high-performance liquid chromatography-tandem mass spectrometric analysis of separated proteins. This multiplexed technology resulted in the isolation and identification of 136 distinct mannose binding proteins from rice. A comparative analysis demonstrates very little overlap of identified proteins between the respective tissues, and confirms the correctly compartmentalized presence of a significant number of proteins from largely tissue-specific biochemical pathways. Over 30% of the identified proteins with a previously annotated function are directly involved in sugar metabolism, including several highly expressed known rice lectins. Direct comparison of the peptide sequences identified in this study to those peptides identified in the most comprehensive survey of the rice proteome to date indicates that our current data represents a significant enrichment of proteins unique to this dataset. Nearly 15% of the identified proteins, identified on the basis of exact peptide matching to sequences in the rice genomic database, represent proteins without a previously known functional annotation, indicating the potential of this combined chromatographic approach to assign a preliminary function to novel proteins in a high-throughput fashion.

Binding, Competitive↗

Synergistic computational and experimental proteomics approaches for more accurate detection of active serine hydrolases in yeast.

An analysis of the structurally and catalytically diverse serine hydrolase protein family in the Saccharomyces cerevisiae proteome was undertaken using two independent but complementary, large-scale approaches. The first approach is based on computational analysis of serine hydrolase active site structures; the second utilizes the chemical reactivity of the serine hydrolase active site in complex mixtures. These proteomics approaches share the ability to fractionate the complex proteome into functional subsets. Each method identified a significant number of sequences, but 15 proteins were identified by both methods. Eight of these were unannotated in the Saccharomyces Genome Database at the time of this study and are thus novel serine hydrolase identifications. Three of the previously uncharacterized proteins are members of a eukaryotic serine hydrolase family, designated as Fsh (family of serine hydrolase), identified here for the first time. OVCA2, a potential human tumor suppressor, and DYR-SCHPO, a dihydrofolate reductase from Schizosaccharomyces pombe, are members of this family. Comparing the combined results to results of other proteomic methods showed that only four of the 15 proteins were identified in a recent large-scale, "shotgun" proteomic analysis and eight were identified using a related, but similar, approach (neither identifies function). Only 10 of the 15 were annotated using alternate motif-based computational tools. The results demonstrate the precision derived from combining complementary, function-based approaches to extract biological information from complex proteomes. The chemical proteomics technology indicates that a functional protein is being expressed in the cell, while the computational proteomics technology adds details about the specific type of function and residue that is likely being labeled. The combination of synergistic methods facilitates analysis, enriches true positive results, and increases confidence in novel identifications. This work also highlights the risks inherent in annotation transfer and the use of scoring functions for determination of correct annotations.

Amino Acid Sequence↗

Between-gel reproducibility of the human cerebrospinal fluid proteome.

This manuscript describes the between-gel reproducibility of the two-dimensional gel electrophoresis analysis of the human lumbar cerebrospinal fluid (CSF) proteome. This reproducibility study is a necessary component for our long-term research program that uses comparative proteomics to analyze lumbar CSF samples in a study of human idiopathic low back pain. A Protein-Plus Dodeca Cell electrophoresis apparatus and PDQuest software were used to measure the level of between-gel reproducibility of the CSF proteome. One pooled CSF sample was used to evaluate the level of within-sample, between-gel reproducibility, and a set of seven different CSF samples (CSF-1 to 7) was used to test the level of within-group and between-group variability. Differentially expressed proteins (six CSF samples versus the designated control, CSF-3) were characterized with mass spectrometry. The number of spots found in the pooled CSF sample was 490 +/- 30 (n = 10 gels); the percentage of protein spots found in those 10 gels was 92 +/- 6%, with a coefficient of variation of 6%; and a positive coefficient of correlation (r = 0.82) was found. In order to test the proof-of-principle, that set of seven CSF samples served as a test of our ability to perform reproducibility comparative proteomics, and to detect differentially expressed proteins within that set of test samples. One sample (CSF-3) served as the control for the other six to locate the differentially expressed proteins. A comparison of fifteen differentially expressed proteins found in that set of test CSF samples correlated with pathology. Matrix-assisted laser desorption/ionization-time-of-flight and electrospray ionization quadrupole ion trap mass spectrometry were used to characterize thirteen of those fifteen differentially expressed proteins. These results (reproducibility, protein characterization, set of test samples, and proof-of-principle) suggest that the analysis of human CSF two-dimensional gels can achieve a high level of within-sample and between-sample reproducibility, and that PDQuest software can measure the relative protein abundance in the human CSF proteome.

Adult↗

Redox proteomics in some age-related neurodegenerative disorders or models thereof.

Neurodegenerative diseases cause memory loss and cognitive impairment. Results from basic and clinical scientific research suggest a complex network of mechanisms involved in the process of neurodegeneration. Progress in treatment of such disorders requires researchers to better understand the functions of proteins involved in neurodegenerative diseases, to characterize their role in pathogenic disease mechanisms, and to explore their roles in the diagnosis, treatment, and prevention of neurodegenerative diseases. A variety of conditions of neurodegenerative diseases often lead to post-translational modifications of proteins, including oxidation and nitration, which might be involved in the pathogenesis of neurodegenerative diseases. Redox proteomics, a subset of proteomics, has made possible the identification of specifically oxidized proteins in neurodegenerative disorders, providing insight into a multitude of pathways that govern behavior and cognition and the response of the nervous system to injury and disease. Proteomic analyses are particularly suitable to elucidate post-translational modifications, expression levels, and protein-protein interactions of thousands of proteins at a time. Complementing the valuable information generated through the integrative knowledge of protein expression and function should enable the development of more efficient diagnostic tools and therapeutic modalities. Here we review redox proteomic studies of some neurodegenerative diseases.

Aged↗

Recent developments in proteomics: implications for the study of cardiac hypertrophy and failure.

The key components to the molecular understanding of the pathophysiology of various forms of heart failure involve global and/or large-scale identifications of proteins, their patterns of expression, posttranslational modifications, and functional characterization. Particularly, proteins involved in the induction of cardiac (mal)adaptive hypertrophic growth, interstitial fibrosis, and contractile dysfunction are of interest. In general, with the accumulation of vast amounts of DNA sequences in databases, researchers have become aware that merely having complete sequences of genomes and transcriptional changes for thousands of genes simultaneously will not be sufficient to elucidate, in molecular terms, the etiology and pathophysiology of cardiovascular disease. In the last decade, a new technology called proteomics has become available that allows biological and (patho)physiological questions to be approached exclusively from the protein perspective. Proteomics may enable us to map the entire complement of proteins expressed by the heart at any time and condition. This approach creates the unique possibility to identify, by differential analysis, protein alterations associated with the etiology of heart disease and its progression, outcome, and response to therapy. To illustrate the true power of proteomics, most of the currently available methodologies are first reviewed, including their limitations. This review also deals with the current status and the perspectives of proteomics applications in research on heart failure in general. Furthermore, examples of our recent data on global protein profiling of the pressure-overloaded rat right ventricle and of endothelin-1-stimulated cultures of neonatal rat cardiac myocytes are provided. The last section is devoted to the continuous advances in proteomic technologies, including protein separation methods, mass spectrometric instrumentation, computational analysis, and bioinformatic tools, together with integrative databases.

Cardiomegaly↗

MitoP2, an integrated database on mitochondrial proteins in yeast and man.

The aim of the MitoP2 database (http://ihg.gsf.de/mitop2) is to provide a comprehensive list of mitochondrial proteins of yeast and man. Based on the current literature we created an annotated reference set of yeast and human proteins. In addition, data sets relevant to the study of the mitochondrial proteome are integrated and accessible via search tools and links. They include computational predictions of signalling sequences, and summarize results from proteome mapping, mutant screening, expression profiling, protein-protein interaction and cellular sublocalization studies. For each individual approach, specificity and sensitivity for allocating mitochondrial proteins was calculated. By providing the evidence for mitochondrial candidate proteins the MitoP2 database lends itself to the genetic characterization of human mitochondriopathies.

Computational Biology↗

Subproteomes of soluble and structure-bound Helicobacter pylori proteins analyzed by two-dimensional gel electrophoresis and mass spectrometry.

Helicobacter pylori is one of the most common bacterial pathogens and causes a variety of diseases, such as peptic ulcer or gastric cancer. Despite intensive study of this human pathogen in the last decades, knowledge about its membrane proteins and, in particular, those which are putative components of the type IV secretion system encoded by the cag pathogenicity island (PAI) remains limited. Our aim is to establish a dynamic two-dimensional electrophoresis-polyacrylamide gel electrophoresis (2-DE-PAGE) database with multiple subproteomes of H. pylori (http://www.mpiib-berlin.mpg.de/2D-PAGE) which facilitates identification of bacterial proteins important in pathogen-host interactions. Using a proteomic approach, we investigated the protein composition of two H. pylori fractions: soluble proteins and structure-bound proteins (including membrane proteins). Both fractions differed markedly in the overall protein composition as determined by 2-DE. The 50 most abundant protein spots in each fraction were identified by peptide mass fingerprinting. We detected four cag PAI proteins, numerous outer membrane proteins (OMPs), the vacuolating cytotoxin VacA, other potential virulence factors, and few ribosomal proteins in the structure-bound fraction. In contrast, catalase (KatA), gamma-glutamyltranspeptidase (Ggt), and the neutrophil-activating protein NapA were found almost exclusively in the soluble protein fraction. The results presented here are an important complement to genome sequence data, and the established 2-D PAGE maps provide a basis for comparative studies of the H. pylori proteome. Such subproteomes in the public domain will be effective instruments for identifying new virulence factors and antigens of potential diagnostic and/or curative value against infections with this important pathogen.

Amino Acid Sequence↗

On the proper use of mass accuracy in proteomics.

Mass measurement is the main outcome of mass spectrometry-based proteomics yet the potential of recent advances in accurate mass measurements remains largely unexploited. There is not even a clear definition of mass accuracy in the proteomics literature, and we identify at least three uses of this term: anecdotal mass accuracy, statistical mass accuracy, and the maximum mass deviation (MMD) allowed in a database search. We suggest using the second of these terms as the generic one. To make the best use of the mass precision offered by modern instruments we propose a series of simple steps involving recalibration of the data on "internal standards" contained in every proteomics data set. Each data set should be accompanied by a plot of mass errors from which the appropriate MMD can be chosen. More advanced uses of high mass accuracy include an MMD that depends on the signal abundance of each peptide. Adapting search engines to high mass accuracy in the MS/MS data is also a high priority. Proper use of high mass accuracy data can make MS-based proteomics one of the most "digital" and accurate post-genomics disciplines.

Mass Spectrometry↗

Detecting protein sequence conservation via metric embeddings.

MOTIVATION: Comparing two protein databases is a fundamental task in biosequence annotation. Given two databases, one must find all pairs of proteins that align with high score under a biologically meaningful substitution score matrix, such as a BLOSUM matrix (Henikoff and Henikoff, 1992). Distance-based approaches to this problem map each peptide in the database to a point in a metric space, such that peptides aligning with higher scores are mapped to closer points. Many techniques exist to discover close pairs of points in a metric space efficiently, but the challenge in applying this work to proteomic comparison is to find a distance mapping that accurately encodes all the distinctions among residue pairs made by a proteomic score matrix. Buhler (2002) proposed one such mapping but found that it led to a relatively inefficient algorithm for protein-protein comparison. RESULTS: This work proposes a new distance mapping for peptides under the BLOSUM matrices that permits more efficient similarity search. We first propose a new distance function on peptides derived from a given score matrix. We then show how to map peptides to bit vectors such that the distance between any two peptides is closely approximated by the Hamming distance (i.e. number of mismatches) between their corresponding bit vectors. We combine these two results with the LSH-ALL-PAIRS-SIM algorithm of Buhler (2002) to produce an improved distance-based algorithm for proteomic comparison. An initial implementation of the improved algorithm exhibits sensitivity within 5% of that of the original LSH-ALL-PAIRS-SIM, while running up to eight times faster.

Algorithms↗

Modeling a whole organ using proteomics: the avian bursa of Fabricius.

While advances in proteomics have improved proteome coverage and enhanced biological modeling, modeling function in multicellular organisms requires understanding how cells interact. Here we used the chicken bursa of Fabricius, a common experimental system for B cell function, to model organ function from proteomics data. The bursa has two major functional cell types: B cells and the supporting stromal cells. We used differential detergent fractionation-multidimensional protein identification technology (DDF-MudPIT) to identify 5198 proteins from all cellular compartments. Of these, 1753 were B cell specific, 1972 were stroma specific and 1473 were shared between the two. By modeling programmed cell death (PCD), cell differentiation and proliferation, and transcriptional activation, we have improved functional annotation of chicken proteins and placed chicken-specific death receptors into the PCD process using phylogenetics. We have identified 114 transcription factors (TFs); 42 of the bursal B cell TFs have not been reported before in any B cells. We have also improved the structural annotation of a newly sequenced genome by confirming the in vivo expression of 4006 "predicted", and 6623 ab initio, ORFs. Finally, we have developed a novel method for facilitating structural annotation, "expressed peptide sequence tags" (ePSTs) and demonstrate its utility by identifying 521 potential novel proteins from the chicken "unassigned chromosome".

Amino Acid Sequence↗

Isobaric tags for relative and absolute quantitation (iTRAQ) reproducibility: Implication of multiple injections.

We analyzed 10 isobaric tags for relative and absolute quantitation (iTRAQ) experiments using three different model organisms across the domains of life: Saccharomyces cerevisiae KAY446, Sulfolobussolfataricus P2, and Synechocystis sp. PCC6803. A double database search strategy was employed to minimize the rate of false positives to less than 3% for all organisms. The reliability of proteins with single-peptide identification was also assessed using the search strategy, coupled with multiple analyses of samples into LC-MS/MS. The outcomes of the three LC-MS/MS analyses provided higher proteome coverage with an average increment in total proteins identified of 6%, 33%, and 50% found in S. cerevisiae, S. solfataricus, and Synechocystis sp., respectively. The iTRAQ quantification values were found to be highly reproducible across the injections, with an average coefficient of variation (CV) of 0.09 (scattering from 0.14 to 0.04) calculated based on log mean average ratio for all three organisms. Hence, we recommend multiple analyses of iTRAQ samples for greater proteome coverage and precise quantification.

Amino Acid Sequence↗

Beyond the genome--SMi Conference 27-28 January 2003, London, UK.

Over a decade of astonishing developments, genomics and proteomics have promised a fundamentally new approach to drug discovery. Although there has been an undeniable increase in the range of potential targets available, this has not led to an increased output of the drug discovery pipeline into the clinic. With tighter markets and increasing competition, the major pharmaceutical companies are under intense pressure to achieve rapid, concrete delivery of those early promises, but there remain acute problems in the genes-to-drugs pipeline. This meeting showcased a range of novel approaches from proteomics and bioinformatics to address these problems. A common theme in the range of proteomics offerings was the prioritization of potential novel targets on the basis of their accessibility to drugs and their functional link to disease phenotypes. Informatics and in silico offerings also concentrated on fast, accurate, drug-focused workflows built on large integrative databases and novel data-mined algorithms.

Animals↗

ProMate: a structure based prediction program to identify the location of protein-protein binding sites.

Is the whole protein surface available for interaction with other proteins, or are specific sites pre-assigned according to their biophysical and structural character? And if so, is it possible to predict the location of the binding site from the surface properties? These questions are answered quantitatively by probing the surfaces of proteins using spheres of radius of 10 A on a database (DB) of 57 unique, non-homologous proteins involved in heteromeric, transient protein-protein interactions for which the structures of both the unbound and bound states were determined. In structural terms, we found the binding site to have a preference for beta-sheets and for relatively long non-structured chains, but not for alpha-helices. Chemically, aromatic side-chains show a clear preference for binding sites. While the hydrophobic and polar content of the interface is similar to the rest of the surface, hydrophobic and polar residues tend to cluster in interfaces. In the crystal, the binding site has more bound water molecules surrounding it, and a lower B-factor already in the unbound protein. The same biophysical properties were found to hold for the unbound and bound DBs. All the significant interface properties were combined into ProMate, an interface prediction program. This was followed by an optimization step to choose the best combination of properties, as many of them are correlated. During optimization and prediction, the tested proteins were not used for data collection, to avoid over-fitting. The prediction algorithm is fully automated, and is used to predict the location of potential binding sites on unbound proteins with known structures. The algorithm is able to successfully predict the location of the interface for about 70% of the proteins. The success rate of the predictor was equal whether applied on the unbound DB or on the disjoint bound DB. A prediction is assumed correct if over half of the predicted continuous interface patch is indeed interface. The ability to predict the location of protein-protein interfaces has far reaching implications both towards our understanding of specificity and kinetics of binding, as well as in assisting in the analysis of the proteome.

Algorithms↗

Evaluation of multidimensional chromatography coupled with tandem mass spectrometry (LC/LC-MS/MS) for large-scale protein analysis: the yeast proteome.

Highly complex protein mixtures can be directly analyzed after proteolysis by liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS). In this paper, we have utilized the combination of strong cation exchange (SCX) and reversed-phase (RP) chromatography to achieve two-dimensional separation prior to MS/MS. One milligram of whole yeast protein was proteolyzed and separated by SCX chromatography (2.1 mm i.d.) with fraction collection every minute during an 80-min elution. Eighty fractions were reduced in volume and then re-injected via an autosampler in an automated fashion using a vented-column (100 microm i.d.) approach for RP-LC-MS/MS analysis. More than 162,000 MS/MS spectra were collected with 26,815 matched to yeast peptides (7,537 unique peptides). A total of 1,504 yeast proteins were unambiguously identified in this single analysis. We present a comparison of this experiment with a previously published yeast proteome analysis by Yates and colleagues (Washburn, M. P.; Wolters, D.; Yates, J. R., III. Nat. Biotechnol. 2001, 19, 242-7). In addition, we report an in-depth analysis of the false-positive rates associated with peptide identification using the Sequest algorithm and a reversed yeast protein database. New criteria are proposed to decrease false-positives to less than 1% and to greatly reduce the need for manual interpretation while permitting more proteins to be identified.

Cations↗

Structural domains, protein modules, and sequence similarities enrich our understanding of the Shewanella oneidensis MR-1 proteome.

The protein coding sequences of S. oneidensis MR-1 were analyzed, and new annotations were given to 491 gene products, 306 of which were previously of unknown function. New information was mainly brought in from structural domain predictions for S. oneidensis proteins of the SUPERFAM database (http://supfam.mrc-lmb.cam.ac.uk/SUPERFAMILY/) and newly identified and experimentally verified functions of homologous proteins. Proteins encoded by fused genes were identified and separated into modules, protein units of at least 83 aa with independent functions and distinct evolutionary histories. A reannotation of the fused gene products was done to assign functions to the appropriate module within the protein. Groups of sequence-similar proteins of S. oneidensis were assembled. The fused gene products were represented by their modular entities for the grouping process. The protein groups were analyzed for their size and functions, and they were used to indicate activities that are of importance to the environmental adaptation of this organism. Making use of several approaches not commonly used in annotation, we have been able to enrich our understanding of the functions encoded by the S. oneidensis genome.

Bacterial Proteins↗

De novo sequencing of neuropeptides using reductive isotopic methylation and investigation of ESI QTOF MS/MS fragmentation pattern of neuropeptides with N-terminal dimethylation.

A stable-isotope dimethyl labeling strategy was previously shown to be a useful tool for quantitative proteomics. More recently, N-terminal dimethyl labeling was also reported for peptide sequencing in combination with database searching. Here, we extend these previous studies by incorporating N-terminal isotopic dimethylation for de novo sequencing of neuropeptides directly from tissue extract without any genomic information. We demonstrated several new sequencing applications of this method in addition to the identification of the N-terminal residue using the enhanced a(1) ion. The isotopic labeling also provides easier and more confident de novo sequencing of peptides by comparing similar MS/MS fragmentation patterns of the isotopically labeled peptide pairs. The current study on neuropeptides shows several distinct fragmentation patterns after N-terminal dimethylation which have not been reported previously. The y((n-1)) ion is enhanced in multiply charged peptides and is weak or missing in singly charged peptides. The MS/MS spectra of singly charged peptides are simplified due to the enhanced N-terminal fragments and suppressed internal fragments. The neutral loss of dimethylamine is also observed. The mechanisms for the above fragmentations are proposed. Finally, the structures of the immonium ion and related ions of N(alpha), N(epsilon)-tetramethylated lysine and N(epsilon)-dimethylated lysine are explored.

Amino Acid Sequence↗