Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Evidence of glutathione transferase complexing and signaling in the model nematode Caenorhabditis elegans using a pull-down proteomic assay.

Phage display techniques using random peptide interactions have supported the role of mammalian glutathione transferase (GST) as part of a signalling pathway for both oxidative stress and an apoptosis pathway. Little is known about the interaction of nonmammalian GST with other proteins. GSTs have been implicated in the development of chronic nematode infections by neutralising cytotoxic products arising from host immune initiated reactive oxygen species (ROS) assault. In this study we attached one of the key GSTs expressed in the model nematode Caenorhabditis elegans to an affinity support matrix and directly identified major interacting proteins by two-dimensional electrophoresis and peptide mass fingerprinting before and following oxidative stress. Nematode GST does not appear to be a stand-alone enzyme and interacts with many types of proteins in both normal and ROS stress conditions. Pull-down proteomic presents a flexible, label free, rapid and economical assay without specialised ligand fishing equipment to identify protein binding partners.

Amino Acid Sequence↗

High throughput quantitative glycomics and glycoform-focused proteomics of murine dermis and epidermis.

Despite recent advances in our understanding of the significance of the protein glycosylation, the throughput of protein glycosylation analysis is still too low to be applied to the exhaustive glycoproteomic analysis. Aiming to elucidate the N-glycosylation of murine epidermis and dermis glycoproteins, here we used a novel approach for focused proteomics. A gross N-glycan profiling (glycomics) of epidermis and dermis was first elucidated both qualitatively and quantitatively upon N-glycan derivatization with novel, stable isotope-coded derivatization reagents followed by MALDI-TOF(/TOF) analysis. This analysis revealed distinct features of the N-glycosylation profile of epidermis and dermis for the first time. A high abundance of high mannose type oligosaccharides was found to be characteristic of murine epidermis glycoproteins. Based on this observation, we performed high mannose type glycoform-focused proteomics by direct tryptic digestion of protein mixtures and affinity enrichment. We identified 15 glycoproteins with 19 N-glycosylation sites that carry high mannose type glycans by off-line LC-MALDI-TOF/TOF mass spectrometry. Moreover the relative quantity of microheterogeneity of different glycoforms present at each N-glycan binding site was determined. Glycoproteins identified were often contained in lysosomes (e.g. cathepsin L and gamma-glutamyl hydrolase), lamellar granules (e.g. glucosylceramidase and cathepsin D), and desmosomes (e.g. desmocollin 1, desmocollin 3, and desmoglein). Lamellar granules are organelles found in the terminally differentiating cells of keratinizing epithelia, and desmosomes are intercellular junctions in vertebrate epithelial cells, thus indicating that N-glycosylation of tissue-specific glycoproteins may contribute to increase the relative proportion of high mannose glycans. The striking roles of lysosomal enzymes in epidermis during lipid remodeling and desquamation may also reflect the observed high abundance of high mannose glycans.

Amino Acid Sequence↗

A statistical framework for combining and interpreting proteomic datasets.

MOTIVATION: To identify accurately protein function on a proteome-wide scale requires integrating data within and between high-throughput experiments. High-throughput proteomic datasets often have high rates of errors and thus yield incomplete and contradictory information. In this study, we develop a simple statistical framework using Bayes' law to interpret such data and combine information from different high-throughput experiments. In order to illustrate our approach we apply it to two protein complex purification datasets. RESULTS: Our approach shows how to use high-throughput data to calculate accurately the probability that two proteins are part of the same complex. Importantly, our approach does not need a reference set of verified protein interactions to determine false positive and false negative error rates of protein association. We also demonstrate how to combine information from two separate protein purification datasets into a combined dataset that has greater coverage and accuracy than either dataset alone. In addition, we also provide a technique for estimating the total number of proteins which can be detected using a particular experimental technique. AVAILABILITY: A suite of simple programs to accomplish some of the above tasks is available at www.unm.edu/~compbio/software/DatasetAssess

Algorithms↗

Mycobacterium tuberculosis functional network analysis by global subcellular protein profiling.

Trends in increased tuberculosis infection and a fatality rate of approximately 23% have necessitated the search for alternative biomarkers using newly developed postgenomic approaches. Here we provide a systematic analysis of Mycobacterium tuberculosis (Mtb) by directly profiling its gene products. This analysis combines high-throughput proteomics and computational approaches to elucidate the globally expressed complements of the three subcellular compartments (the cell wall, membrane, and cytosol) of Mtb. We report the identifications of 1044 proteins and their corresponding localizations in these compartments. Genome-based computational and metabolic pathways analyses were performed and integrated with proteomics data to reconstruct response networks. From the reconstructed response networks for fatty acid degradation and lipid biosynthesis pathways in Mtb, we identified proteins whose involvements in these pathways were not previously suspected. Furthermore, the subcellular localizations of these expressed proteins provide interesting insights into the compartmentalization of these pathways, which appear to traverse from cell wall to cytoplasm. Results of this large-scale subcellular proteome profile of Mtb have confirmed and validated the computational network hypothesis that functionally related proteins work together in larger organizational structures.

Automation↗

VirGen: a comprehensive viral genome resource.

VirGen is a comprehensive viral genome resource that organizes the 'sequence space' of viral genomes in a structured fashion. It has been developed with the objective of serving as an annotated and curated database comprising complete genome sequences of viruses, value-added derived data and data mining tools. The current release (v1.1) contains 559 complete genomes in addition to 287 putative genomes of viruses belonging to eight viral families for which the host range includes animals and plants. Viral genomes in VirGen are annotated using sequence-based Bioinformatics approaches. The genomic data is also curated to identify 'alternate names' of viral proteins, where available. VirGen archives the results of comparisons of genomes, proteomes and individual proteins within and between viral species. It is the first resource to provide phylogenetic trees of viral species computed using whole-genome sequence data. The module of predicted B-cell antigenic determinants in VirGen is an attempt to link the genome to its vaccinome. Comparative genome analysis data facilitate the study of genome organization and evolution of viruses, which would have implications in applied research to identify candidates for the design of vaccines and antiviral drugs. VirGen is a relational database and is available at http://bioinfo. ernet.in/virgen/virgen.html.

Antigens, Viral↗

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals↗

Virtual Expert Mass Spectrometrist v3.0: an integrated tool for proteome analysis.

The number of tools described in the literature for analysis of proteome data is growing fast. However, most tools are not able to communicate or exchange data with other tools. In Virtual Expert Mass Spectrometrist (VEMS) v3.0 an effort has been made to interface and export to already existing tools. In this chapter, an outline of how to use the VEMS program to search tandem mass spectrometry data against databases is described. Additionally, examples on how to extend the analysis with other external tools are given.

Calibration↗

Precise peptide sequencing and protein quantification in the human proteome through in vivo lysine-specific mass tagging.

Proteomics studies demand new scalable and automatable MS-based methods with higher specificity and accuracy. Here we describe an accurate and efficient method for both precise quantification and comprehensive de novo identification of peptide sequences in complex mixtures. The unique feature of this method is based on the incorporation of deuterium-labeled (heavy) lysines into proteins through in vivo cell culturing, which introduces specific mass tags at the carboxyl termini of proteolytic peptides when cleaved by certain proteases. The mass shift between the unlabeled and the deuterated lysine (lys-d4) assigns a mass signature to all lysine-containing peptides in any pool of proteolytic peptides. Lys-d4 tags can also serve as internal markers in MS/MS fragment spectra when they are buried in some peptide sequences due to miscleavages. This signal specificity circumvents the mass accuracy limitations in determining particular amino acid residues for de novo sequencing. Further, this strategy of lysine-specific tagging was successfully implemented to measure the differential protein expression of human skin fibroblast cells in response to heat shock.

Amino Acid Sequence↗

From XML to RDF: how semantic web technologies will change the design of 'omic' standards.

With the ongoing rapid increase in both volume and diversity of 'omic' data (genomics, transcriptomics, proteomics, and others), the development and adoption of data standards is of paramount importance to realize the promise of systems biology. A recent trend in data standard development has been to use extensible markup language (XML) as the preferred mechanism to define data representations. But as illustrated here with a few examples from proteomics data, the syntactic and document-centric XML cannot achieve the level of interoperability required by the highly dynamic and integrated bioinformatics applications. In the present article, we discuss why semantic web technologies, as recommended by the World Wide Web consortium (W3C), expand current data standard technology for biological data representation and management.

Algorithms↗

The German cDNA network: cDNAs, functional genomics and proteomics.

Among the greatest challenges facing biology today is the exploitation of huge amounts of genomic data, and their conversion into functional information about the proteins encoded. For example, the large-scale cDNA sequencing project of the German cDNA Consortium is providing vast numbers of open reading frames (ORFs) encoding novel proteins of completely unknown function. As a first step towards their characterization we have tagged over 500 of these with the green fluorescent protein (GFP), and examined the subcellular localizations of these fusion proteins in living cells. These data have allowed us to classify the proteins into subcellular groups which determines the next step towards a detailed functional characterization. To make further use of these GFP-tagged constructs, a series of functional assays have been designed and implemented to assess the effect of these novel proteins on processes such as cell growth, cell death, and protein transport. Functional assays with such a large set of molecules is only possible by automation. Therefore, we have developed, and adapted, functional assays for use by robotic liquid handling stations and reading stations. A transport assay allows to identify proteins which localize to distinct organelles of the secretory pathway and have the potential to be new regulators in protein transport, a proliferation assay helps identifying proteins that stimulate or repress mitosis. Further assays to monitor the effects of the proteins in apoptosis and signal transduction pathways are in progress. Integrating the functional information that is generated in the assays with data from expression profiling and further functional genomics and proteomics approaches, will ultimately allow us to identify functional networks of proteins in a morphological context, and will greatly contribute to our understanding of cell function.

Cloning, Molecular↗

Progress in the definition of a reference human mitochondrial proteome.

Owing to the complexity of higher eukaryotic cells, a complete proteome is likely to be very difficult to achieve. However, advantage can be taken of the cell compartmentalization to build organelle proteomes, which can moreover be viewed as specialized tools to study specifically the biology and "physiology" of the target organelle. Within this frame, we report here the construction of the human mitochondrial proteome, using placenta as the source tissue. Protein identification was carried out mainly by peptide mass fingerprinting. The optimization steps in two-dimensional electrophoresis needed for proteome research are discussed. However, the relative paucity of data concerning mitochondrial proteins is still the major limiting factor in building the corresponding proteome, which should be a useful tool for researchers working on human mitochondria and their deficiencies.

Databases as Topic↗

Proteome analysis of mesencephalic tissues: evidence for Parkinson's disease.

Proteome analysis is a powerful methodology to investigate protein expression in tissues involved in diseases not linked to particular genetic defects. To date, this technique has a limited number of applications in the field of neurodegenerative disorders. We decided therefore to investigate by this approach autoptic mesencephalic tissues of patients with idiopathic Parkinson's disease as well as control specimens from healthy subjects.

Blotting, Western↗

Pox proteomics: mass spectrometry analysis and identification of Vaccinia virion proteins.

BACKGROUND: Although many vaccinia virus proteins have been identified and studied in detail, only a few studies have attempted a comprehensive survey of the protein composition of the vaccinia virion. These projects have identified the major proteins of the vaccinia virion, but little has been accomplished to identify the unknown or less abundant proteins. Obtaining a detailed knowledge of the viral proteome of vaccinia virus will be important for advancing our understanding of orthopoxvirus biology, and should facilitate the development of effective antiviral drugs and formulation of vaccines. RESULTS: In order to accomplish this task, purified vaccinia virions were fractionated into a soluble protein enriched fraction (membrane proteins and lateral bodies) and an insoluble protein enriched fraction (virion cores). Each of these fractions was subjected to further fractionation by either sodium dodecyl sulfate-polyacrylamide gel electophoresis, or by reverse phase high performance liquid chromatography. The soluble and insoluble fractions were also analyzed directly with no further separation. The samples were prepared for mass spectrometry analysis by digestion with trypsin. Tryptic digests were analyzed by using either a matrix assisted laser desorption ionization time of flight tandem mass spectrometer, a quadrupole ion trap mass spectrometer, or a quadrupole-time of flight mass spectrometer (the latter two instruments were equipped with electrospray ionization sources). Proteins were identified by searching uninterpreted tandem mass spectra against a vaccinia virus protein database created by our lab and a non-redundant protein database. CONCLUSION: Sixty three vaccinia proteins were identified in the virion particle. The total number of peptides found for each protein ranged from 1 to 62, and the sequence coverage of the proteins ranged from 8.2% to 94.9%. Interestingly, two vaccinia open reading frames were confirmed as being expressed as novel proteins: E6R and L3L.

Amino Acid Sequence↗

Mass spectrometric identification and microcharacterization of proteins from electrophoretic gels: strategies and applications.

The entire genomic DNA sequences of a number of prokaryotic and eukaryotic species are now available and many more, including the human genome, will be completed in the near future. The state-of-life of a cell at any given time, however, is defined by its protein composition, i.e., its proteome. Gel electrophoresis, mass spectrometry, and bioinformatics will be important tools for protein and proteome analysis in the post-genome era. Protein identification from electrophoretic gels by mass spectrometric peptide mapping or peptide sequencing combined with sequence database searching is established and has been applied to numerous biological systems. We describe current strategies and selected applications in molecular and cell biology. The next challenges are detailed structure/function analyses, which include studying the molecular composition of multiprotein complexes and characterization of secondary modifications of proteins. The advantages and limitations of a number of mass spectrometry-based strategies designed for microcharacterization of low amounts of protein from electrophoretic gels are discussed and illustrated by examples.

Amino Acid Sequence↗

Isolation and proteomic characterization of the major proteins of the nucleolin-binding ribonucleoprotein complexes.

Nucleolin (NCL) is one of the most abundant nucleolar proteins of exponentially growing eukaryotic cells. It is known to interact only transiently with rRNA and preribosomal particles and not to be detectable in mature cytoplasmic ribosomes, and is believed to function as multi-protein complexes during ribosome biogenesis and maturation. However, those multiprotein complexes remain only partially characterized due to the difficulty of conventional protein analysis methods. Here we report isolation of NCL-binding protein complex and its proteomic characterization with the use of an analytical method based on matrix-assisted laser desorption/ionization-time of flight analysis coupled with searching peptide mass databases. The NCL-binding protein complex was isolated by immunoprecipitation with anti-Flag antibody from human kidney 293 cells that were transfected with the Flag-tagged NCL gene, and showed RNA integrity for holding their protein constituents. Interaction between NCL and its binding complex was disrupted by an RNA oligonucleotide with a NCL recognition element, indicating that NCL binds to the ribonucleoprotein (RNP) complex mainly through the sequence specific protein-RNA interaction. We confirmed that an RNA-binding domain of NCL alone was sufficient to hold the entire NCL-binding RNP complex, indicating the strict binding specificity of NCL to the isolated RNP complex in 293 cells. We identified forty ribosomal proteins from both the large and small subunits, and twenty nonribosomal proteins. These results together suggest that the isolated NCL-binding RNP complex is a preribosomal particle present in the nucleolus of 293 cells.

Blotting, Western↗

Nonredundant mass spectrometry: a strategy to integrate mass spectrometry acquisition and analysis.

Protein identification using automated data-dependent tandem mass spectrometry (MS/MS) is now a standard procedure. However, in many cases data-dependent acquisition becomes redundant acquisition as many different peptides from the same protein are fragmented, whilst only a few are needed for unambiguous identification. To increase the quality of information but decrease the amount of information, a nonredundant MS (nrMS) strategy has been developed. With nrMS, data analysis is an integral part of the overall MS acquisition and analysis, and not an endpoint as typically performed. In this nrMS workflow a matrix assisted laser desorption/ionization-time of flight-time of flight (MALDI-TOF/TOF) instrument is used. MS and restricted MS/MS data are searched and identified proteins are used to generate an "exclusion list", after in silico digestion. Peptide fragmentation is then restricted to only the most intense ions not present in the exclusion list. This process is repeated until all peaks are accounted for or the sample is consumed. Compared to nanoLC-MS/MS, nrMS yielded similar results for the analysis of six pooled two-dimensional electrophoresis (2-DE) spots. In comparison to standard data-dependent MALDI-MS/MS for sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) gel band analysis, nrMS dramatically increased the number of identified proteins. It was also found that this new workflow significantly increased sequence coverage by identifying unexpected peptides, which can result from post-translational modifications.

Algorithms↗

Automatic annotation of matrix-assisted laser desorption/ionization N-glycan spectra.

Matrix-assisted laser desorption/ionization-mass spectrometry (MALDI-MS) is the pre-eminent technique for mass mapping of glycans. In order to make this technique practical for high-throughput screening, reliable automatic methods of annotating peaks must be devised. We describe an algorithm called Cartoonist that labels peaks in MALDI spectra of permethylated N-glycans with cartoons which represent the most plausible glycans consistent with the peak masses and the types of glycans being analyzed. There are three main parts to Cartoonist. (i) It selects annotations from a library of biosynthetically plausible cartoons. The library we currently use has about 2800 cartoons, but was constructed using only about 300 archetype cartoons entered by hand. (ii) It determines the precision and calibration of the machine used to generate the spectrum. It does this automatically based on the spectrum itself. (iii) It assigns a confidence score to each annotation. In particular, rather than making a binary yes/no decision when annotating a peak, it makes all plausible annotations and associates them with scores indicating the probability that they are correct.

Algorithms↗

Effect of training datasets on support vector machine prediction of protein-protein interactions.

Knowledge of protein-protein interaction is useful for elucidating protein function via the concept of 'guilt-by-association'. A statistical learning method, Support Vector Machine (SVM), has recently been explored for the prediction of protein-protein interactions using artificial shuffled sequences as hypothetical noninteracting proteins and it has shown promising results (Bock, J. R., Gough, D. A., Bioinformatics 2001, 17, 455-460). It remains unclear however, how the prediction accuracy is affected if real protein sequences are used to represent noninteracting proteins. In this work, this effect is assessed by comparison of the results derived from the use of real protein sequences with that derived from the use of shuffled sequences. The real protein sequences of hypothetical noninteracting proteins are generated from an exclusion analysis in combination with subcellular localization information of interacting proteins found in the Database of Interacting Proteins. Prediction accuracy using real protein sequences is 76.9% compared to 94.1% using artificial shuffled sequences. The discrepancy likely arises from the expected higher level of difficulty for separating two sets of real protein sequences than that for separating a set of real protein sequences from a set of artificial sequences. The use of real protein sequences for training a SVM classification system is expected to give better prediction results in practical cases. This is tested by using both SVM systems for predicting putative protein partners of a set of thioredoxin related proteins. The prediction results are consistent with observations, suggesting that real sequence is more practically useful in development of SVM classification system for facilitating protein-protein interaction prediction.

Algorithms↗