[Analysis of expressed proteome].
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Peptide mass fingerprint (PMF) matching is a high-throughput method used for protein spot identification in connection with two-dimensional gel electrophoresis (2DE). However, the success of PMF matching largely depends on whether the proteins to be identified exist in the database searched. Consequently, it is often necessary to apply other more sophisticated but also time-consuming technologies to generate sequence-tags for definitive protein identification. On the other hand, modern sequencing technologies are generating a large quantity of DNA sequences, first in unfinished form or with low genome coverage due to the time-consuming and thus limiting steps of finishing and annotation. We recently started to sequence the genome of Bacillus megaterium DSM 319, a bacterium of industrial interest. In this study, we demonstrate that a protein database generated from merely three-fold coverage, unfinished genomic sequences of this bacterium allows a fast and reliable protein spot identification solely based on PMF from high-throughput MALDI-TOF MS analysis. We further show that the strain-specific protein database from low coverage genomic sequence greatly outperforms the commonly used cross-species databases constructed from 13 completely sequenced Bacillus strains for protein spot identification via PMF.
Within the Human Proteome Organization (HUPO) Brain Proteome Project, a pilot study was launched with reference samples shipped to nine international laboratories (see Hamacher et al., this Special Issue) to evaluate different proteome approaches in neuroscience and to build up a first version of a brain protein database. One part of the study addresses quantitative proteome alterations between three developmental stages (embryonic day 16; postnatal day 7; 8 weeks) of mouse brains. Five brains per stage were differentially analyzed by 2-D DIGE using internal standardization and overlapping pH gradients (pH 4-7 and 6-9). In total, 214 protein spots showing stage-dependent intensity alterations (> two-fold) were detected, 56 of which were identified. Several of them, e.g. members of the dihydropyrimidinase family, are known to be associated with brain development. To feed the HUPO BPP brain protein database, a robust 2-D LC-MS/MS method was applied to murine postnatal day 7 and human post-mortem brain samples. Using MASCOT and the IPI database, 350 human and 481 mouse proteins could be identified by at least two different peptides. The data are accessible through the PRIDE database (http://www.ebi.ac.uk/pride/).
In this work, a novel approach based on proteomics is applied for the analysis of the three European marine mussel species: Mytilus edulis (ME), Mytilus galloprovincialis (MG) and Mytilus trossulus (MT), which are of interest in biotechnology and food industry. The proteomes of these species are poorly described in databases, are difficult to diagnose, and have a controversial taxonomy, To characterise species-specific peptides, we compared 51 matrix-assisted laser desorption/ioization-time of flight peptide mass maps generated from 6 random selected prominent spots derived from the two-dimensional electrophoresis analysis of foot protein extracts from several individuals. Minor species-specific differences in the peptide maps were detected in only one of the spots, corresponding to tropomyosin. Two peptides were unique to ME and MG individuals, whereas another peptide was present only in MT individuals. The sequence of these peptides was characterised by, nanoelectrospray ionization-ion trap (nanoESI-IT) tandem mass spectrometry (MS/MS) analysis followed by database searching and de novo sequence interpretation. We detected a single T to D amino acid substitution in MT tropomyosin. Unambiguous and highly-specific species identification was then demonstrated by analysing peptide extracts from tropomyosin spots by micro high-performande liquid chromatography (microHPL) ESI-IT mass spectrometry using the selected ion monitoring configuration, focused on these peptides, in continuous MS/MS operation. Our results suggest that proteomics may be successfully applied for the identification of species whose proteome is not present in databases.
The sarcomere is the major structural and functional unit of striated muscle. Approximately 65 different proteins have been associated with the sarcomere, and their exact composition defines the speed, endurance, and biology of each individual muscle. Past analyses relied heavily on electrophoretic and immunohistochemical techniques, which only allow the analysis of a small fraction of proteins at a time. Here we introduce a quantitative label-free, shotgun proteomics approach to differentially quantitate sarcomeric proteins from microgram quantities of muscle tissue in a fast and reliable manner by liquid chromatography and mass spectrometry. The high sequence similarity of some sarcomeric proteins poses a problem for shotgun proteomics because of limitations in subsequent database search algorithms in the exclusive assignment of peptides to specific isoforms. Therefore multiple sequence alignments were generated to improve the identification of isoform specific peptides. This methodology was used to compare the sarcomeric proteome of the extraocular muscle allotype to limb muscle. Extraocular muscles are a unique group of highly specialized muscles with distinct biochemical, physiological, and pathological properties. We were able to quantitate 40 sarcomeric proteins; although the basic sarcomeric proteins in extraocular muscle are similar to those in limb muscle, key proteins stabilizing the connection of the Z-bands to thin filaments and the costamere are augmented in extraocular muscle and may represent an adaptation to the eccentric contractions known to normally occur during eye movements. Furthermore, a number of changes are seen that closely relate to the unique nature of extraocular muscle.
Explore the source record for details and available documents.
Porins are outermembrane beta-barrel proteins. They have varied biological functionality ranging from phage receptors, immunogenicity, pathogenicity to apoptosis. However, only a small number has been structurally and functionally characterised. A validation mechanism and a database of porins would be useful for target selection in proteomics and structural genomics work. Here we report a validation mechanism developed for membrane porins. A database server for porins, PRNDS, has been created containing experimentally proven porins and likely putative porins. Each porin is validated and ranked using a weighted scoring system developed based on six, structure and sequence based criteria. The server also predicts possible porins.
The development of proteomics is a timely one for cardiovascular research. Analyses at the organ, subcellular, and molecular levels have revealed dynamic, complex, and subtle intracellular processes associated with heart and vascular disease. The power and flexibility of proteomic analyses, which facilitate protein separation, identification, and characterization, should hasten our understanding of these processes at the protein level. Properly applied, proteomics provides researchers with cellular protein "inventories" at specific moments in time, making it ideal for documenting protein modification due to a particular disease, condition, or treatment. This is accomplished through the establishment of species- and tissue-specific protein databases, providing a foundation for subsequent proteomic studies. Evolution of proteomic techniques has permitted more thorough investigation into molecular mechanisms underlying cardiovascular disease, facilitating identification not only of modified proteins but also of the nature of their modification. Continued development should lead to functional proteomic studies, in which identification of protein modification, in conjunction with functional data from established biochemical and physiological methods, has the ability to further our understanding of the interplay between proteome change and cardiovascular disease.
Proteomics offers a new set of tools for investigating parasites and parasite-associated disease. In this article, John Barrett, Jim Jefferies and Peter Brophy describe the key technologies involved, including two-dimensional gel electrophoresis, image analysis, biological mass spectroscopy and database searching. The potential applications of proteomics in drug and vaccine discovery are reviewed, as are possible future developments.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
SUMMARY: The Differentially Expressed Protein Database was designed to store the output of comparative proteomics studies and provides a publicly available query and analysis platform for data mining. The database contains information about more than 3000 differentially expressed proteins (DEPs) manually extracted from the published literature, including relevant biological, experimental and methodological elements. Tools for visualization and functional analysis of DEPs are provided via a user-friendly webinterface. AVAILABILITY: http://protchem.hunnu.edu.cn/depd/.
MOTIVATION: Taxonomic analysis of environmental microbial communities is now routinely performed thanks to advances in DNA sequencing. Determining the role of these communities in global biogeochemical cycles requires the identification of their metabolic functions, such as hydrogen oxidation, sulfur reduction, and carbon fixation. These functions can be directly inferred from metagenomics data, but in many environmental applications metabarcoding is still the method of choice. The reconstruction of metabolic functions from metabarcoding data and their integration into coarse-grained representations of biogeochemical cycles remains a difficult bioinformatics problem today. RESULTS: We developed a pipeline, called Tabigecy, which exploits taxonomic affiliations to predict metabolic functions constituting biogeochemical cycles. In a first step, Tabigecy uses the tool EsMeCaTa to predict consensus proteomes from input affiliations. To optimize this process, we generated a precomputed database containing information about 2404 taxa from UniProt. The consensus proteomes are searched using bigecyhmm, a newly developed Python package relying on Hidden Markov Models to identify key enzymes involved in metabolic function of biogeochemical cycles. The metabolic functions are then projected on coarse-grained representation of the cycles. We applied Tabigecy to two salt cavern datasets and validated its predictions with microbial activity and hydrochemistry measurements performed on the samples. The results highlight the utility of the approach to investigate the impact of microbial communities on biogeochemical processes. AVAILABILITY AND IMPLEMENTATION: The Tabigecy pipeline is available at https://github.com/ArnaudBelcour/tabigecy. The Python package bigecyhmm and the precomputed EsMeCaTa database are also separately available at https://github.com/ArnaudBelcour/bigecyhmm and https://doi.org/10.5281/zenodo.13354073, respectively.
Explore the source record for details and available documents.
MOTIVATION: The CEBS data repository is being developed to promote a systems biology approach to understand the biological effects of environmental stressors. CEBS will house data from multiple gene expression platforms (transcriptomics), protein expression and protein-protein interaction (proteomics), and changes in low molecular weight metabolite levels (metabolomics) aligned by their detailed toxicological context. The system will accommodate extensive complex querying in a user-friendly manner. CEBS will store toxicological contexts including the study design details, treatment protocols, animal characteristics and conventional toxicological endpoints such as histopathology findings and clinical chemistry measures. All of these data types can be integrated in a seamless fashion to enable data query and analysis in a biologically meaningful manner. RESULTS: An object model, the SysBio-OM (Xirasagar et al., 2004) has been designed to facilitate the integration of microarray gene expression, proteomics and metabolomics data in the CEBS database system. We now report SysTox-OM as an open source systems toxicology model designed to integrate toxicological context into gene expression experiments. The SysTox-OM model is comprehensive and leverages other open source efforts, namely, the Standard for Exchange of Nonclinical Data (http://www.cdisc.org/models/send/v2/index.html) which is a data standard for capturing toxicological information for animal studies and Clinical Data Interchange Standards Consortium (http://www.cdisc.org/models/sdtm/index.html) that serves as a standard for the exchange of clinical data. Such standardization increases the accuracy of data mining, interpretation and exchange. The open source SysTox-OM model, which can be implemented on various software platforms, is presented here. AVAILABILITY: A universal modeling language (UML) depiction of the entire SysTox-OM is available at http://cebs.niehs.nih.gov and the Rational Rose object model package is distributed under an open source license that permits unrestricted academic and commercial use and is available at http://cebs.niehs.nih.gov/cebsdownloads. Currently, the public toxicological data in CEBS can be queried via a web application based on the SysTox-OM at http://cebs.niehs.nih.gov CONTACT: xirasagars@saic.com SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Integr8 is a new web portal for exploring the biology of organisms with completely deciphered genomes. For over 190 species, Integr8 provides access to general information, recent publications, and a detailed statistical overview of the genome and proteome of the organism. The preparation of this analysis is supported through Genome Reviews, a new database of bacterial and archaeal DNA sequences in which annotation has been upgraded (compared to the original submission) through the integration of data from many sources, including the EMBL Nucleotide Sequence Database, the UniProt Knowledgebase, InterPro, CluSTr, GOA and HOGENOM. Integr8 also allows the users to customize their own interactive analysis, and to download both customized and prepared datasets for their own use. Integr8 is available at http://www.ebi.ac.uk/integr8.
We investigated human alternative protein isoforms of >2600 genes based on full-length cDNA clones and SwissProt. We classified the isoforms and examined their co-occurrence for each gene. Further, we investigated potential relationships between these changes and differential subcellular localization. The two most abundant patterns were the one with different C-terminal regions and the one with an internal insertion, which together account for 43% of the total. Although changes of the N-terminal region are less common than those of the C-terminal region, extension of the C-terminal region is much less common than that of the N-terminal region, probably because of the difficulty of removing stop codons in one isoform. We also found that there are some frequently used combinations of co-occurrence in alternative isoforms. We interpret this as evidence that there is some structural relationship which produces a repertoire of isoformal patterns. Finally, many terminal changes are predicted to cause differential subcellular localization, especially in targeting either peroxisomes or mitochondria. Our study sheds new light on the enrichment of the human proteome through alternative splicing and related events. Our database of alternative protein isoforms is available through the internet.
In classical proteomic studies, the searches in protein databases lead mostly to the identification of protein functions by homology due to the non-exhaustiveness of the protein databases. The quality of the identification depends on the studied organism, its complexity and its representation in the protein databases. Nevertheless, this basic function identification is insufficient for certain applications namely for the development of RNA-based gene-silencing strategies, commonly termed RNA interference (RNAi) in animals and post-transcriptional gene silencing (PTGS) in plants, that require an unambiguous identification of the targeted gene sequence. A PTGS strategy was considered in the study of the infection of Oryza sativa by the Rice Yellow Mottle Virus (RYMV). It is suspected that the RYMV recruits host proteins after its entry into plant cells to form a complex facilitating virus multiplication and spreading. The protein partners of this complex were identified by a classical proteomic approach, nano liquid chromatography tandem mass spectrometry. Among the identified proteins, several were retained for a PTGS strategy. Nevertheless most of the protein candidates appear to be members of multigenic families for which all paralog genes are not present in protein databases. Thus the identification of the real expressed paralog gene with classical protein database searches is impossible. Consequently, as the genome contains all genes and thus all paralog genes, a whole genome search strategy was developed to determine the specific expressed paralog gene. With this approach, the identification of peptides matching only a single gene, called discriminant peptides, allows definitive proof of the expression of this identified gene. This strategy has several requirements: (i) a genome completely sequenced and accessible; (ii) high protein sequence coverage. In the present work, through three examples, we report and validate for the first time a genome database search strategy to specifically identify paralog genes belonging to multigenic families expressed under specific conditions.