Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Comparison of theoretical proteomes: identification of COGs with conserved and variable pI within the multimodal pI distribution.

BACKGROUND: Theoretical proteome analysis, generated by plotting theoretical isoelectric points (pI) against molecular masses of all proteins encoded by the genome show a multimodal distribution for pI. This multimodal distribution is an effect of allowed combinations of the charged amino acids, and not due to evolutionary causes. The variation in this distribution can be correlated to the organisms ecological niche. Contributions to this variation maybe mapped to individual proteins by studying the variation in pI of orthologs across microorganism genomes. RESULTS: The distribution of ortholog pI values showed trimodal distributions for all prokaryotic genomes analyzed, similar to whole proteome plots. Pairwise analysis of pI variation show that a few COGs are conserved within, but most vary between, the acidic and basic regions of the distribution, while molecular mass is more highly conserved. At the level of functional grouping of orthologs, five groups vary significantly from the population of orthologs, which is attributed to either conservation at the level of sequences or a bias for either positively or negatively charged residues contributing to the function. Individual COGs conserved in both the acidic and basic regions of the trimodal distribution are identified, and orthologs that best represent the variation in levels of the acidic and basic regions are listed. CONCLUSION: The analysis of pI distribution by using orthologs provides a basis for resolution of theoretical proteome comparison at the level of individual proteins. Orthologs identified that significantly vary between the major acidic and basic regions maybe used as representative of the variation of the entire proteome.

Bacterial Proteins↗

Exploring and exploiting bacterial proteomes.

The plethora of data now available from bacterial genome sequencing has opened a wealth of new research opportunities. Many of these have been reviewed in preceding chapters. Genomics alone, however, cannot capture a biological snapshot from an organism at a given point in time. The genome itself is static, and it is the changes in expression of genes, leading to the production of functional proteins, which allows an organism to survive and adapt to a constantly changing environment. Proteomics is the term used to describe the global analysis of proteins involved in a particular biological process. Such processes may be analyzed via comparative studies that examine bacterial strain differences, both phenotypic and genetic, bacteria grown under nutrient limiting conditions, growth phase, temperature, or in the presence of chemical compounds, such as antibiotics. Proteomics also provides the researcher with a tool to begin characterizing the functions of the vast proportion of "hypothetical" or "unknown" proteins elucidated from genome sequencing and database comparisons. For example, study of protein-protein, protein-ligand, protein-substrate, and protein-nucleic acid interactions for a given target protein may all help to define the functions of previously unknown proteins. Furthermore, genetic manipulation combined with proteomics technologies can provide an understanding of how gene expression is regulated. This chapter examines the technologies used in proteome analysis and the applications of proteomics to microbiological research, with an emphasis on clinically-relevant bacteria.

Bacterial Proteins↗

Proteome-wide functional classification and identification of prokaryotic transmembrane proteins by transmembrane topology similarity comparison.

We propose a new method for classifying and identifying transmembrane (TM) protein functions in proteome-scale by applying a single-linkage clustering method based on TM topology similarity, which is calculated simply from comparing the lengths of loop regions. In this study, we focused on 87 prokaryotic TM proteomes consisting of 31 proteobacteria, 22 gram-positive bacteria, 19 other bacteria, and 15 archaea. Prior to performing the clustering, we first categorized individual TM protein sequences as "known," "putative" (similar to "known" sequences), or "unknown" by using the homology search and the sequence similarity comparison against SWISS-PROT to assess the current status of the functional annotation of the TM proteomes based on sequence similarity only. More than three-quarters, that is, 75.7% of the TM protein sequences are functionally "unknown," with only 3.8% and 20.5% of them being classified as "known" and "putative," respectively. Using our clustering approach based on TM topology similarity, we succeeded in increasing the rate of TM protein sequences functionally classified and identified from 24.3% to 60.9%. Obtained clusters correspond well to functional superfamilies or families, and the functional classification and identification are successfully achieved by this approach. For example, in an obtained cluster of TM proteins with six TM segments, 109 sequences out of 119 sequences annotated as "ATP-binding cassette transporter" are properly included and 122 "unknown" sequences are also contained.

Algorithms↗

Secreted protein prediction system combining CJ-SPHMM, TMHMM, and PSORT.

To increase the coverage of secreted protein prediction, we describe a combination strategy. Instead of using a single method, we combine Hidden Markov Model (HMM)-based methods CJ-SPHMM and TMHMM with PSORT in secreted protein prediction. CJ-SPHMM is an HMM-based signal peptide prediction method, while TMHMM is an HMM-based transmembrane (TM) protein prediction algorithm. With CJ-SPHMM and TMHMM, proteins with predicted signal peptide and without predicted TM regions are taken as putative secreted proteins. This HMM-based approach predicts secreted protein with Ac (Accuracy) at 0.82 and Cc (Correlation coefficient) at 0.75, which are similar to PSORT with Ac at 0.82 and Cc at 0.76. When we further complement the HMM-based method, i.e., CJ-SPHMM + TMHMM with PSORT in secreted protein prediction, the Ac value is increased to 0.86 and the Cc value is increased to 0.81. Taking this combination strategy to search putative secreted proteins from the International Protein Index (IPI) maintained at the European Bioinformatics Institute (EBI), we constructed a putative human secretome with 5235 proteins. The prediction system described here can also be applied to predicting secreted proteins from other vertebrate proteomes.

Computational Biology↗

A common open representation of mass spectrometry data and its application to proteomics research.

A broad range of mass spectrometers are used in mass spectrometry (MS)-based proteomics research. Each type of instrument possesses a unique design, data system and performance specifications, resulting in strengths and weaknesses for different types of experiments. Unfortunately, the native binary data formats produced by each type of mass spectrometer also differ and are usually proprietary. The diverse, nontransparent nature of the data structure complicates the integration of new instruments into preexisting infrastructure, impedes the analysis, exchange, comparison and publication of results from different experiments and laboratories, and prevents the bioinformatics community from accessing data sets required for software development. Here, we introduce the 'mzXML' format, an open, generic XML (extensible markup language) representation of MS data. We have also developed an accompanying suite of supporting programs. We expect that this format will facilitate data management, interpretation and dissemination in proteomics research.

Database Management Systems↗

Database for renal collecting duct regulatory and transporter proteins.

The mammalian kidney collecting duct plays an important role in the fine regulation of Na, K, water, and acid-base balance. Functional genomic and proteomic studies of the kidney offer new opportunities in the understanding of renal physiology and pathophysiology, and the collecting duct is an appropriate target tissue because of the relative simplicity of its cells and the ease of isolating or culturing large numbers of collecting duct cells. Study of the collecting duct includes assessment of gene expression and protein regulation and abundance. For example, DNA and protein microarrays can be used to quantitate gene expression and protein regulation and abundance under varying physiological conditions. An Internet-accessible database has been devised for major collecting duct proteins involved in transport and regulation of cellular processes. The individual proteins included in this database are those culled from literature searches and from previously published studies involving cDNA arrays and serial analysis of gene expression (SAGE). Design of microarray targets for the study of kidney collecting duct tissues is facilitated by the database, which includes links to curated base pair and amino acid sequence data, relevant literature, and related databases. Use of the database is illustrated by a search for water channel proteins, aquaporins, and by a subsequent search for vasopressin receptors. Links are shown to the literature and to sequence data for human, rat, and mouse, as well as to relevant web-based resources. Extension of the database is dynamic and is done through a maintenance interface. This permits creation of new categories, updating of existing entries, and addition of new ones.

Animals↗

Identification of alkaline proteins that are differentially expressed in an overgrowth-mediated growth arrest and cell death of Escherichia coli by proteomic methodologies.

The available Escherichia coli genome sequences offer an opportunity to further expand our understanding of this bacterium. In the current study, we present a rapid method for the isolation of bacterial alkaline proteins using acid incubation, purification and protein array by 2-DE, followed by protein identification using MS. Fifty-seven proteins were randomly chosen, in which 55 were identified by a database searching of MS data. The searching results showed that most of these alkaline proteins were involved in special functions within the cell, suggesting that alkaline proteome is an ideal fraction for an understanding of their special functions. Furthermore, alkaline proteomes were compared between the period of majority live bacteria (18-h culture), the period of similar amount of live and dead bacteria (30-h culture) and the period of majority dead bacteria (48-h culture). Six proteins were identified as differentially expressed targets, in which putative transcriptional regulator and superoxide dismutase genes were cloned and expressed for antiserum preparations. The antisera were applied for the confirmation of results obtained from 2-DE. The presented data clearly reveal that alkaline proteome analysis by 2-DE with MS plays an important role in the understanding of protein functions within the cell, and six alkaline proteins are determined as key ones in an overgrowth-mediated growth cycle of E. coli.

Cell Death↗

Proteomic analysis of steady-state nuclear hormone receptor coactivator complexes.

We report our initial efforts in the analysis of endogenous nuclear receptor coactivator complexes as a research bridging strand of the Nuclear Receptor Signaling Atlas (NURSA) (www.NURSA.org). A proteomic approach is used to systematically isolate a variety of coactivator complexes using HeLa cells as a model cell line and to identify the coactivator-associated proteins with mass spectrometry. We have isolated and identified seven coactivator complexes including the p160 steroid receptor coactivator family, cAMP response element binding protein-binding protein, p300, coactivator of activating protein-1 and estrogen receptors, and E6 papillomavirus-associated protein. The newly identified coactivator-associated proteins provide unbiased clues and links for understanding of the endogenous hormone receptor coregulator network and its regulation. We hope that the electronic availability of these data to the general scientific community will facilitate generation and testing of new hypotheses to further our understanding of nuclear receptor signaling and coactivator functions.

Antibodies↗

The dual origin of the yeast mitochondrial proteome.

We propose a scheme for the origin of mitochondria based on phylogenetic reconstructions with more than 400 yeast nuclear genes that encode mitochondrial proteins. Half of the yeast mitochondrial proteins have no discernable bacterial homologues, while one-tenth are unequivocally of alpha-proteobacterial origin. These data suggest that the majority of genes encoding yeast mitochondrial proteins are descendants of two different genomic lineages that have evolved in different modes. First, the ancestral free-living alpha-proteobacterium evolved into an endosymbiont of an anaerobic host. Most of the ancestral bacterial genes were lost, but a small fraction of genes supporting bioenergetic and translational processes were retained and eventually transferred to what became the host nuclear genome. In a second, parallel mode, a larger number of novel mitochondrial genes were recruited from the nuclear genome to complement the remaining genes from the bacterial ancestor. These eukaryotic genes, which are primarily involved in transport and regulatory functions, transformed the endosymbiont into an ATP-exporting organelle.

Alphaproteobacteria↗

Functional proteomics and correlated signaling pathway of the thermophilic bacterium Bacillus stearothermophilus TLS33 under cold-shock stress.

The thermophilic bacterium Bacillus stearothermophilus TLS33 was examined under cold-shock stress by a proteomic approach to gain a better understanding of the protein synthesis and complex regulatory pathways of bacterial adaptation. After downshift in the temperature from 65 degrees C, the optimal growth temperature for this bacterium, to 37 degrees C and 25 degrees C for 2 h, we used the high-throughput techniques of proteomic analysis combining 2-DE and MS to identify 53 individual proteins including differentially expressed proteins. The bioinformatics database was used to search the biological functions of proteins and correlate these with gene homology and metabolic pathways in cell protection and adaptation. Eight cold-shock-induced proteins were shown to have markedly different protein expression: glucosyltransferase, anti-sigma B (sigma(B)) factor, Mrp protein homolog, dihydroorthase, hypothetical transcriptional regulator in FeuA-SigW intergenic region, RibT protein, phosphoadenosine phosphosulfate reductase and prespore-specific transcriptional activator RsfA. Interestingly, six of these cold-shock-induced proteins are correlated with the signal transduction pathway of bacterial sporulation. This study aims to provide a better understanding of the functional adaptation of this bacterium to environmental cold-shock stress.

Amino Acid Sequence↗

Preliminary 2-D chromatographic investigation of the human stem cell proteome.

Stem cells represent a promising tool for the treatment of various hematopoietic diseases. In order to identify stem cell-specific proteins, the proteome of human stem cells from umbilical cord blood was explored for the first time. For this purpose, the crude lysate of 4 x 10(5) CD34+ cells was subjected to in solution trypsin digestion. The resulting peptides were then separated via cation exchange followed by reversed phase chromatography and analyzed by nanospray MS/MS. Database search revealed a total of 215 proteins which could be reliably identified. To obtain a more complete picture of the human stell cell proteome and to also access low abundant proteins, pooling of more than one CD34+ preparations seems necessary in order to increase the cell number and thus the protein content.

Antigens, CD34↗

Increased frequency of cysteine, tyrosine, and phenylalanine residues since the last universal ancestor.

Analysis of extant proteomes has the potential of revealing how amino acid frequencies within proteins have evolved over biological time. Evidence is presented here that cysteine, tyrosine, and phenylalanine residues have substantially increased in frequency since the three primary lineages diverged more than three billion years ago. This inference was derived from a comparison of amino acid frequencies within conserved and non-conserved residues of a set of proteins dating to the last universal ancestor in the face of empirical knowledge of the relative mutability of these amino acids. The under-representation of these amino acids within last universal ancestor proteins relative to their modern descendants suggests their late introduction into the genetic code. Thus, it appears that extant ancient proteins contain evidence pertaining to early events in the formation of biological systems.

Amino Acid Sequence↗

Counting the zinc-proteins encoded in the human genome.

Metalloproteins are proteins capable of binding one or more metal ions, which may be required for their biological function, or for regulation of their activities or for structural purposes. Genome sequencing projects have provided a huge number of protein primary sequences, but, even though several different elaborate analyses and annotations have been enabled by a rich and ever-increasing portfolio of bioinformatic tools, metal-binding properties remain difficult to predict as well as to investigate experimentally. Consequently, the present knowledge about metalloproteins is only partial. The present bioinformatic research proposes a strategy to answer the question of how many and which proteins encoded in the human genome may require zinc for their physiological function. This is achieved by a combination of approaches, which include: (i) searching in the proteome for the zinc-binding patterns that, on their turn, are obtained from all available X-ray data; (ii) using libraries of metal-binding protein domains based on multiple sequence alignments of known metalloproteins obtained from the Pfam database; and (iii) mining the annotations of human gene sequences, which are based on any type of information available. It is found that 1684 proteins in the human proteome are independently identified by all three approaches as zinc-proteins, 746 are identified by two, and 777 are identified by only one method. By assuming that all proteins identified by at least two approaches are truly zinc-binding and inspecting the proteins identified by a single method, it can be proposed that ca. 2800 human proteins are potentially zinc-binding in vivo, corresponding to 10% of the human proteome, with an uncertainty of 400 sequences. Available functional information suggests that the large majority of human zinc-binding proteins are involved in the regulation of gene expression. The most abundant class of zinc-binding proteins in humans is that of zinc-fingers, with Cys4 and Cys2His2 being the most common types of coordination environment.

Computational Biology↗

A reagent resource to identify proteins and peptides of interest for the cancer community: a workshop report.

On the basis of discussions with representatives from all sectors of the cancer research community, the National Cancer Institute (NCI) recognizes the immense opportunities to apply proteomics technologies to further cancer research. Validated and well characterized affinity capture reagents (e.g. antibodies, aptamers, and affibodies) will play a key role in proteomics research platforms for the prevention, early detection, treatment, and monitoring of cancer. To discuss ways to develop new resources and optimize current opportunities in this area, the NCI convened the "Proteomic Technologies Reagents Resource Workshop" in Chicago, IL on December 12-13, 2005. The workshop brought together leading scientists in proteomics research to discuss model systems for evaluating and delivering resources for reagents to support MS and affinity capture platforms. Speakers discussed issues and identified action items related to an overall vision for and proposed models for a shared proteomics reagents resource, applications of affinity capture methods in cancer research, quality control and validation of affinity capture reagents, considerations for target selection, and construction of a reagents database. The meeting also featured presentations and discussion from leading private sector investigators on state-of-the-art technologies and capabilities to meet the user community's needs. This workshop was developed as a component of the NCI's Clinical Proteomics Technologies Initiative for Cancer, a coordinated initiative that includes the establishment of reagent resources for the scientific community. This workshop report explores various approaches to develop a framework that will most effectively fulfill the needs of the NCI and the cancer research community.

Biomedical Research↗

Additional paper: computational resources for metabolomics.

Metabolomics, a comprehensive extension of traditional targeted metabolite analysis, has recently attracted much attention as the biological jigsaw puzzle's missing piece that can complement transcriptome and proteome analysis. This tutorial survey introduces practical web resources with special emphasis on the computational aspects involved in processing and navigating metabolome data. The introduced materials are also accessible from the author's web directory (Atomic Reconstruction of Metabolism or ARM).

Algorithms↗

Peptide mass fingerprinting: identification of proteins by MALDI-TOF.

MALDI-TOF peptide mass fingerprinting (PMF) is the fastest and cheapest method of protein identification; the studied genome is sequenced and annotated, and the protein is amenable to separation and detection in 2D gel electrophoresis. In plant proteomics there are two main difficulties: few plant genomes are sequenced, and major contaminants are non-plant specific. This chapter describes the classical "bottom-up" method (i.e., from peptide to protein identification) of gel cutting, in-gel digestion, peptide recovery and purification, MALDI-TOF mass spectrometry, and critical survey of protein database queries.

Acrylic Resins↗

Abundance of intrinsically unstructured proteins in P. falciparum and other apicomplexan parasite proteomes.

Preliminary sequence analysis of Plasmodium falciparum has shown that the proteome of this organism is enriched in intrinsically unstructured proteins (IUPs), which are either completely disordered or contain large disordered regions. IUPs have been characterized as a unique class of proteins that plays an important role in biology and disease. In this study, the IUP contents in the proteomes of apicomplexan parasites, especially the proteome of P. falciparum and its various life cycle stages, have been evaluated with DisEMBL-1.4. Compared with other proteomes, apicomplexan species are extremely abundant in proteins containing long disordered regions, and the IUP contents in mammalian Plasmodium species are higher than in most other apicomplexan parasites. The proteome of the P. falciparum sporozoite appears to be distinct from the other life cycle stages in having an even higher content of disordered proteins. The abundance of IUPs in the P. falciparum proteome correlates with its enrichment in repetitive sequences. The structural plasticity of IUPs, which allows promiscuous binding interactions, may favour parasite survival both by inhibiting the generation of effective high affinity antibody responses and by facilitating the interactions with host molecules necessary for attachment and invasion of host cells.

Animals↗

Informatics for protein identification by mass spectrometry.

High throughput protein analysis (i.e., proteomics) first became possible when sensitive peptide mass mapping techniques were developed, thereby allowing for the possibility of identifying and cataloging most 2D gel electrophoresis spots. Shortly thereafter a few groups pioneered the idea of identifying proteins by using peptide tandem mass spectra to search protein sequence databases. Hence, it became possible to identify proteins from very complex mixtures. One drawback to these latter techniques is that it is not entirely straightforward to make matches using tandem mass spectra of peptides that are modified or have sequences that differ slightly from what is present in the sequence database that is being searched. This has been part of the motivation behind automated de novo sequencing programs that attempt to derive a peptide sequence regardless of its presence in a sequence database. The sequence candidates thus generated are then subjected to homology-based database search programs (e.g., BLAST or FASTA). These homology search programs, however, were not developed with mass spectrometry in mind, and it became necessary to make minor modifications such that mass spectrometric ambiguities can be taken into account when comparing query and database sequences. Finally, this review will discuss the important issue of validating protein identifications. All of the search programs will produce a top ranked answer; however, only the credulous are willing to accept them carte blanche.

Amino Acid Sequence↗