Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Cell signalling - the proteomics of it all.

A challenge for biomedical scientists today is to arrive at an understanding of cellular behavior on a global scale. The advent of DNA microarrays has greatly facilitated discovery of gene expression profiles associated with different cellular states. The problem of understanding cellular signaling at the level of the interacting proteins is in some ways more challenging. Ashman et al. discuss the current methods available for studying protein interactions on a global scale, as well as directions for the future. Technical hurdles exist at many stages, from the isolation of protein complexes, to the determination of their composition, to the software and databases needed to analyze the results of large-scale, high-throughput datasets. Ashman et al. suggest that, with advances in technology and cooperation among academia and industry, a global protein interaction map that underlies cellular behavior will emerge as an essential resource for basic and applied research.

Computational Biology↗

High-throughput peptide mass fingerprinting of soybean seed proteins: automated workflow and utility of UniGene expressed sequence tag databases for protein identification.

Identification of anonymous proteins from two-dimensional (2-D) gels by peptide mass fingerprinting is one area of proteomics that can greatly benefit from a simple, automated workflow to minimize sample contamination and facilitate high-throughput sample processing. In this investigation we outline a workflow employing robotic automation at each step subsequent to 2-D gel electrophoresis. As proof-of-concept, 96 protein spots from a 2-D gel were analyzed using this approach. Whole protein (1 mg) from mature, dry soybean (Glycine max [L.] Merr.) cv. Jefferson seed was resolved by high resolution 2-D gel electrophoresis. Approximately 150 proteins were observed after staining with Coomassie Blue. The rather low number of detected proteins was due to the fact that the dynamic range of protein expression was greater than 100-fold. The most abundant proteins were seed storage proteins which in total represented over 60% of soybean seed protein. Using peptide mass fingerprinting 44 protein spots were identified. Identification of soybean proteins was greatly aided by the use of annotated, contiguous Expressed Sequence Tag (EST) databases which are available for public access (UniGene, ftp.ncbi.nih.gov/repository/UniGene/). Searches were orders of magnitude faster when compared to searches of unannotated EST databases and resulted in a higher frequency of valid, high-scoring matches. Some abundant, non seed storage proteins identified in this investigation include an isoelectric series of sucrose binding proteins, alcohol dehydrogenase and seed maturation proteins. This survey of anonymous seed proteins will serve as the basis for future comparative analysis of seed-filling in soybean as well as comparisons with other soybean varieties.

Databases, Genetic↗

Improved peptide identification in proteomics by two consecutive stages of mass spectrometric fragmentation.

MS-based proteomics usually involves the fragmentation of tryptic peptides (tandem MS or MS(2)) and their identification by searching protein sequence databases. In ion trap instruments fragments can be further fragmented and analyzed, a process termed MS/MS/MS or MS(3). Here, we report that efficient ion capture in a linear ion trap leads to MS(3) acquisition times and spectra quality similar to those for MS(2) experiments with conventional 3D ion traps. Fragmentation of N- or C-terminal ions resulted in informative and low-background spectra, even at subfemtomol levels of peptide. Typically C-terminal ions are chosen for further fragmentation, and the MS(3) spectrum greatly constrains the C-terminal amino acids of the peptide sequence. MS(3) spectra allow resolution of ambiguities in identification, a crucial problem in proteomics. Because of the sensitivity and rapid scan rates of the linear ion trap, several MS(3) spectra per peptide can be obtained even when sequencing very complex mixtures. We calculate the probability that an experimental MS(3) spectrum originates from fragmentation of a given N- or C-terminal ion of a peptide under consideration. This MS(3) identification score can be combined with the MS(2) scores of the precursor peptide from existing search engines. When MS(3) is performed on the linear ion trap-Fourier transform mass spectrometer combination, accurate peptide masses further increase confidence in peptide identification.

Algorithms↗

Differences between predicted and observed sequences in Saccharomyces cerevisiae.

We recently studied the protein composition of a Saccharomyces cerevisiae wine yeast strain (K310) of enological interest. About 2,500 spots of 8-250 kDa observed molecular mass were resolved by two-dimensional gel electrophoresis. Experimental molecular masses and isoelectric points were calculated for most of them. Twenty-seven proteins were subjected to Edman microsequencing. N-terminal sequences of 12/27 proteins were determined, whereas internal sequences of 6/27 proteins were obtained following in situ proteolysis. Comparison between the experimental data and those reported in the SWISS-PROT database revealed some differences between genotypic and phenotypic sequences. These are indicative of the changes a protein can undergo with respect to the primary structure coded by the genomic DNA. Our results highlight the need to complement genomic analysis with detailed proteomics in order to refine the vast amount of information provided by DNA sequencing and to find an exact correlation between genome and proteome.

Databases, Factual↗

Processed N-termini of mature proteins in higher eukaryotes and their major contribution to dynamic proteomics.

N-terminal-ubiquitinylation (NTU) is a newly discovered protein degradation pathway initiated by ubiquitin-tagging of the N-terminal alpha-amino group. We have used data from recent genomic studies, especially those on humans, to up-date and re-interpret biochemical data to identify the sequence features associated with NTU. We compared a mini-proteome for which experimental protein sequence is available with large-scale genomic data. We conclude that N-alpha-acetylation involves less than 30%, and not the widely assumed 90%, of the proteins encoded by any higher eukaryote genome, greatly increasing thereby the number of possible targets for NTU-mediated degradation. Next, straightforward rules linking the first N-terminal residues of any nascent polypeptides to the nature of their processed N-termini are established and dedicated prediction tool is made available at . We provide strong arguments indicating that the nature of the processed N-terminus is a major determinant factor of the half-life of the protein. We finally reveal that one third of the nuclear-encoded proteins starting with an unprocessed and unblocked methionine are at least one order of magnitude less stable than is average in higher eukaryotes. This appears to be the first common feature of proteins undergoing N-terminal ubiquitinylation. Hence, a pool of about 3000 proteins in each proteome could be unstable per se and tagged for rapid degradation via NTU.

Acetylation↗

Common interchange standards for proteomics data: Public availability of tools and schema.

The Proteomics Standards Initiative (PSI) aims to define community standards for data representation in proteomics and to facilitate data comparision, exchange and verification. To this end, a Level 1 Molecular Interaction XML data exchange format has been developed which has been accepted for publication and is freely available at the PSI website (http.//psidev.sf.net/). Several major protein interaction databases are already making data available in this format. A draft XML interchange format for mass spectrometry data has been written and is currently undergoing evaluation whilst work is ongoing to develop a proteomics data integration model, MIAPE.

Computational Biology↗

MannDB - a microbial database of automated protein sequence analyses and evidence integration for protein characterization.

BACKGROUND: MannDB was created to meet a need for rapid, comprehensive automated protein sequence analyses to support selection of proteins suitable as targets for driving the development of reagents for pathogen or protein toxin detection. Because a large number of open-source tools were needed, it was necessary to produce a software system to scale the computations for whole-proteome analysis. Thus, we built a fully automated system for executing software tools and for storage, integration, and display of automated protein sequence analysis and annotation data. DESCRIPTION: MannDB is a relational database that organizes data resulting from fully automated, high-throughput protein-sequence analyses using open-source tools. Types of analyses provided include predictions of cleavage, chemical properties, classification, features, functional assignment, post-translational modifications, motifs, antigenicity, and secondary structure. Proteomes (lists of hypothetical and known proteins) are downloaded and parsed from Genbank and then inserted into MannDB, and annotations from SwissProt are downloaded when identifiers are found in the Genbank entry or when identical sequences are identified. Currently 36 open-source tools are run against MannDB protein sequences either on local systems or by means of batch submission to external servers. In addition, BLAST against protein entries in MvirDB, our database of microbial virulence factors, is performed. A web client browser enables viewing of computational results and downloaded annotations, and a query tool enables structured and free-text search capabilities. When available, links to external databases, including MvirDB, are provided. MannDB contains whole-proteome analyses for at least one representative organism from each category of biological threat organism listed by APHIS, CDC, HHS, NIAID, USDA, USFDA, and WHO. CONCLUSION: MannDB comprises a large number of genomes and comprehensive protein sequence analyses representing organisms listed as high-priority agents on the websites of several governmental organizations concerned with bio-terrorism. MannDB provides the user with a BLAST interface for comparison of native and non-native sequences and a query tool for conveniently selecting proteins of interest. In addition, the user has access to a web-based browser that compiles comprehensive and extensive reports. Access to MannDB is freely available at http://manndb.llnl.gov/.

Algorithms↗

A functional proteomics approach to signal transduction.

The purpose of this review is to highlight how proteomics techniques can be used to answer specific questions related to signal transduction in a wide variety of systems. In our laboratory, we utilize proteomic technologies to elucidate signal transduction pathways involved in smooth muscle contraction and relaxation, cell growth and tumorigenesis, and the pathogenesis of malaria. We see the real application of this technology as a tool to enhance the power of existing approaches such as classical yeast and mouse genetics, tissue culture, protein expression systems, and site-directed mutagenesis. Our basic approach is to examine only those proteins that differ by some variable from the control sample. In this way, the number of proteins to be processed by electrophoresis, Edman degradation, or mass spectrometry is greatly reduced. In addition, since only those proteins that change in response to a given biological treatment are analyzed, the experimental outcome provides information about specific signaling pathways. Examples of typical experiments in our laboratory are measurement of changes in protein phosphorylation in response to treatment of cells with growth factors or specific drugs, characterization of proteins associated with a bait protein in a "pull-down" experiment, or measurement of changes in protein expression. Frequently, in these experiments, it is necessary to define complex protein mixtures. To achieve this goal, we utilize a variety of techniques to isolate specific types of proteins or "subproteomes" for further analysis. In this review, we discuss strategies used in our laboratory for studying signaling pathways, including subproteome isolation, proteome mining, and analysis of the phosphoproteome.

Amino Acid Sequence↗

Integrating 'top-down" and "bottom-up" mass spectrometric approaches for proteomic analysis of Shewanella oneidensis.

Here we present a comprehensive method for proteome analysis that integrates both intact protein measurement ("top-down") and proteolytic fragment characterization ("bottom-up") mass spectrometric approaches, capitalizing on the unique capabilities of each method. This integrated approach was applied in a preliminary proteomic analysis of Shewanella oneidensis, a metal-reducing microbe of potential importance to the field of bioremediation. Cellular lysates were examined directly by the "bottom-up" approach as well as fractionated via anion-exchange liquid chromatography for integrated studies. A portion of each fraction was proteolytically digested, with the resulting peptides characterized by on-line liquid chromatography/tandem mass spectrometry. The remaining portion of each fraction containing the intact proteins was examined by high-resolution Fourier transform mass spectrometry. This "top-down" technique provided direct measurement of the molecular masses for the intact proteins and thereby enabled confirmation of post-translational modifications, signal peptides, and gene start sites of proteins detected in the "bottom-up" experiments. A total of 868 proteins from virtually every functional class, including hypotheticals, were identified from this organism.

Amino Acid Sequence↗

High-efficiency on-line solid-phase extraction coupling to 15-150-microm-i.d. column liquid chromatography for proteomic analysis.

The ability to manipulate and effectively utilize small proteomic samples is important for analyses using liquid chromatography (LC) in combination with mass spectrometry (MS) and becomes more challenging for very low flow rates due to extra column volume effects on separation quality. Here we report on the use of commercial switching valves (150-microm channels) for implementing the on-line coupling of capillary LC columns operated at 10,000 psi with relatively large solid-phase extraction (SPE) columns. With the use of optimized column connections, switching modes, and SPE column dimensions, high-efficiency on-line SPE-capillary and nanoscale LC separations were obtained demonstrating peak capacities of approximately 1000 for capillaries having inner diameters between 15 and 150 microm. The on-line coupled SPE columns increased the sample processing capacity by approximately 400-fold for sample solution volume and approximately 10-fold for sample mass. The proteomic applications of this on-line SPE-capillary LC system were evaluated for analysis of both soluble and membrane protein tryptic digests. Using an ion trap tandem MS it was typically feasible to identify 1100-1500 unique peptides in a 5-h analysis. Peptides extracted from the SPE column and then eluted from the LC column covered a hydrophilicity/hydrophobicity range that included an estimated approximately 98% of all tryptic peptides. The SPE-capillary LC implementation also facilitates automation and enables use of both disposable SPE columns and electrospray emitters, providing a robust basis for automated proteomic analyses.

Amino Acid Sequence↗

Proteomic analysis of a thermostable superoxide dismutase from Bacillus stearothermophilus TLS33.

Thermophilic bacterium Bacillus stearothermophilus TLS33 isolated from a hot spring in Chiang Mai, Thailand produces an extracellular superoxide dismutase (SOD). SOD is a free radical metabolizing enzyme that protects the cell membrane from damage by the highly reactive superoxide free radicals. To identify the secreted SOD, we used the systematically proteomic approaches of two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) analysis and database searching. The bacterium was grown in a medium containing 0.1% w/v yeast extract and 0.1% w/v tryptone in 100% v/v base mixture at 65 degrees C for 72 h, by assessing their growth by protein and SOD activity. The bacterium produced the highest SOD activity at 65 degrees C for 48 h and the extracellular SOD was run on 2-D PAGE using broad range pH 3-10 immobilized pH gradients (IPGs) and narrow range pH 4-7 IPGs. The isoelectric point and molecular mass of the extracellular SOD were approximately 5.8 and 28 kDa, respectively. In addition, the NH(2)-terminal amino acid sequence was found to be P-F-E-L-P-A-L-P-Y-P-Y-D-A-L-E-P-P-I-I-D, which had a homology of approximately 85% to the Mn-SOD family and 65% to the Fe-SOD family.

Amino Acid Sequence↗

Parasite genome databases and web-based resources.

In the last decade, high-throughput genome sequencing and complementary techniques such as microarray and proteomics have generated, and will continue to generate, ever-increasing amounts of data. These technologies of gene discovery, expression, and functional analysis have been applied to a vast array of organisms, including parasites. In most instances, the data are freely available via the Internet, and researchers are becoming increasingly reliant on up-to-date, centralized data repositories to complement wet bench science. This chapter presents an overview of resources relevant to researchers with an interest in para-site genomics and biology. After briefly touching on some of the publicly available nucleotide and protein sequence as well as domain databases, the focus turns to parasite genome projects and associated Web-based resources. A list of parasite sequencing projects current at the time of writing, including relevant Web site addresses, is provided. The available resources range from network sites and project pages at sequencing institutes to databases that integrate and curate sequence data and associated annotation with diverse biological datasets. Particular attention is given to three databases, GeneDB (http://www.genedb.org/), PlasmoDB (http://plasmodb. org/), and tigr db, detailing the scope of each database and the tools available for data querying and retrieval.

Animals↗

2DDB - a bioinformatics solution for analysis of quantitative proteomics data.

BACKGROUND: We present 2DDB, a bioinformatics solution for storage, integration and analysis of quantitative proteomics data. As the data complexity and the rate with which it is produced increases in the proteomics field, the need for flexible analysis software increases. RESULTS: 2DDB is based on a core data model describing fundamentals such as experiment description and identified proteins. The extended data models are built on top of the core data model to capture more specific aspects of the data. A number of public databases and bioinformatical tools have been integrated giving the user access to large amounts of relevant data. A statistical and graphical package, R, is used for statistical and graphical analysis. The current implementation handles quantitative data from 2D gel electrophoresis and multidimensional liquid chromatography/mass spectrometry experiments. CONCLUSION: The software has successfully been employed in a number of projects ranging from quantitative liquid-chromatography-mass spectrometry based analysis of transforming growth factor-beta stimulated fi-broblasts to 2D gel electrophoresis/mass spectrometry analysis of biopsies from human cervix. The software is available for download at SourceForge.

Computational Biology↗

Proteome analysis using selective incorporation of isotopically labeled amino acids.

A method is described for identifying intact proteins from genomic databases using a combination of accurate molecular mass measurements and partial amino acid content. An initial demonstration was conducted for proteins isolated from Escherichia coli (E. coli) using a multiple auxotrophic strain of K12. Proteins extracted from the organism grown in natural isotopic abundance minimal medium and also minimal medium containing isotopically labeled leucine (Leu-D10), were mixed and analyzed by capillary isoelectric focusing (CIEF) coupled with Fourier transform ion cyclotron resonance mass spectrometry (FTICR). The incorporation of the isotopically labeled Leu residue has no effect on the CIEF separation of the protein, therefore both versions of the protein are observed within the same FTICR spectrum. The difference in the molecular mass of the natural isotopic abundance and Leu-D10 isotopically labeled proteins is used to determine the number of Leu residues present in that particular protein. Knowledge of the molecular mass and number of Leu residues present can be used to unambiguously identify the intact protein. Preliminary results show the efficacy of this method for unambiguously identifying proteins isolated from E. coli.

Amino Acids↗

Establishment of a two-dimensional electrophoresis map for Neospora caninum tachyzoites by proteomics.

Expressed proteins and antigens from Neospora caninum tachyzoites were studied by two-dimensional gel electrophoresis and immunoblot analysis combined with matrix-assisted laser desorption/ionization-time of flight mass spectrometry. Thirty-one spots corresponding to 20 different proteins were identified from N. caninum tachyzoites by peptide mass fingerprinting. Six proteins were identified from a N. caninum database (NTPase, 14-3-3 protein homologue, NcMIC1, NCDG1, NcGRA1 and NcGRA2), and 11 proteins were identified in closely related species using the T. gondii database (HSP70, HSP60, pyruvate kinase, tubulin alpha- and beta-chain, putative protein disulfide isomerase, enolase, actin, fructose-1,6-bisphosphatase, lactate dehydrogenase and glyceradehyde-3-phosphate dehydrogenase). One hundred and two antigen spots were observed using pH 4-7 IPG strips on immunoblot profiles. Among them, 17 spots corresponding to 11 antigenic proteins were identified from a N. caninum protein map. This study involved the construction of in-depth protein maps for N. caninum tachyzoites, which will be of value for studies of its pathogenesis, drug and vaccine development, and phylogenetic studies.

Animals↗

Studying fertilization in cell-free extracts: focusing on membrane/lipid raft functions and proteomics.

Xenopus oocytes, eggs, and embryos serve as an ideal model system to study several aspects of animal development (e.g., gametogenesis, fertilization, embryogenesis, and organogenesis). In particular, the Xenopus system has been extensively employed not only as a "living cell" system but also as a "cell-free" or "reconstitutional" system. In this chapter, we describe a protocol for studying the molecular mechanism of egg fertilization with the use of cell-free extracts and membrane/lipid rafts prepared from unfertilized, metaphase II-arrested Xenopus eggs. By using this experimental system, we have reconstituted a series of signal transduction events associated with egg fertilization, such as sperm-egg membrane interaction, activation of Src tyrosine kinase and phospholipase Cgamma, production of inositol trisphosphate, transient calcium release, and cell cycle transition. This type of reconstitutional system may allow us to perform focused proteomics (e.g., rafts) as well as global protein analysis (i.e., whole egg proteome) of fertilization in a cell-free manner. As one of these proteomics approaches, we provide a protocol for molecular identification of Xenopus egg raft proteins using mass spectrometry and database mining.

Animals↗