Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Protein probabilities in shotgun proteomics: evaluating different estimation methods using a semi-random sampling model.

The calculation of protein probabilities is one of the most intractable problems in large-scale proteomic research. Current available estimating methods, for example, ProteinProphet, PROT_PROBE, Poisson model and two-peptide hits, employ different models trying to resolve this problem. Until now, no efficient method is used for comparative evaluation of the above methods in large-scale datasets. In order to evaluate these various methods, we developed a semi-random sampling model to simulate large-scale proteomic data. In this model, the identified peptides were sampled from the designed proteins and their cross-correlation scores were simulated according to the results from reverse database searching. The simulated result of 18 control proteins was consistent with the experimental one, demonstrating the efficiency of our model. According to the simulated results of human liver sample, ProteinProphet returned slightly higher probabilities and lower specificity than real cases. PROT_PROBE was a more efficient method with higher specificity. Predicted results from a Poisson model roughly coincide with real datasets, and the method of two-peptide hits seems solid but imprecise. However, the probabilities of identified proteins are strongly correlated with several experimental factors including spectra number, database size and protein abundance distribution.

Chromatography, Liquid↗

Psychrophilicity of Bacillus psychrosaccharolyticus: a proteomic study.

Psychrophilicity of Gram-positive bacterium, Bacillus psychrosaccharolyticus was investigated in a proteomic approach. One hundred and thirty-one protein spots were analyzed by electrospray ionization-quadrupole-time of flight-tandem mass spectrometry and identified using an unpublished translated contig database as well as a nonredundant Gram-positive bacteria protein database from NCBI because of the lack of a genome sequence of this organism. Results focused on proteomic behavior of cold-response show that global up-regulation of metabolic functions and protective mechanism by stress responses might play a major role in psychrophilicity of B. psychrosaccharolyticus.

Bacillus↗

Genome annotating proteomics pipelines: available tools.

Proteomics based on tandem mass spectrometry is a powerful tool for identifying novel biomarkers and drug targets. Previously, a major bottleneck in high-throughput proteomics has been that the computational techniques needed to reliably identify proteins from proteomic data lagged behind the ability to collect the immense quantity of data generated. This is no longer the case, as fully automated pipelines for peptide and protein identification exist, and these are publicly and privately accessible. Such pipelines can automatically and rapidly generate high-confidence protein identifications from large datasets in a searchable format covering multiple experimental runs. However, the main challenge for the community now is to use these resources as they are, by taking full advantage of the pooling of information, so that the next barrier in our understanding of biology may be broken. There are currently two pipelines in the public domain that provide such potential: PeptideAtlas and the Genome Annotating Proteomic Pipeline. This review will introduce their features in the context of high-throughput proteomics, and provide indicative results as to their usefulness and usability through a side-by-side comparison of results obtained when processing a set of human plasma samples.

Animals↗

A predictive model for identifying proteins by a single peptide match.

MOTIVATION: Tandem mass-spectrometry of trypsin digests, followed by database searching, is one of the most popular approaches in high-throughput proteomics studies. Peptides are considered identified if they pass certain scoring thresholds. To avoid false positive protein identification, > or = 2 unique peptides identified within a single protein are generally recommended. Still, in a typical high-throughput experiment, hundreds of proteins are identified only by a single peptide. We introduce here a method for distinguishing between true and false identifications among single-hit proteins. The approach is based on randomized database searching and usage of logistic regression models with cross-validation. This approach is implemented to analyze three bacterial samples enabling recovery 68-98% of the correct single-hit proteins with an error rate of < 2%. This results in a 22-65% increase in number of identified proteins. Identifying true single-hit proteins will lead to discovering many crucial regulators, biomarkers and other low abundance proteins. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Identification of peptides in antimicrobial fractions of cheese extracts by electrospray ionization ion trap mass spectrometry coupled to a two-dimensional liquid chromatographic separation.

Electrospray ionization ion trap mass spectrometry (ESI-ITMS) coupled to a two-dimensional liquid chromatographic separation was applied to the identification of peptides in antimicrobial fractions of the aqueous extracts of nine Italian cheese varieties. In particular, the chromatographic fractions collected during a preliminary fast protein liquid chromatography (FPLC) separation on the cheese extracts were assayed for antimicrobial activity towards Lactobacillus sakei A15. Active fractions were subsequently analyzed by reversed-phase high-performance liquid chromatography electrospray ionization sequential mass spectrometry (HPLC/ESI)-ITMSn, with n up to 3. Peptide identification was then performed starting from a conventional proteomics approach based on tandem mass spectrometric (MS/MS) analysis followed by database searching. In many cases this strategy had to be integrated by a careful correlation between spectral information and predicted peptide fragmentation, in order to reach unambiguous identifications. When even this integrated approach failed, MS3 measurements provided decisive information on the amino acid sequence of some peptides, through fragmentation of pendant groups along the peptide chain. As a result, 45 peptides, all arising from hydrolysis of milk caseins, were identified in nine antimicrobial FPLC fractions of aqueous extracts obtained from five of the nine cheese varieties considered. Many of them corresponded to peptides already known to exhibit biological activity.

Amino Acid Sequence↗

Identifying the proteome: software tools.

The interest in proteomics has recently increased dramatically and proteomic methods are now applied to many problems in cell biology. The method of choice in proteomics for identifying and characterizing proteins is mass spectrometry combined with database searching. Software tools have been improved to increase the sensitivity of protein identification and methods for evaluating the search results have been incorporated

Databases, Factual↗

Role of accurate mass measurement (+/- 10 ppm) in protein identification strategies employing MS or MS/MS and database searching.

We describe the impact of advances in mass measurement accuracy, +/- 10 ppm (internally calibrated), on protein identification experiments. This capability was brought about by delayed extraction techniques used in conjunction with matrix-assisted laser desorption ionization (MALDI) on a reflectron time-of-flight (TOF) mass spectrometer. This work explores the advantage of using accurate mass measurement (and thus constraint on the possible elemental composition of components in a protein digest) in strategies for searching protein, gene, and EST databases that employ (a) mass values alone, (b) fragment-ion tagging derived from MS/MS spectra, and (c) de novo interpretation of MS/MS spectra. Significant improvement in the discriminating power of database searches has been found using only molecular weight values (i.e., measured mass) of > 10 peptide masses. When MALDI-TOF instruments are able to achieve the +/- 0.5-5 ppm mass accuracy necessary to distinguish peptide elemental compositions, it is possible to match homologous proteins having > 70% sequence identity to the protein being analyzed. The combination of a +/- 10 ppm measured parent mass of a single tryptic peptide and the near-complete amino acid (AA) composition information from immonium ions generated by MS/MS is capable of tagging a peptide in a database because only a few sequence permutations > 11 AA's in length for an AA composition can ever be found in a proteome. De novo interpretation of peptide MS/MS spectra may be accomplished by altering our MS-Tag program to replace an entire database with calculation of only the sequence permutations possible from the accurate parent mass and immonium ion limited AA compositions. A hybrid strategy is employed using de novo MS/MS interpretation followed by text-based sequence similarity searching of a database.

Animals↗

Comparative genomics of the pennate diatom Phaeodactylum tricornutum.

Diatoms are one of the most important constituents of phytoplankton communities in aquatic environments, but in spite of this, only recently have large-scale diatom-sequencing projects been undertaken. With the genome of the centric species Thalassiosira pseudonana available since mid-2004, accumulating sequence information for a pennate model species appears a natural subsequent aim. We have generated over 12,000 expressed sequence tags (ESTs) from the pennate diatom Phaeodactylum tricornutum, and upon assembly into a nonredundant set, 5,108 sequences were obtained. Significant similarity (E < 1E-04) to entries in the GenBank nonredundant protein database, the COG profile database, and the Pfam protein domains database were detected, respectively, in 45.0%, 21.5%, and 37.1% of the nonredundant collection of sequences. This information was employed to functionally annotate the P. tricornutum nonredundant set and to create an internet-accessible queryable diatom EST database. The nonredundant collection was then compared to the putative complete proteomes of the green alga Chlamydomonas reinhardtii, the red alga Cyanidioschyzon merolae, and the centric diatom T. pseudonana. A number of intriguing differences were identified between the pennate and the centric diatoms concerning activities of relevance for general cell metabolism, e.g. genes involved in carbon-concentrating mechanisms, cytosolic acetyl-Coenzyme A production, and fructose-1,6-bisphosphate metabolism. Finally, codon usage and utilization of C and G relative to gene expression (as measured by EST redundance) were studied, and preferences for utilization of C and CpG doublets were noted among the P. tricornutum EST coding sequences.

Animals↗

Listeria monocytogenes virulence and pathogenicity, a food safety perspective.

Several virulence factors of Listeria monocytogenes have been identified and extensively characterized at the molecular and cell biologic levels, including the hemolysin (listeriolysin O), two distinct phospholipases, a protein (ActA), several internalins, and others. Their study has yielded an impressive amount of information on the mechanisms employed by this facultative intracellular pathogen to interact with mammalian host cells, escape the host cell's killing mechanisms, and spread from one infected cell to others. In addition, several molecular subtyping tools have been developed to facilitate the detection of different strain types and lineages of the pathogen, including those implicated in common-source outbreaks of the disease. Despite these spectacular gains in knowledge, the virulence of L. monocytogenes as a foodborne pathogen remains poorly understood. The available pathogenesis and subtyping data generally fail to provide adequate insight about the virulence of field isolates and the likelihood that a given strain will cause illness. Possible mechanisms for the apparent prevalence of three serotypes (1/2a, 1/2b, and 4b) in human foodborne illness remain unidentified. The propensity of certain strain lineages (epidemic clones) to be implicated in common-source outbreaks and the prevalence of serotype 4b among epidemic-associated stains also remain poorly understood. This review first discusses current progress in understanding the general features of virulence and pathogenesis of L. monocytogenes. Emphasis is then placed on areas of special relevance to the organism's involvement in human foodborne illness, including (i) the relative prevalence of different serotypes and serotype-specific features and genetic markers; (ii) the ability of the organism to respond to environmental stresses of relevance to the food industry (cold, salt, iron depletion, and acid); (iii) the specific features of the major known epidemic-associated lineages; and (iv) the possible reservoirs of the organism in animals and the environment and the pronounced impact of environmental contamination in the food processing facilities. Finally, a discussion is provided on the perceived areas of special need for future research of relevance to food safety, including (i) theoretical modeling studies of niche complexity and contamination in the food processing facilities; (ii) strain databases for comprehensive molecular typing; and (iii) contributions from genomic and proteomic tools, including DNA microarrays for genotyping and expression signatures. Virulence-related genomic and proteomic signatures are expected to emerge from analysis of the genomes at the global level, with the support of adequate epidemiologic data and access to relevant strains.

Consumer Product Safety↗

MALDI quadrupole time-of-flight mass spectrometry: a powerful tool for proteomic research.

A MALDI QqTOF mass spectrometer has been used to identify proteins separated by one-dimensional or two-dimensional gel electrophoresis at the femtomole level. The high mass resolution and the high mass accuracy of this instrument in both MS and MS/MS modes allow identification of a protein either by peptide mass fingerprinting of the protein digest or from tandem mass spectra acquired by collision-induced dissociation of individual peptide precursors. A peptide mass map of the digest and tandem mass spectra of multiple peptide precursor ions can be acquired from the same sample in the course of a single experiment. Database searching and acquisition of MS and MS/MS spectra can be combined in an interactive fashion, increasing the information value of the analytical data. The approach has demonstrated its usefulness in the comprehensive characterization of protein in-gel digests, in the dissection of complex protein mixtures, and in sequencing of a low molecular weight integral membrane protein. Proteins can be identified in all types of sequence databases, including an EST database. Thus, MALDI QqTOF mass spectrometry promises to have remarkable potential for advancing proteomic research.

Amino Acid Sequence↗

Resources for integrative systems biology: from data through databases to networks and dynamic system models.

In systems biology, biologically relevant quantitative modelling of physiological processes requires the integration of experimental data from diverse sources. Recent developments in high-throughput methodologies enable the analysis of the transcriptome, proteome, interactome, metabolome and phenome on a previously unprecedented scale, thus contributing to the deluge of experimental data held in numerous public databases. In this review, we describe some of the databases and simulation tools that are relevant to systems biology and discuss a number of key issues affecting data integration and the challenges these pose to systems-level research.

Animals↗

Proteome analysis of rat polymorphonuclear leukocytes: a two-dimensional electrophoresis/mass spectrometry approach.

The development of a two-dimensional (2-D) map of rat polymorphonuclear (PMN) leukocytes is here reported for the first time. The map is built up by utilizing a wide immobilized pH gradient (IPG), pH 3-10, in the first dimension and also a narrower IPG pH 4.5-8.5 gradient. In addition, the map is constructed by adopting the most recent protocols in 2-D mapping, which call for reduction and alkylation of the sample prior to the start of any electrophoretic step, including the IPG dimension. Fifty-two major protein spots have been so far identified by utilizing both matrix assisted laser desorption/ionization-time of flight (MALDI-TOF) and electrospray quadrupole (Q)-TOF mass spectrometry. A large number of house-keeping and cytoskeleton proteins were detected, together with proteins which are specific to PMN organelles or related to PMN functions such as phagocytosis and chemotaxis. The results obtained demonstrate the possibility of obtaining a single 2-D gel based proteomic map of PMN with representative proteins from different cellular compartments, also including membrane components, allowing the study of PMN protein expression on a proteome-wide scale. The aim of this project is to build an extensive database of such proteins, to be utilized for future studies where the expression of PMN proteins is used as a disease- or drug treatment marker.

Animals↗

Abundant protein domains occur in proportion to proteome size.

BACKGROUND: Conserved domains in proteins have crucial roles in protein interactions, DNA binding, enzyme activity and other important cellular processes. It will be of interest to determine the proportions of genes containing such domains in the proteomes of different eukaryotes. RESULTS: The average proportion of conserved domains in each of five eukaryote genomes was calculated. In pairwise genome comparisons, the ratio of genes containing a given conserved domain in the two genomes on average reflected the ratio of the predicted total gene numbers of the two genomes. These ratios have been verified using a repository of databases and one of its subdivisions of conserved domains. CONCLUSIONS: Many conserved domains occur as a constant proportion of proteome size across the five sequenced eukaryotic genomes. This raises the possibility that this proportion is maintained because of functional constraints on interacting domains. The universality of the ratio in the five eukaryotic genomes attests to its potential importance.

Animals↗

Peptidomic and proteomic analyses of the systemic immune response of Drosophila.

Insects have developed an efficient host defense against microorganisms, which involves humoral and cellular mechanisms. Numerous data highlight similarities between defense responses of insects and innate immunity of mammals. The fruit fly, Drosophila melanogaster, is a favorable model system for the analysis of the first line defense against microorganisms. Taking advantages of improvements in mass spectrometry (MS), two-dimensional (2D) gel electrophoresis and bioinformatics, differential analyses of blood content (hemolymph) from immune-challenged versus control Drosophila were performed. Two strategies were developed: (i) peptidomic analyses through matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) MS and high performance liquid chromatography for molecules below 15 kDa, and (ii) proteomic studies based on 2D gel electrophoresis, MALDI-TOF fingerprinting and database searches, for compounds of greater molecular masses. The peptidomic strategy led to the detection of a large number of peptides induced in the hemolymph of challenged flies as compared to controls. Of these, 28 were characterized, amongst which were antimicrobial peptides. The 2D gel electrophoresis strategy led to the detection of 70 spots differentially regulated by at least fivefold after microbial infection. This approach yielded the identity of a series of proteins that were related to the Drosophila immune response, such as proteases, protease inhibitors, prophenoloxydase-activating enzymes, serpins and a Gram-negative binding protein-like protein. This strategy also brought to light new candidates with a potential function in the immune response (odorant-binding protein, peptidylglycine alpha-hydroxylating monooxygenase and transferrin). Interestingly, several molecules resulting from the cleavage of proteins were detected after a fungal infection. Together, peptidomic and proteomic analyses represent new tools to characterize molecules involved in the innate immune reactions of Drosophila.

Animals↗

Parallel processing of large datasets from NanoLC-FTICR-MS measurements.

A new approach for automatic parallel processing of large mass spectral datasets in a distributed computing environment is demonstrated to significantly decrease the total processing time. The implementation of this novel approach is described and evaluated for large nanoLC-FTICR-MS datasets. The speed benefits are determined by the network speed and file transfer protocols only and allow almost real-time analysis of complex data (e.g., a 3-gigabyte raw dataset is fully processed within 5 min). Key advantages of this approach are not limited to the improved analysis speed, but also include the improved flexibility, reproducibility, and the possibility to share and reuse the pre- and postprocessing strategies. The storage of all raw data combined with the massively parallel processing approach described here allows the scientist to reprocess data with a different set of parameters (e.g., apodization, calibration, noise reduction), as is recommended by the proteomics community. This approach of parallel processing was developed in the Virtual Laboratory for e-Science (VL-e), a science portal that aims at allowing access to users outside the computer research community. As such, this strategy can be applied to all types of serially acquired large mass spectral datasets such as LC-MS, LC-MS/MS, and high-resolution imaging MS results.

Algorithms↗

Proteomics and Beyond: a report on the 3rd Annual Spring Workshop of the HUPO-PSI 21-23 April 2006, San Francisco, CA, USA.

The theme of the third annual Spring workshop of the HUPO-PSI was "proteomics and beyond" and its underlying goal was to reach beyond the boundaries of the proteomics community to interact with groups working on the similar issues of developing interchange standards and minimal reporting requirements. Significant developments in many of the HUPO-PSI XML interchange formats, minimal reporting requirements and accompanying controlled vocabularies were reported, with many of these now feeding into the broader efforts of the Functional Genomics Experiment (FuGE) data model and Functional Genomics Ontology (FuGO) ontologies.

Animals↗