Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

MultiTag: multiple error-tolerant sequence tag search for the sequence-similarity identification of proteins by mass spectrometry.

The characterization of proteomes by mass spectrometry is largely limited to organisms with sequenced genomes. To identify proteins from organisms with unsequenced genomes, database sequences from related species must be employed for sequence-similarity protein identifications. Peptide sequence tags (Mann, 1994) have been used successfully for the identification of proteins in sequence databases using partially interpreted tandem mass spectra of tryptic peptides. We have extended the ability of sequence tag searching to the identification of proteins whose sequences are yet unknown but are homologous to known database entries. The MultiTag method presented here assigns statistical significance to matches of multiple error-tolerant sequence tags to a database entry and ranks alignments by their significance. The MultiTag approach has the distinct advantage over other sequence-similarity approaches of being able to perform sequence-similarity identifications using only very short (2-4) amino acid residue stretches of peptide sequences, rather than complete peptide sequences deduced by de novo interpretation of tandem mass spectra. This feature facilitates the identification of low abundance proteins, since noisy and low-intensity tandem mass spectra can be utilized.

Alcohol Dehydrogenase↗

Toward a comprehensive atlas of the physical interactome of Saccharomyces cerevisiae.

Defining protein complexes is critical to virtually all aspects of cell biology. Two recent affinity purification/mass spectrometry studies in Saccharomyces cerevisiae have vastly increased the available protein interaction data. The practical utility of such high throughput interaction sets, however, is substantially decreased by the presence of false positives. Here we created a novel probabilistic metric that takes advantage of the high density of these data, including both the presence and absence of individual associations, to provide a measure of the relative confidence of each potential protein-protein interaction. This analysis largely overcomes the noise inherent in high throughput immunoprecipitation experiments. For example, of the 12,122 binary interactions in the general repository of interaction data (BioGRID) derived from these two studies, we marked 7504 as being of substantially lower confidence. Additionally, applying our metric and a stringent cutoff we identified a set of 9074 interactions (including 4456 that were not among the 12,122 interactions) with accuracy comparable to that of conventional small scale methodologies. Finally we organized proteins into coherent multisubunit complexes using hierarchical clustering. This work thus provides a highly accurate physical interaction map of yeast in a format that is readily accessible to the biological community.

Cluster Analysis↗

The Protein Information Resource.

The Protein Information Resource (PIR) is an integrated public resource of protein informatics that supports genomic and proteomic research and scientific discovery. PIR maintains the Protein Sequence Database (PSD), an annotated protein database containing over 283 000 sequences covering the entire taxonomic range. Family classification is used for sensitive identification, consistent annotation, and detection of annotation errors. The superfamily curation defines signature domain architecture and categorizes memberships to improve automated classification. To increase the amount of experimental annotation, the PIR has developed a bibliography system for literature searching, mapping, and user submission, and has conducted retrospective attribution of citations for experimental features. PIR also maintains NREF, a non-redundant reference database, and iProClass, an integrated database of protein family, function, and structure information. PIR-NREF provides a timely and comprehensive collection of protein sequences, currently consisting of more than 1 000 000 entries from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and PDB. The PIR web site (http://pir.georgetown.edu) connects data analysis tools to underlying databases for information retrieval and knowledge discovery, with functionalities for interactive queries, combinations of sequence and text searches, and sorting and visual exploration of search results. The FTP site provides free download for PSD and NREF biweekly releases and auxiliary databases and files.

Amino Acid Sequence↗

Automated construction of structural motifs for predicting functional sites on protein structures.

Structural genomics initiatives are beginning to rapidly generate vast numbers of protein structures. For many of the structures, functions are not yet determined and high-throughput methods for determining function are necessary. Although there has been extensive work in function prediction at the sequence level, predicting function at the structure level may provide better sensitivity and predictive value. We describe a method to predict functional sites by automatically creating three dimensional structural motifs from amino acid sequence motifs. These structural motifs perform comparably well with manually generated structural motifs and perform better than sequence motifs. Automatically generated structural motifs can be used for structural-genomic scale function prediction on protein structures.

Amino Acid Motifs↗

Evaluation of proteome reference maps for cross-species identification of proteins by peptide mass fingerprinting.

We tested whether proteome reference maps established for one species can be used for cross-species protein identification by comparing two-dimensional protein gel patterns and protein identification data of two closely related bacterial strains and four plant species. First, proteome profiles of two strains of the fully sequenced bacterium Sinorhizobium meliloti were compared as an example of close relatedness, high reproducibility and sequence availability. Secondly, the proteome profiles of three legumes (Medicago truncatula, Melilotus alba and Trifolium subterraneum), and the nonlegume rice (Oryza sativa) were analysed to test cross-species similarities. In general, we found stronger similarities in gel patterns of the arrayed proteins between the two bacterial strains and between the plant species than could be expected from the sequence similarities. However, protein identity could not be concluded from their gel position, not even when comparing strains of the same species. Surprisingly, in the bacterial strains peptide mass fingerprinting was more reliable for species-specific protein identification than N-terminal sequencing. While peptide masses were found to be unreliable for cross-species protein identification, we present useful criteria to determine confident matching against species-specific expressed sequence tag databases. In conclusion, we present evidence that cautions the use of proteome reference maps and peptide mass fingerprinting for cross-species protein identification.

Electrophoresis, Gel, Two-Dimensional↗

New perspectives in the Escherichia coli proteome investigation.

Escherichia coli is a model organism for biochemical and biological studies as it is one of the best characterised prokaryote. Two-dimensional polyacrylamide gel electrophoresis, computer image analysis and different protein identification techniques gave rise, in 1995, to the Escherichia coli SWISS-2D PAGE database (http://www.expasy.ch/ch2d/). In the E. coli 3.5-10 SWISS-2D PAGE map, 40% of the E. coli proteome was displayed. The present study demonstrated that the use of narrow range pH gradients is able to potentially display up to a few copies of protein per E. coli cell. Moreover, the six new E. coli SWISS-2D PAGE maps (pH 4-5, 4.5-5.5, 5-6, 5.5-6.7, 6-9 and 6-11) presented here displayed altogether more than 70% of the entire E. coli proteome.

Databases, Protein↗

CHOMPER: a bioinformatic tool for rapid validation of tandem mass spectrometry search results associated with high-throughput proteomic strategies.

Current efforts aimed at developing high-throughput proteomics focus on increasing the speed of protein identification. Although improvements in sample separation, enrichment, automated handling, mass spectrometric analysis, as well as data reduction and database interrogation strategies have done much to increase the quality, quantity and efficiency of data collection, significant bottlenecks still exist. Various separation techniques have been coupled with tandem mass spectrometric (MS/MS) approaches to allow a quicker analysis of complex mixtures of proteins, especially where a high number of unambiguous protein identifications are the exception, rather than the rule. MS/MS is required to provide structural / amino acid sequence information on a peptide and thus allow protein identity to be inferred from individual peptides. Currently these spectra need to be manually validated because: (a) the potential of false positive matches i.e., protein not in database, and (b) observed fragmentation trends may not be incorporated into current MS/MS search algorithms. This validation represents a significant bottleneck associated with high-throughput proteomic strategies. We have developed CHOMPER, a software program which reduces the time required to both visualize and confirm MS/MS search results and generate post-analysis reports and protein summary tables. CHOMPER extracts the identification information from SEQUEST MS/MS search result files, reproduces both the peptide and protein identification summaries, provides a more interactive visualization of the MS/MS spectra and facilitates the direct submission of manually validated identifications to a database.

Algorithms↗

Structural proteome of human colostral fat globule membrane proteins.

Milk fat globule membrane (MFGM) contains proteins derived from the apical membrane of secreting epithelial cells of the mammary gland. Between 2-4% of total human milk protein content is associated with the fat globule fraction, as MFGM proteins. While MFGM proteins have very low classical nutritional value, they play important roles in various cell processes and defence mechanisms for the newborn. To date, fewer than 30 human MFGM proteins have been identified and characterized, either by immunological methods or by Edman sequencing and mass spectrometry. This study aimed to update the structural proteome of human colostral MFGM proteins and to create an annotated two-dimensional electrophoresis (2-DE) MFGM protein database available on-line. More than one hundred 2-DE spots derived from human colostral MFGM proteins were investigated by matrix-assisted laser desorption/ionization-time of flight mass spectrometry and proteins were identified by three different software packages available on the web (PeptIdent, MS-Fit and ProFound); uncertain identifications were solved by nanoelectrospray ionization-ion trap mass spectrometry using SEQUEST software.

Databases, Protein↗

Use of lysozyme as a standard for evaluating the effectiveness of a proteomics process.

Automated sequencing of unknowns in bottom-up proteomics makes the data produced susceptible to process control errors, which can be propagated into mistakes in analyte identification. Inclusion of an unintrusive internal standard, such as lysozyme, allows monitoring all phases of the proteomics process including sample preparation, enzymatic digestion, HPLC, mass spectrometry, and database searching. By using this internal standard, digestion issues including rearrangements, semi-tryptic fragments, and modifications were monitored. In addition, control of the HPLC process including column performance was achieved. The use of the lysozyme standard allowed easy optimization of mass spectral conditions including data dependent and collision induced dissociation settings. The use of this internal standard in a study of differential protein expression in rat serum samples is presented.

Animals↗

Search and discovery strategies for biotechnology: the paradigm shift.

Profound changes are occurring in the strategies that biotechnology-based industries are deploying in the search for exploitable biology and to discover new products and develop new or improved processes. The advances that have been made in the past decade in areas such as combinatorial chemistry, combinatorial biosynthesis, metabolic pathway engineering, gene shuffling, and directed evolution of proteins have caused some companies to consider withdrawing from natural product screening. In this review we examine the paradigm shift from traditional biology to bioinformatics that is revolutionizing exploitable biology. We conclude that the reinvigorated means of detecting novel organisms, novel chemical structures, and novel biocatalytic activities will ensure that natural products will continue to be a primary resource for biotechnology. The paradigm shift has been driven by a convergence of complementary technologies, exemplified by DNA sequencing and amplification, genome sequencing and annotation, proteome analysis, and phenotypic inventorying, resulting in the establishment of huge databases that can be mined in order to generate useful knowledge such as the identity and characterization of organisms and the identity of biotechnology targets. Concurrently there have been major advances in understanding the extent of microbial diversity, how uncultured organisms might be grown, and how expression of the metabolic potential of microorganisms can be maximized. The integration of information from complementary databases presents a significant challenge. Such integration should facilitate answers to complex questions involving sequence, biochemical, physiological, taxonomic, and ecological information of the sort posed in exploitable biology. The paradigm shift which we discuss is not absolute in the sense that it will replace established microbiology; rather, it reinforces our view that innovative microbiology is essential for releasing the potential of microbial diversity for biotechnology penetration throughout industry. Various of these issues are considered with reference to deep-sea microbiology and biotechnology.

Biotechnology↗

Genome annotation of Anopheles gambiae using mass spectrometry-derived data.

BACKGROUND: A large number of animal and plant genomes have been completely sequenced over the last decade and are now publicly available. Although genomes can be rapidly sequenced, identifying protein-coding genes still remains a problematic task. Availability of protein sequence data allows direct confirmation of protein-coding genes. Mass spectrometry has recently emerged as a powerful tool for proteomic studies. Protein identification using mass spectrometry is usually carried out by searching against databases of known proteins or transcripts. This approach generally does not allow identification of proteins that have not yet been predicted or whose transcripts have not been identified. RESULTS: We searched 3,967 mass spectra from 16 LC-MS/MS runs of Anopheles gambiae salivary gland homogenates against the Anopheles gambiae genome database. This allowed us to validate 23 known transcripts and 50 novel transcripts. In addition, a novel gene was identified on the basis of peptides that matched a genomic region where no gene was known and no transcript had been predicted. The amino termini of proteins encoded by two predicted transcripts were confirmed based on N-terminally acetylated peptides sequenced by tandem mass spectrometry. Finally, six sequence polymorphisms could be annotated based on experimentally obtained peptide sequences. CONCLUSION: The peptide sequences from this study were mapped onto the genomic sequence using the distributed annotation system available at Ensembl and can be visualized in the context of all other existing annotations. The strategy described in this paper can be used to correct and confirm genome annotations and permit discovery of novel proteins in a high-throughput manner by mass spectrometry.

Animals↗

The amniotic fluid cell proteome.

Proteomic analysis of amniotic fluid cells may lead to the discovery of novel markers for embryonic abnormalities. A two-dimensional database for proteins of normal human amniotic fluid cells was constructed. The amniotic fluid cell extract was analyzed by two-dimensional gel electrophoresis and the proteins were identified by matrix-assisted laser desorption ionisation-time of flight-mass spectrometry. The database comprises 432 different gene products, which are in the majority enzymes, structural proteins, heat shock proteins, and proteins related to signal transduction. The obtained data show that the amniotic fluid population maybe either heterogeneous, originating from different fetal compartments and embryo tissues or is still pluripotent. Many proteins which are known to belong to certain cell types were found in the amnion cell fluid. This indicates that some types of fetal cells are already differentiated at the time of amniocentesis (about the 16(th) week of gestation). Moreover, the finding of proteins highly expressed in embryonic stem cells suggests that amniotic fluid could be used as a cell pool for transplantation therapy.

Amniotic Fluid↗

Systematic comparison of a two-dimensional ion trap and a three-dimensional ion trap mass spectrometer in proteomics.

The utility and advantages of the recently introduced two-dimensional quadrupole ion trap mass spectrometer in proteomics over the traditional three-dimensional ion trap mass spectrometer have not been systematically characterized. Here we rigorously compared the performance of these two platforms by using over 100,000 tandem mass spectra acquired with identical complex peptide mixtures and acquisition parameters. Specifically we compared four factors that are critical for a successful proteomic study: 1) the number of proteins identified, 2) sequence coverage or the number of peptides identified for every protein, 3) the data base matching SEQUEST X(corr) and S(p) score, and 4) the quality of the fragment ion series of peptides. We found a 4-6-fold increase in the number of peptides and proteins identified on the two-dimensional ion trap mass spectrometer as a direct result of improvement in all the other parameters examined. Interestingly more than 70% of the doubly and triply charged peptides, but not the singly charged peptides, showed better quality of fragmentation spectra on the two-dimensional ion trap. These results highlight specific advantages of the two-dimensional ion trap over the conventional three-dimensional ion traps for protein identification in proteomic experiments.

Computational Biology↗