Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Sodium dodecyl sulfate versus acid-labile surfactant gel electrophoresis: comparative proteomic studies on rat retina and mouse brain.

A long-chain derivative of 1,3-dioxolane sodium propyloxy sulfate, with similar denaturing and electrophoretic properties as SDS, and facilitated protein identification following polyacrylamide gel electrophoresis (PAGE) for Coomassie-stained protein bands, has been tested. Comparative acid-labile surfactant/sodium dodecyl sulfate two-dimensional (ALS/SDS 2-D)-PAGE experiments of lower abundant proteins from the proteomes of regenerating rat retina and mouse brain show that peptide recovery for mass spectrometry (MS) mapping is significantly enhanced using ALS leading to more successful database searches. ALS may influence some procedures in proteomic analysis such as the determination of protein content and methods need to be adjusted to that effect. The promising results of the use of ALS in bioanalytics call for detailed physicochemical investigations of surfactant properties.

Acids↗

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics↗

PlasmoDB: exploring genomics and post-genomics data of the malaria parasite, Plasmodium falciparum.

The recent completion of the genome sequence of Plasmodium falciparum 3D7 provides the foundation for genome-wide analysis of the parasite. In addition to DNA and gene sequence data, postgenomic methods including microarray-based transcript profiling and high-throughput proteomics are now accessible to Plasmodium researchers. The Plasmodium Genome database ( ) was developed to provide rapid and convenient access to the terabytes of genomic-scale data now being generated around the world. All data are available in a relational framework, permitting convenient downloading, browsing, and analysis. Combinatorial use of data analysis tools enables powerful data mining queries, such as combining gene and protein expression data to monitor changes through various life-cycle stages. Functional predictions can be used to explore potential targets for antimalarial drug development. This report outlines the use of PlasmoDB to examine redox-active functions in Plasmodium.

Animals↗

Impact of ion trap tandem mass spectra variability on the identification of peptides.

Peptide identification based on tandem mass spectrometry and database searching algorithms has become one of the central technologies in proteomics. At the heart of this technology is the ability to reproducibly acquire high-quality tandem mass spectra for database interrogation. The variability in tandem mass spectra generation is often assumed to be minimal, and peptide identifications are typically based on a single tandem mass spectrum. In this paper, we characterize the variance of scores derived from replicate tandem mass spectra using several database search algorithms and demonstrate the effects of spectral variability on the correct identification of peptides. We show that the variance associated with the collection of tandem mass spectra can be substantial leading to sizable errors in search algorithm scores ( approximately 5-25% RSD) and ultimately incorrect assignments. Processing strategies are discussed to minimize the impact of tandem mass spectra variability on peptide identification.

Algorithms↗

ProteomeWeb: a web-based interface for the display and interrogation of proteomes.

The analysis of proteomes, i.e., the proteins expressed by biological organisms under a given set of conditions at a given time, requires separating complex protein mixtures into discrete protein components, measuring their relative abundances, and identifying the individual protein components. Many types of data are generated during the course of proteome analysis, including graphic images of the protein profiles, flat files containing numeric data, spreadsheets for assimilating numeric data, and relational database tables for integrating data from multiple experiments. As part of a project to describe the proteomes of microbes of interest to the U.S. Department of Energy, a World-Wide Web-based interface has been developed for the display of protein profiles generated by two-dimensional gel electrophoresis. The web interface is capable of obtaining protein identifications on the fly, interrogating the quantitative data in the context of available genome sequence information, and relating the proteome data to existing metabolic pathway databases. Analysis of protein expression profiles is expedited, providing the capability to efficiently determine the gene locations for proteins modulated in abundance in response to different growth conditions and to locate the positions of the proteins within specific metabolic pathways. The proteome of the archaeon Methanococcus jannaschii, a microbe for which the complete genome sequence is available, is used to demonstrate the capabilities of this evolving web interface (http://proteomeweb.anl.gov).

Amino Acid Sequence↗

Interpretation of shotgun proteomic data: the protein inference problem.

The shotgun proteomic strategy based on digesting proteins into peptides and sequencing them using tandem mass spectrometry and automated database searching has become the method of choice for identifying proteins in most large scale studies. However, the peptide-centric nature of shotgun proteomics complicates the analysis and biological interpretation of the data especially in the case of higher eukaryote organisms. The same peptide sequence can be present in multiple different proteins or protein isoforms. Such shared peptides therefore can lead to ambiguities in determining the identities of sample proteins. In this article we illustrate the difficulties of interpreting shotgun proteomic data and discuss the need for common nomenclature and transparent informatic approaches. We also discuss related issues such as the state of protein sequence databases and their role in shotgun proteomic analysis, interpretation of relative peptide quantification data in the presence of multiple protein isoforms, the integration of proteomic and transcriptional data, and the development of a computational infrastructure for the integration of multiple diverse datasets.

Amino Acid Sequence↗

Novel gene and gene model detection using a whole genome open reading frame analysis in proteomics.

BACKGROUND: Defining the location of genes and the precise nature of gene products remains a fundamental challenge in genome annotation. Interrogating tandem mass spectrometry data using genomic sequence provides an unbiased method to identify novel translation products. A six-frame translation of the entire human genome was used as the query database to search for novel blood proteins in the data from the Human Proteome Organization Plasma Proteome Project. Because this target database is orders of magnitude larger than the databases traditionally employed in tandem mass spectra analysis, careful attention to significance testing is required. Confidence of identification is assessed using our previously described Poisson statistic, which estimates the significance of multi-peptide identifications incorporating the length of the matching sequence, number of spectra searched and size of the target sequence database. RESULTS: Applying a false discovery rate threshold of 0.05, we identified 282 significant open reading frames, each containing two or more peptide matches. There were 627 novel peptides associated with these open reading frames that mapped to a unique genomic coordinate placed within the start/stop points of previously annotated genes. These peptides matched 1,110 distinct tandem MS spectra. Peptides fell into four categories based upon where their genomic coordinates placed them relative to annotated exons within the parent gene. CONCLUSION: This work provides evidence for novel alternative splice variants in many previously annotated genes. These findings suggest that annotation of the genome is not yet complete and that proteomics has the potential to further add to our understanding of gene structures.

Alternative Splicing↗

Transfusion medicine in the era of genomics and proteomics.

Viewing recent trends in transfusion medicine (TM), the authors make predictions about possible future developments within this specialty including greater cost-effectiveness and blood safety resulting from increased automation; techniques in genetics replacing serological typing in many standard assays; and TM service playing a major R&D role together with clinical services in the emerging cell-based therapeutics. To achieve this, the TM laboratory of the future will need to have available extensive skills in immunogenetics and database expertise; emerging techniques in genomics and proteomics will need to be integrated with classic immunohematology approaches; and collaborative networks of TM laboratories will need to raise their profiles as a competent partner in the ongoing clinical biotechnology revolution. Blood product safety is profiled to highlight some of these developments. Until recently, avoiding pathogen transmission has focused primarily on excluding at-risk donors and testing donor blood for pathogen markers. Newer trends in pathogen-inactivation procedures could alter the protein composition of the blood product, potentially causing unintended immune reactions that could outweigh their benefits in further reducing a very low current risk of pathogen transmission. By combining proteomics and immunohematology, those manufacturing processes least likely to generate posttranslational protein modifications will need to be identified.

Automation↗

Mapping the platelet proteome: a report of the ISTH Platelet Physiology Subcommittee.

Proteomic technology has the potential to transform the way we analyze platelet biology, through the determination of platelet protein composition and its modification upon stimulation and with disease. We are a considerable way from achieving these goals, however, because of significant limitations in current methodology. It is therefore important to consider the extent to which these aims can be met and the way that proteomic data should be presented and used. These issues are discussed in the present paper by the Platelet Physiology Subcommittee of the ISTH Scientific Standardisation Committee (SSC). It is recommended that proteomic information be combined with data from other experimental approaches to establish a database on protein expression and function in platelets.

Blood Platelets↗

The path to enlightenment: making sense of genomic and proteomic information.

Whereas genomics describes the study of genome, mainly represented by its gene expression on the DNA or RNA level, the term proteomics denotes the study of the proteome, which is the protein complement encoded by the genome. In recent years, the number of proteomic experiments increased tremendously. While all fields of proteomics have made major technological advances, the biggest step was seen in bioinformatics. Biological information management relies on sequence and structure databases and powerful software tools to translate experimental results into meaningful biological hypotheses and answers. In this resource article, I provide a collection of databases and software available on the Internet that are useful to interpret genomic and proteomic data. The article is a toolbox for researchers who have genomic or proteomic datasets and need to put their findings into a biological context.

Computational Biology↗

HUPO Publications Committee Meeting: 21 April 2006, San Francisco, CA, USA.

This meeting was convened with the aim of bringing together representatives from scientific journals, granting authorities, software and instrumentation manufacturers, data producers and database providers to discuss the implementation and adoption of the HUPO-PSI data standards and how these can be best used to support the publication and dissemination of proteomics data. The current status of data formats and reporting requirements was reviewed and the attendees agreed that the use of data standards was essential as the field of proteomics grows and matures.

Databases, Protein↗

Improved classification of mass spectrometry database search results using newer machine learning approaches.

Manual analysis of mass spectrometry data is a current bottleneck in high throughput proteomics. In particular, the need to manually validate the results of mass spectrometry database searching algorithms can be prohibitively time-consuming. Development of software tools that attempt to quantify the confidence in the assignment of a protein or peptide identity to a mass spectrum is an area of active interest. We sought to extend work in this area by investigating the potential of recent machine learning algorithms to improve the accuracy of these approaches and as a flexible framework for accommodating new data features. Specifically we demonstrated the ability of boosting and random forest approaches to improve the discrimination of true hits from false positive identifications in the results of mass spectrometry database search engines compared with thresholding and other machine learning approaches. We accommodated additional attributes obtainable from database search results, including a factor addressing proton mobility. Performance was evaluated using publically available electrospray data and a new collection of MALDI data generated from purified human reference proteins.

Amino Acid Sequence↗

Web and database software for identification of intact proteins using "top down" mass spectrometry.

For the identification and characterization of proteins harboring posttranslational modifications (PTMs), a "top down" strategy using mass spectrometry has been forwarded recently but languishes without tailored software widely available. We describe a Web-based software and database suite called ProSight PTM constructed for large-scale proteome projects involving direct fragmentation of intact protein ions. Four main components of ProSight PTM are a database retrieval algorithm (Retriever), MySQL protein databases, a file/data manager, and a project tracker. Retriever performs probability-based identifications from absolute fragment ion masses, automatically compiled sequence tags, or a combination of the two, with graphical rendering and browsing of the results. The database structure allows known and putative protein forms to be searched, with prior or predicted PTM knowledge used during each search. Initial functionality is illustrated with a 36-kDa yeast protein identified from a processed cell extract after automated data acquisition using a quadrupole-FT hybrid mass spectrometer. A +142-Da delta(m) on glyceraldehyde-3-phosphate dehydrogenase was automatically localized between Asp90 and Asp192, consistent with its two cystine residues (149 and 153) alkylated by acrylamide (+71 Da each) during the gel-based sample preparation. ProSight PTM is the first search engine and Web environment for identification of intact proteins (https://prosightptm.scs.uiuc.edu/).

Amino Acid Sequence↗

Scale-free behavior in protein domain networks.

Several technical, social, and biological networks were recently found to demonstrate scale-free and small-world behavior instead of random graph characteristics. In this work, the topology of protein domain networks generated with data from the ProDom, Pfam, and Prosite domain databases was studied. It was found that these networks exhibited small-world and scale-free topologies with a high degree of local clustering accompanied by a few long-distance connections. Moreover, these observations apply not only to the complete databases, but also to the domain distributions in proteomes of different organisms. The extent of connectivity among domains reflects the evolutionary complexity of the organisms considered.

Amino Acid Sequence↗

Proteomics: quantitative and physical mapping of cellular proteins.

Genome sequencing provides a wealth of information on predicted gene products (mostly proteins), but the majority of these have no known function. Two-dimensional gel electrophoresis and mass spectrometry have, coupled with searches in protein and EST databases, transformed the protein-identification process. The proteome is the expressed protein complement of a genome and proteomics is functional genomics at the protein level. Proteomics can be divided into expression proteomics, the study of global changes in protein expression, and cell-map proteomics, the systematic study of protein-protein interactions through the isolation of protein complexes.

Biotechnology↗

Characterization of the human salivary proteome by capillary isoelectric focusing/nanoreversed-phase liquid chromatography coupled with ESI-tandem MS.

Saliva is a readily available body fluid with great diagnostic potential. The foundation for saliva-based diagnostics, however, is the development of a complete catalog of secreted and "leaked" proteins detectable in saliva. By employing a capillary isoelectric focusing-based multidimensional separation platform coupled with electrospray ionization tandem mass spectrometry (MS), a total of 5338 distinct peptides were sequenced, leading to the identification of 1381 distinct proteins. A search of bacterial protein sequences also identified many peptides unique to several organisms and unique to the NCBI nonredundant database. To the best of our knowledge, this proteome study represents the largest catalog of proteins measured from a single saliva sample to date. Data analysis was performed on individual MS/MS spectra using the highly specific peptide identification algorithm, OMSSA. Searches were conducted against a decoyed SwissProt human database to control the false-positive rate at 1%. Furthermore, the well-curated SwissProt sequences represent perhaps the least redundant human protein sequence database (12,484 records versus the 50,009 records found in the International Protein Index human database), therefore minimizing multiple protein inferences from single peptides. This combined bioanalytical and bioinformatic approach has established a solid foundation for building up the human salivary proteome for the realization of the diagnostic potential of saliva.

Chromatography, Liquid↗

Assembling an arsenal: origin and evolution of the snake venom proteome inferred from phylogenetic analysis of toxin sequences.

We analyzed the origin and evolution of snake venom toxin families represented in both viperid and elapid snakes by means of phylogenetic analysis of the amino acid sequences of the toxins and related nonvenom proteins. Out of eight toxin families analyzed, five provided clear evidence of recruitment into the snake venom proteome before the diversification of the advanced snakes (Kunitz-type protease inhibitors, CRISP toxins, galactose-binding lectins, M12B peptidases, nerve growth factor toxins), and one was equivocal (cystatin toxins). In two others (phospholipase A(2) and natriuretic toxins), the nonmonophyly of venom toxins demonstrates that presence of these proteins in elapids and viperids results from independent recruitment events. The ANP/BNP natriuretic toxins are likely to be basal, whereas the CNP/BPP toxins are Viperidae only. Similarly, the lectins were recruited twice. In contrast to the basal recruitment of the galactose-binding lectins, the C-type lectins were shown to be Viperidae only, with the alpha-chains and beta-chains resulting from an early duplication event. These results provide strong additional evidence that venom evolved once, at the base of the advanced snake radiation, rather than multiple times in different lineages, with these toxins also present in the venoms of the "colubrid" snake families. Moreover, they provide a first insight into the composition of the earliest ophidian venoms and point the way toward a research program that could elucidate the functional context of the evolution of the snake venom proteome.

Animals↗