Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

The International Rice Information System. A platform for meta-analysis of rice crop data.

Ambiguous germplasm identification; difficulty in tracing pedigree information; and lack of integration between genetic resources, characterization, breeding, evaluation, and utilization data are constraints in developing knowledge-intensive crop improvement programs. To address these constraints, the International Crop Information System (www.icis.cgiar.org), a database system for the management and integration of global information on genetic resources and crop improvement for any crop, was developed by genetic resource specialists, crop scientists, and information technicians associated with the Consultative Group for International Agricultural Research and collaborative partners. The International Rice Information System (www.iris.irri.org) is the rice (Oryza species) implementation of the International Crop Information System. New components are now being added to the International Rice Information System to handle the diversity of rice functional genomics data including genomic sequence data, molecular genetic data, expression data, and proteomic information. Users access information in the database through stand-alone programs and Web interfaces, which offer specialized applications and customized views to researchers with different interests.

Breeding↗

A comprehensive dictionary of protein accession codes for complete protein accession identifier alias resolving.

In mass spectrometry-based proteomics, protein identification results usually consist of peptide sequences and database-dependent accession identifiers of the matching proteins. Often certain annotations are only available in particular databases that in turn must be queried by a certain identifier. In order to simplify and unify the tracing of identified proteins back to their original annotation information, a system capable of set-oriented mapping the different accession identifiers of proteins derived from multiple sequence database sources has been developed. This allows unification of the access to protein information and tracing to other online resources providing additional information as well as resolving cross-references of protein identifications. The interface of seqDB is available via http://www.protein-ms.de following the link to seqDB.

Database Management Systems↗

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine↗

Informatics solutions for high-throughput proteomics.

The success of mass-spectrometry-based proteomics as a method for analyzing proteins in biological samples is accompanied by challenges owning to demands for increased throughput. These challenges arise from the vast volume of data generated by proteomics experiments combined with the heterogeneity in data formats, processing methods, software tools and databases that are involved in the translation of spectral data into relevant and actionable information for scientists. Informatics aims to provide answers to these challenges by transferring existing solutions from information management to proteomics and/or by generating novel computational methods for automation of proteomics data processing.

Databases, Protein↗

Multiple parameter cross-species protein identification using MultiIdent--a world-wide web accessible tool.

Recent increases in the number of genome sequencing projects means that the amount of protein sequence in databases is increasing at an astonishing pace. In proteome studies, this is facilitating the identification of proteins from molecularly well-defined organisms. However, in studies of proteins from the majority of organisms, proteins must be identified by comparing analytical data to sequences in databases from other species. This process is known as cross-species protein identification. Here we present a new program, MultiIdent, which uses multiple protein parameters such as amino acid composition, peptide masses, sequence tags, estimated protein pI and mass, to achieve cross-species protein identification. The program is structured so that protein amino acid composition, which is highly conserved across species boundaries, first generates a set of candidate proteins. These proteins are then queried with other protein parameters such as sequence tags and peptide masses. A final list of database entries which considers all analytical parameters is presented, ranked by an integrated score. We illustrate the power of the approach with the identification of a set of standard proteins, and the identification of proteins from dog heart separated by two-dimensional gel electrophoresis. The MultiIdent program is available on the world-wide web at: http://www.expasy.ch/sprot/multiident.h tml.

Amino Acid Sequence↗

Proteome analysis of primary neurons and astrocytes from rat cerebellum.

Neurons and astrocytes are predominant cell types in brain and have distinguished morphological and functional features. Although several proteomics studies were carried out on the brain, work on individual brain cells is limited. Generating individual proteomes of neurons and astrocytes, however, is mandatory to assign protein expression to cell types rather than to tissues. We aimed to provide maps of rat primary neurons and astrocytes using two-dimensional gel electrophoresis with subsequent in-gel digestion, followed by MALDI-TOF/TOF. 428 protein spots corresponding to 226 individual proteins in neurons and 406 protein spots representing 228 proteins in astrocytes were unambiguously identified. Proteome data include proteins from several cascades differentially expressed in neurons and astrocytes, and specific expressional patterns of antioxidant, signaling, chaperone, cytoskeleton, nucleic acid binding, proteasomal, and metabolic proteins are demonstrated. We herein present a reference database of primary rat primary neuron and astrocyte proteomes and provide an analytical tool for these structures. The concomitant expressional patterns of several protein classes are given and potential neuronal and astrocytic marker candidates are presented.

Animals↗

Simple modification of a protein database for mass spectral identification of N-linked glycopeptides.

We describe an algorithm which modifies a protein database such that during a database search deamidation is limited to asparagines strictly contained within the N-glycosylation consensus sequence. The modified database was evaluated using a dataset created from the shotgun proteomic analysis of N-linked glycopeptides from human blood serum. We demonstrate that the application of the modified database eliminates incorrect glycopeptide assignments, reduces the peptide false-discovery rate, and eliminates the need for manual validation of glycopeptide identifications.

Algorithms↗

Analysis of the proteome in the human pituitary.

The pituitary is the master endocrine gland responsible for the regulation of various physiologic and metabolic processes. Proteomics offers an efficient means for a comprehensive analysis of pituitary protein expression. This paper reports on the application of proteomics for the mapping of major proteins in a normal (control) pituitary. Pituitary proteins were separated by two-dimensional gel electrophoresis with immobilized pH 3-10 gradient strips. Major protein spots that were visualized in the two-dimensional gel by silver staining were excised, and the proteins in these spots were digested with trypsin. The tryptic digests were analyzed by mass spectrometry, and the mass spectrometric data were used to identify the proteins through searches of the SWISS-PROT or NCBInr protein sequence databases. The majority of the proteins were identified on the basis of peptide mass fingerprinting data obtained by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry. Several proteins were also characterized based on product-ion spectra measured by post-source decay analysis and/or liquid chromatography-electrospray-quadrupole ion trap mass spectrometry. To date, 62 prominent protein spots, corresponding to 38 different proteins, were identified. The identified proteins include important pituitary hormones, structural proteins, enzymes, and other proteins. The protein identification data were used to establish a two-dimensional reference database of the human pituitary, which can be accessed over the Internet (http://www.utmem.edu/proteomics). This database will serve as a tool for further proteomics studies of pituitary protein expression in health and disease.

Computational Biology↗

Arabidopsis thaliana proteomics: from proteome to genome.

Proteomics has become an important approach for investigating cellular processes and network functions. Significant improvements have been made during the last few years in technologies for high-throughput proteomics, both at the level of data analysis software and mass spectrometry hardware. As proteomics technologies advance and become more widely accessible, efforts of cataloguing and quantifying full proteomes are underway to complement other genomics approaches, such as RNA and metabolite profiling. Of particular interest is the application of proteome data to improve genome annotation and to include information on post-translational protein modifications with the annotation of the corresponding gene. This type of analysis requires a paradigm shift because amino acid sequences must be assigned to peptides without relying on existing protein databases. In this review, advances and current limitations of full proteome analysis are briefly highlighted using the model plant Arabidopsis thaliana as an example. Strategies to identify peptides are also discussed on the basis of MS/MS data in a protein database-independent approach.

Arabidopsis↗

BioBuilder as a database development and functional annotation platform for proteins.

BACKGROUND: The explosion in biological information creates the need for databases that are easy to develop, easy to maintain and can be easily manipulated by annotators who are most likely to be biologists. However, deployment of scalable and extensible databases is not an easy task and generally requires substantial expertise in database development. RESULTS: BioBuilder is a Zope-based software tool that was developed to facilitate intuitive creation of protein databases. Protein data can be entered and annotated through web forms along with the flexibility to add customized annotation features to protein entries. A built-in review system permits a global team of scientists to coordinate their annotation efforts. We have already used BioBuilder to develop Human Protein Reference Database http://www.hprd.org, a comprehensive annotated repository of the human proteome. The data can be exported in the extensible markup language (XML) format, which is rapidly becoming as the standard format for data exchange. CONCLUSIONS: As the proteomic data for several organisms begins to accumulate, BioBuilder will prove to be an invaluable platform for functional annotation and development of customizable protein centric databases. BioBuilder is open source and is available under the terms of LGPL.

Computational Biology↗

Proteomic analysis of native metabotropic glutamate receptor 5 protein complexes reveals novel molecular constituents.

We used a proteomic approach to identify novel proteins that may regulate metabotropic glutamate receptor 5 (mGluR5) responses by direct or indirect protein interactions. This approach does not rely on the heterologous expression of proteins and offers the advantage of identifying protein interactions in a native environment. The mGluR5 protein was immunoprecipitated from rat brain lysates; co-immunoprecipitating proteins were analyzed by mass spectrometry and identified peptides were matched to protein databases to determine the correlating parent proteins. This proteomic approach revealed the interaction of mGluR5 with known regulatory proteins, as well as novel proteins that reflect previously unidentified molecular constituents of the mGluR5-signaling complex. Immunoblot analysis confirmed the interaction of high confidence proteins, such as phosphofurin acidic cluster sorting protein 1, microtubule-associated protein 2a and dynamin 1, as mGluR5-interacting proteins. These studies show that a proteomic approach can be used to identify candidate interacting proteins. This approach may be particularly useful for neurobiology applications where distinct protein interactions within a signaling complex can dramatically alter the outcome of the response to neurotransmitter release, or the disruption of normal protein interactions can lead to severe neurological and psychiatric disorders.

Algorithms↗

The CyberCell Database (CCDB): a comprehensive, self-updating, relational database to coordinate and facilitate in silico modeling of Escherichia coli.

The CyberCell Database (CCDB: http://redpoll. pharmacy.ualberta.ca/CCDB) is a comprehensive, web-accessible database designed to support and coordinate international efforts in modeling an Escherichia coli cell on a computer. The CCDB brings together both observed and derived quantitative data from numerous independent sources covering many aspects of the genomic, proteomic and metabolomic character of E.coli (strain K12). The database is self-updating but also supports 'community' annotation, and provides an extensive array of viewing, querying and search options including a powerful, easy-to-use relational data extraction system.

Computational Biology↗

Comparative proteomics of apoptosis initiation induced by 5-fluorouracil in human gastric cancer.

5-Fluorouracil is the first choice chemotherapeutic drug for patients with gastric cancer, but the mechanism that 5-fluorouracil plays the anti-tumor role remains unclear. The aim of this study was to clarify correlated [corrected] proteins induced by 5-fluorouracil in the apoptosis-initiation of human gastric cancer (MGC-803) cells. The time point of apoptosis-initiation induced by 5-fluorouracil in MGC-803 cells was determinated using 5-fluorouracil-withdrawal. Two-dimensional electrophoreses (2-DE) were employed to compare the differentials of protein expressions of the MGC-803 cells at the apoptosis-initiation phase and those of the MGC-803 cells untreated with 5-fluorouracil. The differential proteins included 14 upregulated proteins and 8 downregulated proteins. They indicated a more-than-doubled alteration. These proteins were digested in gels by trypsin and the mass of generated peptides were measured by matrix assisted laser desorption ionization time of flight mass spectrometry (MALDI-TOF-MS). The data obtained from peptide mass fingerprinting (PMF) were searched out using the internet available database mascot (http://www.matrixscience.com). The results showed that proteomics analyses have evidenced that many kinds of proteins are involved in the apoptosis initiation of human gastric cancer MGC-803 cells. These proteins are related to metabolism, oxidation, cytoskeleton and signal transduction and other aspects of cells. In conclusion, the experiment model of apoptosis-initiation of human gastric cancer MGC-803 cells induced by 5-fluorouracil based on proteomic analysis has been established, giving an impetus to researches of the mechanism of apoptosis in human gastric cancer, and laying a foundation for the selection of potential drug precursors specific for inducing apoptosis-initiation in human gastric cancer.

Adenocarcinoma↗

Using standard positions and image fusion to create proteome maps from collections of two-dimensional gel electrophoresis images.

Databases for two-dimensional protein gels pose new challenges in extracting meaningful information from large numbers of experiments. In order to create expression profiles, positions of corresponding protein spots across all gel images have to be established. In larger gel sets errors may accumulate rapidly during this spot matching process, effectively limiting the number of samples available for data mining. Here we present a novel approach for organizing spot data based on the concept of a standard position for a protein species. Standard positions are meaningful average positions that are determined using all occurrences of a protein species. They can be extended to spots that are not annotated via interpolation. The standard position of a spot can serve as a unifying index across all gels in a database, thus allowing creation and analysis of expression profiles that span the whole collection. The standard position gives a much more accurate estimation of a spot's position on a gel than can be obtained using theoretical isoelectric point and molecular weight. Positional indexing is a complement to a priori identifications (e.g. by mass spectrometry or Edman degradation). Moreover it can be used in advance to select spots that are worth identifying because they show relevant expression profiles. Furthermore, we show how to combine all spots that occur on any of the gels into one synthetic but nevertheless realistic-looking image. This composite image is produced such that all spots have their standard positions. It can serve as a proteome reference map for an organism. As an application, we have computed a reference map from 23 gel images of Bacillus subtilis, using an enhanced prerelease version of the gel analysis software Delta2D (DECODON, Greifswald, Germany).

Bacterial Proteins↗

Mass spectrometric identification of proteins and characterization of their post-translational modifications in proteome analysis.

High-throughput DNA sequencing has resulted in increasing input in protein sequence databases. Today more than 20 genomes have been sequenced and many more will be completed in the near future, including the largest of them all, the human genome. Presently, sequence databases contain entries for more than 425.000 protein sequences. However, the cellular functions are determined by the set of proteins expressed in the cell--the proteome. Two-dimensional gel electrophoresis, mass spectrometry and bioinformatics have become important tools in correlating the proteome with the genome. The current dominant strategies for identification of proteins from gels based on peptide mass spectrometric fingerprinting and partial sequencing by mass spectrometry are described. After identification of the proteins the next challenge in proteome analysis is characterization of their post-translational modifications. The general problems associated with characterization of these directly from gel separated proteins are described and the current state of art for the determination of phosphorylation, glycosylation and proteolytic processing is illustrated.

Computational Biology↗

Domain graph of Arabidopsis proteome by comparative analysis.

The domain graph of domains and domain combinations of Arabidopsis thaliana is established based on pfam 14.0 database and analyzed via comparison with 10 eukaryotic, 30 bacterial, and 16 archaeal proteomes. The comparative analysis of the domain graphs provides a useful platform for revealing global insights on the evolution of plant kingdom. More importantly, it is a powerful tool for searching not only the possible new function of both plant-specific and nonspecific domains via specific domain combinations in Arabidopsis thaliana but also the functional role of unknown domains. As an example, we present the functional link between ubiquitin and Myb_DNA-binding domains via Bromodomain as the plant specific evidence for the association between transcription and ubiquitin. We further show that PentatricoPeptide Repeats (PPR) proteins have plant-specific links with a wide variety of domains responsible for RNA binding/metabolism, modulation of protein-protein interactions, ubiquitin-conjugation, cell growth/maintenance, catalysis, and others. This further supports the recently proposed association of PPR proteins with specific RNA transcripts and defined effector proteins. Moreover, the domain graph built from tissue-specific genes is frequently associated with DNA binding domains, suggesting that the differentiation of tissue cell types is contributed mostly by tissue-specific transcriptional process. DOGMA (DOmain Graph via coMparitive analysis for Arabidopsis thaliana) is available on-line with a variety of search tools at http://theory.med.buffalo.edu/DOGMA. The database, which allows user-specified search for plant specific domains and their combinations, will be useful as an additional tool for annotation of the proteins that play specific roles in plants and other organisms.

Arabidopsis↗

A model of random mass-matching and its use for automated significance testing in mass spectrometric proteome analysis.

A rapid and accurate method for testing the significance of protein identities determined by mass spectrometric analysis of protein digests and genome database searching is presented. The method is based on direct computation using a statistical model of the random matching of measured and theoretical proteolytic peptide masses. Protein identification algorithms typically rank the proteins of a genome database according to a score based on the number of matches between the masses obtained by mass spectrometry analysis and the theoretical proteolytic peptide masses of a database protein. The random matching of experimental and theoretical masses can cause false results. A result is significant only if the score characterizing the result deviates significantly from the score expected from a false result. A distribution of the score (number of matches) for random (false) results is computed directly from our model of the random matching, which allows significance testing under any experimental and database search constraints. In order to mimic protein identification data quality in large-scale proteome projects, low-to-high quality proteolytic peptide mass data were generated in silico and subsequently submitted to a database search program designed to include significance testing based on direct computation. This simulation procedure demonstrates the usefulness of direct significance testing for automatically screening for samples that must be subjected to peptide sequence analysis by e.g. tandem mass spectrometry in order to determine the protein identity.

Algorithms↗

Identification of an evolutionary conserved SURF-6 domain in a family of nucleolar proteins extending from human to yeast.

The mammalian SURF-6 protein is localized in the nucleolus, yet its function remains elusive in the recently characterized nucleolar proteome. We discovered by searching the Protein families database that a unique evolutionary conserved SURF-6 domain is present in the carboxy-terminal of a novel family of eukaryotic proteins extending from human to yeast. By using the enhanced green fluorescent protein as a fusion protein marker in mammalian cells, we show that proteins from distantly related taxonomic groups containing the SURF-6 domain are localized in the nucleolus. Deletion sequence analysis shows that multiple regions of the SURF-6 protein are capable of nucleolar targeting independently of the evolutionary conserved domain. We identified that the Saccharomyces cerevisiae member of the SURF-6 family, named rrp14 or ykl082c, has been categorized in yeast databases to interact with proteins involved in ribosomal biogenesis and cell polarity. These results classify SURF-6 as a new family of nucleolar proteins in the eukaryotic kingdom and point out that SURF-6 has a distinct domain within the known nucleolar proteome that may mediate complex protein-protein interactions for analogous processes between yeast and mammalian cells.

Amino Acid Sequence↗