Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

IntNetDB v1.0: an integrated protein-protein interaction network database generated by a probabilistic model.

BACKGROUND: Although protein-protein interaction (PPI) networks have been explored by various experimental methods, the maps so built are still limited in coverage and accuracy. To further expand the PPI network and to extract more accurate information from existing maps, studies have been carried out to integrate various types of functional relationship data. A frequently updated database of computationally analyzed potential PPIs to provide biological researchers with rapid and easy access to analyze original data as a biological network is still lacking. RESULTS: By applying a probabilistic model, we integrated 27 heterogeneous genomic, proteomic and functional annotation datasets to predict PPI networks in human. In addition to previously studied data types, we show that phenotypic distances and genetic interactions can also be integrated to predict PPIs. We further built an easy-to-use, updatable integrated PPI database, the Integrated Network Database (IntNetDB) online, to provide automatic prediction and visualization of PPI network among genes of interest. The networks can be visualized in SVG (Scalable Vector Graphics) format for zooming in or out. IntNetDB also provides a tool to extract topologically highly connected network neighborhoods from a specific network for further exploration and research. Using the MCODE (Molecular Complex Detections) algorithm, 190 such neighborhoods were detected among all the predicted interactions. The predicted PPIs can also be mapped to worm, fly and mouse interologs. CONCLUSION: IntNetDB includes 180,010 predicted protein-protein interactions among 9,901 human proteins and represents a useful resource for the research community. Our study has increased prediction coverage by five-fold. IntNetDB also provides easy-to-use network visualization and analysis tools that allow biological researchers unfamiliar with computational biology to access and analyze data over the internet. The web interface of IntNetDB is freely accessible at http://hanlab.genetics.ac.cn/IntNetDB.htm. Visualization requires Mozilla version 1.8 (or higher) or Internet Explorer with installation of SVGviewer.

Algorithms↗

Sequence, annotation, and analysis of synteny between rice chromosome 3 and diverged grass species.

Rice (Oryza sativa L.) chromosome 3 is evolutionarily conserved across the cultivated cereals and shares large blocks of synteny with maize and sorghum, which diverged from rice more than 50 million years ago. To begin to completely understand this chromosome, we sequenced, finished, and annotated 36.1 Mb ( approximately 97%) from O. sativa subsp. japonica cv Nipponbare. Annotation features of the chromosome include 5915 genes, of which 913 are related to transposable elements. A putative function could be assigned to 3064 genes, with another 757 genes annotated as expressed, leaving 2094 that encode hypothetical proteins. Similarity searches against the proteome of Arabidopsis thaliana revealed putative homologs for 67% of the chromosome 3 proteins. Further searches of a nonredundant amino acid database, the Pfam domain database, plant Expressed Sequence Tags, and genomic assemblies from sorghum and maize revealed only 853 nontransposable element related proteins from chromosome 3 that lacked similarity to other known sequences. Interestingly, 426 of these have a paralog within the rice genome. A comparative physical map of the wild progenitor species, Oryza nivara, with japonica chromosome 3 revealed a high degree of sequence identity and synteny between these two species, which diverged approximately 10,000 years ago. Although no major rearrangements were detected, the deduced size of the O. nivara chromosome 3 was 21% smaller than that of japonica. Synteny between rice and other cereals using an integrated maize physical map and wheat genetic map was strikingly high, further supporting the use of rice and, in particular, chromosome 3, as a model for comparative studies among the cereals.

Arabidopsis↗

Proteomic analysis of differential protein expression induced by ultraviolet light radiation in HeLa cells.

Cells treated with ultraviolet (UV) radiation undergo cell cycle arrest at the S-phase and G1/S boundary, allowing DNA repair to occur. Several proteins such as replication protein A and DNA-dependent protein kinase have been suggested to be involved in UV-induced inhibition of DNA replication. However, the role of these proteins in inhibiting DNA replication remains unknown. Other proteins may play important roles in modulating functions of these proteins in response to UV-irradiation. To understand the broad range of proteins involved in this inhibition, we carried out a systematic study to identify specific proteins involved in UV-induced replication arrest using two-dimensional gel electrophoresis and mass spectrometry. Unique changes in protein expression level for 31 proteins were observed over a 24-hour time course, including calgizzarin, cyclophilin A, and macrophage migration inhibitory factor. The expression level changes of these proteins are dynamically correlated to DNA replication activity, suggesting involvement of these proteins in modulating DNA replication and repair activities. This proteomic approach provides opportunities to gain insights into the mechanism by which DNA replication is inhibited.

DNA Repair↗

Proteomic characterization of wheat amyloplasts using identification of proteins by tandem mass spectrometry.

We describe the initial characterization of the wheat amyloplast proteome, consisting of the identification and classification of 171 proteins. Whole amyloplasts and purified amyloplast membranes were prepared from wheat (Triticum aestivum). Protein extracts were examined by one-dimensional and two-dimensional electrophoresis, followed by high performance liquid chromatography-tandem mass spectrometry of separated proteins. Tandem mass spectrometry data of individual peptides was then searched by SEQUEST, using a database containing known protein sequences from both wheat and other homologous cereal crops. Using this approach we identified 108 proteins from whole amyloplasts and 63 proteins from purified amyloplast membranes. The majority of protein identifications were derived from protein sequences from cereal crops other than wheat, for which relatively little gene sequence data is available. The highest percentage of protein identifications obtained from any individual species was 46% of the total number of proteins identified, using sequence data found in our proprietary rice (Oryza sativa) genome database.

Chromatography, High Pressure Liquid↗

TcruziDB: an integrated, post-genomics community resource for Trypanosoma cruzi.

TcruziDB (http://TcruziDB.org) is an integrated post-genomics database for the parasitic organism, Trypanosoma cruzi, the causative agent of Chagas' disease. TcruziDB was established in 2003 as a flat-file database with tools for mining the unannotated sequence reads and preliminary contig assemblies emerging from the Tri-Tryp genome consortium (TIGR/SBRI/Karolinska). Today, TcruziDB houses the recently published assembled genomic contigs and annotation provided by the genome consortium in a relational database supported by the Genomics Unified Schema (GUS) architecture. The combination of an annotated genome and a relational architecture has facilitated the integration of genomic data with expression data (proteomic and EST) and permitted the construction of automated analysis pipelines. TcruziDB has accepted, and will continue to accept the deposition of genomic and functional genomic datasets contributed by the research community.

Animals↗

Analysis of proteins and proteomes by mass spectrometry.

A decade after the discovery of electrospray and matrix-assisted laser desorption ionization (MALDI), methods that finally allowed gentle ionization of large biomolecules, mass spectrometry has become a powerful tool in protein analysis and the key technology in the emerging field of proteomics. The success of mass spectrometry is driven both by innovative instrumentation designs, especially those operating on the time-of-flight or ion-trapping principles, and by large-scale biochemical strategies, which use mass spectrometry to detect the isolated proteins. Any human protein can now be identified directly from genome databases on the basis of minimal data derived by mass spectrometry. As has already happened in genomics, increased automation of sample handling, analysis, and the interpretation of results will generate an avalanche of qualitative and quantitative proteomic data. Protein-protein interactions can be analyzed directly by precipitation of a tagged bait followed by mass spectrometric identification of its binding partners. By these and similar strategies, entire protein complexes, signaling pathways, and whole organelles are being characterized. Posttranslational modifications remain difficult to analyze but are starting to yield to generic strategies.

Chromatography, Liquid↗

Bio-support vector machines for computational proteomics.

MOTIVATION: One of the most important issues in computational proteomics is to produce a prediction model for the classification or annotation of biological function of novel protein sequences. In order to improve the prediction accuracy, much attention has been paid to the improvement of the performance of the algorithms used, few is for solving the fundamental issue, namely, amino acid encoding as most existing pattern recognition algorithms are unable to recognize amino acids in protein sequences. Importantly, the most commonly used amino acid encoding method has the flaw that leads to large computational cost and recognition bias. RESULTS: By replacing kernel functions of support vector machines (SVMs) with amino acid similarity measurement matrices, we have modified SVMs, a new type of pattern recognition algorithm for analysing protein sequences, particularly for proteolytic cleavage site prediction. We refer to the modified SVMs as bio-support vector machine. When applied to the prediction of HIV protease cleavage sites, the new method has shown a remarkable advantage in reducing the model complexity and enhancing the model robustness.

Algorithms↗

Nano-high-performance liquid chromatography in combination with nano-electrospray ionization Fourier transform ion-cyclotron resonance mass spectrometry for proteome analysis.

Fourier transform ion-cyclotron resonance (FTICR) mass spectrometry offers several advantages for the analysis of biological samples, including excellent mass resolution, ultra-high mass measurement accuracy, high sensitivity, and wide mass range. We report the application of a nano-HPLC system coupled to an FTICR mass spectrometer equipped with nanoelectrospray source (nano-HPLC/nano-ESI-FTICRMS) for proteome analysis. Protein identification in proteomics is usually conducted by accurately determining peptide masses resulting from enzymatic protein digests and comparing them with theoretically digested protein sequences from databases. A tryptic in-solution digest of bovine serum albumin was used to optimize experimental conditions and data processing. Spots from Coomassie Blue and silver-stained two-dimensional (2D) gels of human thyroid tissue were excised, in-gel digested with trypsin, and subsequently analyzed by nano-HPLC/nano-ESI-FTICRMS. Additionally, we analyzed 1D-gel bands of membrane preparations of COS-6 cells from African green monkey kidney as an example of more complex protein mixtures. Nano-HPLC was performed using 1-mm reverse-phase C-18 columns for pre-concentration of the samples and reverse-phase C-18 capillary columns for separation, applying water/acetonitrile gradient elution conditions at flow rates of 200 nL/min. Mass measurement accuracies smaller than 3 ppm were routinely obtained. Different methods for processing the raw data were compared in order to identify a maximum number of peptides with the highest possible degree of automation. Parallel identification of proteins from complex mixtures down to low-femtomole levels makes nano-HPLC/nano-ESI-FTICRMS an attractive approach for proteome analysis.

Animals↗

Proteome annotations and identifications of the human pulmonary fibroblast.

We hereby report on a three year project initiative undertaken by our research team encompassing large-scale protein expression profiling and annotations of human primary lung fibroblast cells. An overview is given of proteomic studies of the fibroblast target cell involved in several diseases such as asthma, idiopatic pulmonary disease, and COPD. It has been the objective within our research team to map and identify the protein expressions occurring in both activated-, as well as resting cell states. The JGGL database www.2DDB.org has been built around these data, allowing advanced hypothesis building using the interactive query bioinformatic tools developed. Gene ontology has been applied to these annotations, classifying and correlating protein expressions to function. The localization as well as the biological processes involved for the annotations are being presented including an annotation-, and sequence-identification strategy, resulting in close to 2000 protein identities. Both gel based, high resolution 2D-gels, and liquid-phase separation (three-dimensional HPLC), as well as the combination of gel- and LC-based approaches (1D-gels and nano-capillary LC, reversed-phase) were utilized. Protein sequencing and structure identities were acquired by a combination of MALDI-, and electrospray-mass spectrometry techniques. Phenotypical and morphological characterizations were also made for this human disease target cell in both stimulated- and resting-cell states. The use of functional assays that demonstrate the key regulating role of growth factors and cytokine stimuli such as PDGF, TGF-beta, and EGF and the effect of ECM molecules such as Biglycan, are also presented and discussed.

Amino Acid Sequence↗

Proteomics of Chlamydomonas reinhardtii light-harvesting proteins.

With the recent development of techniques for analyzing transmembrane thylakoid proteins by two-dimensional gel electrophoresis, systematic approaches for proteomic analyses of membrane proteins became feasible. In this study, we established detailed two-dimensional protein maps of Chlamydomonas reinhardtii light-harvesting proteins (Lhca and Lhcb) by extensive tandem mass spectrometric analysis. We predicted eight distinct Lhcb proteins. Although the major Lhcb proteins were highly similar, we identified peptides which were unique for specific lhcbm gene products. Interestingly, lhcbm6 gene products were resolved as multiple spots with different masses and isoelectric points. Gene tagging experiments confirmed the presence of differentially N-terminally processed Lhcbm6 proteins. The mass spectrometric data also revealed differentially N-terminally processed forms of Lhcbm3 and phosphorylation of a threonine residue in the N terminus. The N-terminal processing of Lhcbm3 leads to the removal of the phosphorylation site, indicating a potential novel regulatory mechanism. At least nine different lhca-related gene products were predicted by comparison of the mass spectrometric data against Chlamydomonas expressed sequence tag and genomic databases, demonstrating the extensive variability of the C. reinhardtii Lhca antenna system. Out of these nine, three were identified for the first time at the protein level. This proteomic study demonstrates the complexity of the light-harvesting proteins at the protein level in C. reinhardtii and will be an important basis of future functional studies addressing this diversity.

Amino Acid Sequence↗

Experimental analysis of the Arabidopsis mitochondrial proteome highlights signaling and regulatory components, provides assessment of targeting prediction programs, and indicates plant-specific mitochondrial proteins.

A novel insight into Arabidopsis mitochondrial function was revealed from a large experimental proteome derived by liquid chromatography-tandem mass spectrometry. Within the experimental set of 416 identified proteins, a significant number of low-abundance proteins involved in DNA synthesis, transcriptional regulation, protein complex assembly, and cellular signaling were discovered. Nearly 20% of the experimentally identified proteins are of unknown function, suggesting a wealth of undiscovered mitochondrial functions in plants. Only approximately half of the experimental set is predicted to be mitochondrial by targeting prediction programs, allowing an assessment of the benefits and limitations of these programs in determining plant mitochondrial proteomes. Maps of putative orthology networks between yeast, human, and Arabidopsis mitochondrial proteomes and the Rickettsia prowazekii proteome provide detailed insights into the divergence of the plant mitochondrial proteome from those of other eukaryotes. These show a clear set of putative cross-species orthologs in the core metabolic functions of mitochondria, whereas considerable diversity exists in many signaling and regulatory functions.

Arabidopsis↗

Centralized data analysis of a large interlaboratory proteomics project: a feasibility study.

The human Plasma Proteome Project (PPP) is a large-scale collaboration between many laboratories. One of the most demanding tasks in the PPP involved the analysis of very large amounts of raw MS/MS data produced by the participants. The main approach for managing this task was letting the participants analyze their own data and submit the results to the central PPP repository as lists of identified proteins and peptides. To complement this distributed approach, we also performed centralized analysis of the raw MS/MS data provided by the participants. Due to the data redundancy inherent in such a project, centralized analysis has the potential to reduce the computational effort by reducing redundancy before the analysis. Centralized analysis can also unify the process and take advantage of data sharing among laboratories to improve protein identification and validation. The process we employed included removing low-quality spectra, clustering spectra by mutual similarity, and applying uniform peptide and protein identification procedures. To demonstrate the process, we analyzed 5.28 million MS/MS spectra derived by eight laboratories from tryptic peptides of serum and plasma proteins.

Blood Proteins↗

Plant protein annotation in the UniProt Knowledgebase.

The Swiss-Prot, TrEMBL, Protein Information Resource (PIR), and DNA Data Bank of Japan (DDBJ) protein database activities have united to form the Universal Protein Resource (UniProt) Consortium. UniProt presents three database layers: the UniProt Archive, the UniProt Knowledgebase (UniProtKB), and the UniProt Reference Clusters. The UniProtKB consists of two sections: UniProtKB/Swiss-Prot (fully manually curated entries) and UniProtKB/TrEMBL (automated annotation, classification and extensive cross-references). New releases are published fortnightly. A specific Plant Proteome Annotation Program (http://www.expasy.org/sprot/ppap/) was initiated to cope with the increasing amount of data produced by the complete sequencing of plant genomes. Through UniProt, our aim is to provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information that will allow the plant community to fully explore and utilize the wealth of information available for both plant and non-plant model organisms.

Amino Acid Sequence↗

Application of proteome analysis to seafood authentication.

Part of our work aims at studying the modifications that proteins suffer in foods and use them as markers to estimate the origin and history of the product. Proteomics is a powerful approach to do this: comparison of the two-dimensional (2-D) maps of the intact and treated samples would permit to identify marker spots so that in the future it may be possible to estimate the treatment a foodstuff has suffered by examining its 2-D protein map or just the selected markers. This work summarizes some of our previous studies showing the application of proteomics to the (i) identification of species and muscle tissues, (ii) characterization of post-mortem changes in arctic and tropical species, and (iii) study of the effect of some additives during the processing of fish muscle.

Animals↗

Open mass spectrometry search algorithm.

Large numbers of MS/MS peptide spectra generated in proteomics experiments require efficient, sensitive and specific algorithms for peptide identification. In the Open Mass Spectrometry Search Algorithm (OMSSA), specificity is calculated by a classic probability score using an explicit model for matching experimental spectra to sequences. At default thresholds, OMSSA matches more spectra from a standard protein cocktail than a comparable algorithm. OMSSA is designed to be faster than published algorithms in searching large MS/MS datasets.

Algorithms↗

Neurogenomics: at the intersection of neurobiology and genome sciences.

Neurogenomics is the study of how the genome as a whole contributes to the evolution, development, structure and function of the nervous system. It includes investigations of how genome products (transcriptomes and proteomes) vary in time and space. Neurogenomics differs markedly from the application of genome sciences to other systems, particularly in the spatial category, because anatomy and connectivity are paramount to our understanding of function in the nervous system. We focus here on some of the influences of genomics and its associated technologies on neuroscience. We discuss comparative genomics, gene expression atlases of the brain, network genetics and applications to behavioral phenotypes, and consider the culture, organization and funding of genome-scale projects.

Animals↗

Annotated proteome of a human T-cell lymphoma.

As the reliable identification of proteins by tandem mass spectrometry becomes increasingly common, the full characterization of large data sets of proteins remains a difficult challenge. Our goal was to survey the proteome of a human T-cell lymphoma-derived cell line in a single set of experiments and present an automated method for the annotation of lists of proteins. A downstream application of these data includes the identification of novel pathogenetic and candidate diagnostic markers of T-cell lymphoma. Total protein isolated from cytoplasmic, membrane, and nuclear fractions of the SUDHL-1 T-cell lymphoma cell line was resolved by SDS-PAGE, and the entire gel lanes digested and analyzed by tandem mass spectrometry. Acquired data files were searched against the UniProt protein database using the SEQUEST algorithm. Search results for each subcellular fraction were analyzed using INTERACT and ProteinProphet. All protein identifications with an error rate of less than 10% were directly exported into excel and analyzed using GOMiner (NIH/NCI). The Gene ontology molecular function and cell location data were summarized for the identified proteins and results exported as user-interactive directed acyclic graphs. A total of 1105 unique proteins were identified and fully annotated, including numerous proteins that had not been previously characterized in lymphoma, in functional categories such as cell adhesion, migration, signaling, and stress response. This study demonstrates the utility of currently available bioinformatics tools for the robust identification and annotation of large numbers of proteins in a batchwise fashion.

Algorithms↗

Widespread distribution of antisense transcripts in the Plasmodium falciparum genome.

The availability of the complete genome sequence of Plasmodium falciparum has facilitated high-throughput profiling of its complex life cycle, following the application of micro-array, proteomic, and serial analysis of gene expression (SAGE) technologies in this system. These, in turn, have yielded unprecedented insight into global gene expression, including the foremost demonstration of antisense transcription in the parasite. For example, owing to its inherent ability to sample novel ORFs and to predict transcript orientation, SAGE analysis in asexual forms led to the initial discovery of highly abundant antisense RNAs. To determine the extent of this phenomenon in P. falciparum, we have surveyed the distribution of both sense and antisense transcripts across the asexual transcriptome for the first time. To this end, a relational database integrating SAGE expression data with genome annotation information was constructed. This allowed the comprehensive annotation of a total of 17245 SAGE tags, extending over a 350-fold expression range. Transcripts from approximately 30% of the estimated 3D7 gene loci were present at detectable levels in mixed asexual stages, where loci involved in invasion and immune evasion; and carbohydrate metabolism were highly represented in the sense transcriptome. Approximately 12% of SAGE tags, however, were derived from the non-coding strand of nuclear-encoded ORFs, indicating that endogenous antisense RNAs are widespread in this system. Notably, these antisense transcripts were absent from the mitochondrial genome. Interestingly, we note that sense and antisense tag counts from single loci across the transcriptome were inversely related. Taken together, this data may provide first hints as to the possible function of antisense transcription in this system.

Animals↗