Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Bioinformatic approaches to assigning protein function from novel sequence data.

The current pace of functional genomic initiatives and genome sequencing projects has provided researchers with a bewildering array of sequence and biological data to analyze. The disease system-driven approach to identifying key genes frequently identifies nucleotide and protein sequences for which the gene and protein function are not known in sufficient detail to allow informed follow-up. Using a range of bioinformatic tools and sequence-based clues, most of unassigned sequences can now be annotated. This chapter takes as an example an unannotated expressed sequence tag, describing how to identify its related gene, and how to annotate the encoded protein using sequence, profile, and structure-based annotation methodologies.

Computational Biology↗

PathAligner: metabolic pathway retrieval and alignment.

MOTIVATION: Analysis of metabolic pathways is a central topic in understanding the relationship between genotype and phenotype. The rapid accumulation of biological data provides the possibility of studying metabolic pathways at both the genomic and the metabolic levels. Retrieving metabolic pathways from current biological data sources, reconstructing metabolic pathways from rudimentary pathway components, and aligning metabolic pathways with each other are major tasks. Our motivation was to develop a conceptual framework and computational system that allows the retrieval of metabolic pathway information and the processing of alignments to reveal the similarities between metabolic pathways. RESULTS: PathAligner extracts metabolic information from biological databases via the Internet and builds metabolic pathways with data sources of genes, sequences, enzymes, metabolites etc. It provides an easy-to-use interface to retrieve, display and manipulate metabolic information. PathAligner also provides an alignment method to compare the similarity between metabolic pathways. AVAILABILITY: PathAligner is available at http://bibiserv.techfak.uni-bielefeld.de/pathaligner.

Algorithms↗

Large-scale prediction of protein structure and function from sequence.

The identification of novel drug targets from genomic data involves the large-scale analysis of many protein sequences. Methods for automated structure and function prediction are an essential tool for this purpose. In this review we concentrate on the recent developments in the field of protein structure prediction and how these can be used to gain hints about the function of proteins. The current state-of-the-art is highlighted through recent community-wide experiments aimed at comparing different approaches. For structure prediction this allows the identification of key improvements to increase the crucial sequence to structure alignment needed for accurate models. Function prediction is a rapidly maturing field that is still being benchmarked. Definitions for protein function are presented and available methods, mostly concentrating on functional site descriptors and structural motifs, presented.

Algorithms↗

[Bioinformatics].

Explore the source record for details and available documents.

Computational Biology↗

Protein structure prediction and analysis as a tool for functional genomics.

Bioinformatic analyses of whole genome sequences highlight the problem of identifying the biochemical and cellular functions of the many gene products that are at present uncharacterised. Determination of their three-dimensional structures, either experimentally or by prediction, provides a powerful tool to address function, since it is at this level that biological activity is expressed. Here, we discuss the current approaches to protein structure prediction from sequence data, including the ab initio prediction of new folds, methods of fold recognition and comparative modelling based on homology. The value and limitations of such models are also explored. A major factor for the future will be the growth of the database of experimentally determined protein structures, through structural genomics projects. The prospects for this approach are also discussed, together with our experience in a pilot structural genomics project focused on proteins from Mycobacterium tuberculosis, the cause of tuberculosis (TB).

Computer Simulation↗

Evaluation of prefractionation methods as a preparatory step for multidimensional based chromatography of serum proteins.

Prefractionations of proteins prior to their proteolysis, chromatography, and MS/MS analyses help reduce complexity and increase the yield of protein identifications. A number of methods were evaluated here for prefractionating serum samples distributed to the participating laboratories as part of the human Plasma Proteome Project. These methods include strong cation exchange (SCX) chromatography, slicing of SDS-PAGE gel bands, and liquid-phase IEF of the proteins. The fractionated proteins were trypsinized and the resulting peptides were resolved and analyzed by multidimensional protein identification technology coupled to IT MS/MS. The MS/MS spectra were clustered, combined, and searched against the IPI protein databank using Pep-Miner. The identification results were evaluated for the efficacy of the different prefractionation methodologies to identify larger numbers of proteins at higher confidence and to achieve the best coverage of the proteins with the identified peptides. Prefractionation based on SCX resulted in the largest number of identified proteins, followed by gel slices and then the liquid-phase IEF. An important observation was that each of the methods revealed a set of unique proteins, some identified with high confidence. Therefore, for comprehensive identification of the serum proteins, several different prefractionation approaches should be used in parallel.

Blood Proteins↗

Prediction of protein-protein interaction based on structure.

A great challenge in the proteomics and structural genomics era is to predict protein structure and function from sequence, including the identification of biological partners. The development of a procedure to construct position-specific scoring matrices for the prediction and identification of sequences with putative significant affinity faces this challenge. The local and web applications used for sequence and structure search, sequence alignment, protein modeling, molecule edition and modification, and scoring matrices construction are described in detail. The methodology is based on the information contained in structural databases and takes into account the subtle conformational and sequence details that characterize different structures within a family. Using the matrices, the protein sequence databases can be easily scanned to locate putative partners of biological significance. The success of this methodology opens the way for the prediction of protein-protein interaction at genome scale.

Algorithms↗

Profiling of abundant proteins associated with dichlorodiphenyltrichloroethane resistance in Drosophila melanogaster.

Dichlorodiphenyltrichloroethane (DDT) metabolism-based resistance in Drosophila melanogaster is a complex metabolic system associated with the transcription of detoxification related genes, ion transport, lipid and sugar metabolism pathways. However, little is known about the differences regarding the proteome of field- and laboratory-selected resistant Drosophila genotypes. We investigated the impact of DDT resistance in the abundant proteome of field- and laboratory- selected resistant Drosophila using a two-dimensional gel electrophoresis DDT reference map. Proteomic profiling was performed in two DDT susceptible genotypes (Canton-S and 91-C) and three DDT resistant lines (Rst(2)DDT(91-R), Rst(2)DDT(Wisconsin) and Rst(2)DDT(Hikone-R)). Protein spots were stained with Coomassie blue and compared using PDQuest software. Selected protein spots were cut out and analyzed using matrix assisted laser desorption-time of flight mass spectrometry. Querying the NCBInr. 10.21.2003 database with mass spectrometric data yielded the identity of 21 differentially translated proteins in Rst(2)DDT(91-R), Rst(2)DDT(Wisconsin) and Canton-S representing proteins putatively involved in biochemical pathways such as glycolysis and gluconeogenesis, the pentose phosphate pathway, the Krebs cycle and fatty acid oxidation. We hypothesize that both strategies are aimed to use of the pentose phosphate pathway to increase glucose utilization while Rst(2)DDT(91-R) relies primarily on glycolysis to produce reduced NADP and increase DDT detoxification. DDT exposure in Canton-S induced six proteins, while four proteins were repressed in Rst(2)DDT(Hikone-R). Our data suggest that insecticide resistance appears to impact different metabolic pathways in Drosophila genotypes selected with the same pesticide (DDT).

Animals↗

Statistical model for large-scale peptide identification in databases from tandem mass spectra using SEQUEST.

Recent technological advances have made multidimensional peptide separation techniques coupled with tandem mass spectrometry the method of choice for high-throughput identification of proteins. Due to these advances, the development of software tools for large-scale, fully automated, unambiguous peptide identification is highly necessary. In this work, we have used as a model the nuclear proteome from Jurkat cells and present a processing algorithm that allows accurate predictions of random matching distributions, based on the two SEQUEST scores Xcorr and DeltaCn. Our method permits a very simple and precise calculation of the probabilities associated with individual peptide assignments, as well as of the false discovery rate among the peptides identified in any experiment. A further mathematical analysis demonstrates that the score distributions are highly dependent on database size and precursor mass window and suggests that the probability associated with SEQUEST scores depends on the number of candidate peptide sequences available for the search. Our results highlight the importance of adjusting the filtering criteria to discriminate between correct and incorrect peptide sequences according to the circumstances of each particular experiment.

Chromatography, Liquid↗

Protein expression in a transformed trabecular meshwork cell line: proteome analysis.

PURPOSE: Characterization of the human trabecular meshwork (TM) proteome is hindered by the small mass of intact tissue and the slow growth of cultured cell strains. We have previously characterized a transformed TM cell strain (GTM3) that demonstrates many of the same protein expression and cell signaling systems of nontransformed cell strains. The aim of this study was to initiate a proteomic survey of GTM3 cells as the initial step toward characterization of the complete human TM proteome. METHODS: GTM3 cells were cultured to confluence, harvested and solubilized in urea/Nonidet. The protein extract (600 mug) was focused in immobilized isoelectric focusing (IEF) strips, separated by 10% SDS PAGE, and visualized with colloidal Coomassie Blue. Spots of interest were excised, destained, and the contained proteins subjected to in-gel reduction, derivatization, and tryptic digestion. Tryptic peptides were extracted and analyzed by electrospray LC/MS/MS. Protein identification was made using the TurboSequest search algorithm and a recent version of the nonredundant human protein database downloaded from the National Center for Biotechnology Information (NCBI). RESULTS: Eighty-seven (87) primary proteins and 93 variants of these proteins were identified. A website was created (TM proteome) that combines data such as graphic spot location within the gel, peptide sequence, apparent and calculated pI, apparent and calculated mass, percentage of coverage, and protein informatic website links. CONCLUSIONS: Proteomic analysis of a transformed human TM cell line has been initiated combining preparative two-dimensional PAGE separation, LC/MS/MS analysis of major proteins, and bioinformatic cataloging of the data. Further investigation of data from the transformed cell strain will be used in a comparative fashion for spot identification of analytical proteomic gels of human TM tissue and cultured normal cells. These initial data will form the base from which the characterization of protein expression in the normal and glaucomatous TM can be accomplished.

Cell Line, Transformed↗

A statistical framework to discover true associations from multiprotein complex pull-down proteomics data sets.

Experimental processes to collect and process proteomics data are increasingly complex, and the computational methods to assess the quality and significance of these data remain unsophisticated. These challenges have led to many biological oversights and computational misconceptions. We developed an empirical Bayes model to analyze multiprotein complex (MPC) proteomics data derived from peptide mass spectrometry detections of purified protein complex pull-down experiments. Using our model and two yeast proteomics data sets, we estimated that there should be an average of about 20 true associations per MPC, almost 10 times as high as was previously estimated. For data sets generated to mimic a real proteome, our model achieved on average 80% sensitivity in detecting true associations, as compared with the 3% sensitivity in previous work, while maintaining a comparable false discovery rate of 0.3%. Cross-examination of our results with protein complexes confirmed by various experimental techniques demonstrates that many true associations that cannot be identified by previous approach are identified by our method.

Algorithms↗

Proteomic mass spectra classification using decision tree based ensemble methods.

MOTIVATION: Modern mass spectrometry allows the determination of proteomic fingerprints of body fluids like serum, saliva or urine. These measurements can be used in many medical applications in order to diagnose the current state or predict the evolution of a disease. Recent developments in machine learning allow one to exploit such datasets, characterized by small numbers of very high-dimensional samples. RESULTS: We propose a systematic approach based on decision tree ensemble methods, which is used to automatically determine proteomic biomarkers and predictive models. The approach is validated on two datasets of surface-enhanced laser desorption/ionization time of flight measurements, for the diagnosis of rheumatoid arthritis and inflammatory bowel diseases. The results suggest that the methodology can handle a broad class of similar problems.

Algorithms↗

Proteomic analysis of rat aorta during atherosclerosis induced by high cholesterol diet and injection of vitamin D3.

1. Atherosclerosis (AS) in rats displays important clinical similarities to human AS. 2. After the experimental model of AS in rat was established and using a proteomic approach, we compared the protein profiling of aorta tissues from healthy and AS rats. 3. Using two-dimensional electrophoresis (2-DE), over 1878 protein species were separated; among them, 1239 protein spots were matched between different gels with average matching rate of approximately 66%. Gel analysis and protein characterization have identified 58 protein spots whose abundance is significantly altered in AS rats. 4. By using matrix-associated laser desorption ionization time-of-flight mass spectrometer (MALDI-TOF-MS) and NCBInr database, 46 proteins were successfully identified. Among them, 18 proteins were of increased abundance in diseased tissues including a group of oxidization-related enzymes such as peroxiredoxin2 and NADH dehydrogenase Fe-S protein 6, components of inflammatory pathways such as lamin A, while 28 proteins were of decreased abundance in the diseased state, including CaM-KII inhibitory protein, transferring, fructose-bisphosphate aldolase. 5. We believe that these results would give insights into the cellular and molecular mechanisms involved in AS development and might lead to the discovery of novel diagnostic markers and new therapeutic opportunities.

Animals↗

Identifying cytotoxic T cell epitopes from genomic and proteomic information: "The human MHC project.".

Complete genomes of many species including pathogenic microorganisms are rapidly becoming available and with them the encoded proteins, or proteomes. Proteomes are extremely diverse and constitute unique imprints of the originating organisms allowing positive identification and accurate discrimination, even at the peptide level. It is not surprising that peptides are key targets of the immune system. It follows that proteomes can be translated into immunogens once it is known how the immune system generates and handles peptides. Recent advances have identified many of the basic principles involved. The single most selective event is that of peptide binding to MHC, making it particularly important to establish accurate descriptions and predictions of peptide binding for the most common MHC variants. These predictions should be integrated with those of other steps involved in antigen processing, as these become available. The ability to translate the accumulating primary sequence databases in terms of immune recognition should enable scientists and clinicians to analyze any protein of interest for the presence of potentially immunogenic epitopes. The computational tools to scan entire proteomes should also be developed, as this would enable a rational approach to vaccine development and immunotherapy. Thus, candidate vaccine epitopes might be predicted from the various microbial genome projects, tumor vaccine candidates from mRNA expression profiling of tumors ("transcriptomes") and auto-antigens from the human genome.

Antigen Presentation↗