Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

MASCOT HTML and XML parser: an implementation of a novel object model for protein identification data.

Protein identification using MS is an important technique in proteomics as well as a major generator of proteomics data. We have designed the protein identification data object model (PDOM) and developed a parser based on this model to facilitate the analysis and storage of these data. The parser works with HTML or XML files saved or exported from MASCOT MS/MS ions search in peptide summary report or MASCOT PMF search in protein summary report. The program creates PDOM objects, eliminates redundancy in the input file, and has the capability to output any PDOM object to a relational database. This program facilitates additional analysis of MASCOT search results and aids the storage of protein identification information. The implementation is extensible and can serve as a template to develop parsers for other search engines. The parser can be used as a stand-alone application or can be driven by other Java programs. It is currently being used as the front end for a system that loads HTML and XML result files of MASCOT searches into a relational database. The source code is freely available at http://www.ccbm.jhu.edu and the program uses only free and open-source Java libraries.

Databases, Protein↗

Gene3D: modelling protein structure, function and evolution.

The Gene3D release 4 database and web portal (http://cathwww.biochem.ucl.ac.uk:8080/Gene3D) provide a combined structural, functional and evolutionary view of the protein world. It is focussed on providing structural annotation for protein sequences without structural representatives--including the complete proteome sets of over 240 different species. The protein sequences have also been clustered into whole-chain families so as to aid functional prediction. The structural annotation is generated using HMM models based on the CATH domain families; CATH is a repository for manually deduced protein domains. Amongst the changes from the last publication are: the addition of over 100 genomes and the UniProt sequence database, domain data from Pfam, metabolic pathway and functional data from COGs, KEGG and GO, and protein-protein interaction data from MINT and BIND. The website has been rebuilt to allow more sophisticated querying and the data returned is presented in a clearer format with greater functionality. Furthermore, all data can be downloaded in a simple XML format, allowing users to carry out complex investigations at their own computers.

Databases, Protein↗

Clinical-scale high-throughput human plasma proteome analysis: lung adenocarcinoma.

Clinical proteomics requires the stable and reproducible analysis of a large number of human samples. We report a high-throughput comprehensive protein profiling system comprising a fully automated, on-line, two-dimensional microflow liquid chromatography/tandem mass spectrometry (2-D microLC-MS/MS) system for use in clinical proteomics. A linear ion-trap mass spectrometer (ITMS) also known as a 2-D ITMS instrument, which is characterized by high scan speed, was incorporated into the microLC-MS/MS system in order to obtain highly improved sensitivity and resolution in MS/MS acquisition. This system was used to evaluate bovine serum albumin and human 26S proteasome. Application of these high-throughput microLC conditions and the 2-D ITMS resulted in a 10-fold increase in sensitivity in protein identification. Additionally, peptide fragments from the 26S proteasome were identified three-fold more efficiently than by the conventional 3-D ITMS instrument. In this study, the 2-D microLC-MS/MS system that uses linear 2-D ITMS has been applied for the plasma proteome analysis of a few samples from healthy individuals and lung adenocarcinoma patients. Using the 2-D and 1-D microLC-MS/MS analyses, approximately 250 and 100 different proteins were detected, respectively, in each HSA- and IgG-depleted sample, which corresponds to only 0.4 microL of blood plasma. Automatic operation enabled the completion of a single run of the entire 1-D and 2-D microLC-MS/MS analyses within 11 h. Investigation of the data extracted from the protein identification datasets of both healthy and adenocarcinoma groups revealed that several of the group-specific proteins could be candidate protein disease markers expressed in the human blood plasma. Consequently, it was demonstrated that this high-throughput microLC-MS/MS protein profiling system would be practically applicable to the discovery of protein disease markers, which is the primary objective in clinical plasma proteome projects.

Adenocarcinoma↗

Organism complexity anti-correlates with proteomic beta-aggregation propensity.

We introduce a novel approach to estimate differences in the beta-aggregation potential of eukaryotic proteomes. The approach is based on a statistical analysis of the beta-aggregation propensity of polypeptide segments, which is calculated by an equation derived from first principles using the physicochemical properties of the natural amino acids. Our analysis reveals a significant decreasing trend of the overall beta-aggregation tendency with increasing organism complexity and longevity. A comparison with randomized proteomes shows that natural proteomes have a higher degree of polarization in both low and high beta-aggregation prone sequences. The former originates from the requirement of intrinsically disordered proteins, whereas the latter originates from the necessity of proteins with a stable folded structure.

Amino Acid Sequence↗

MACSIMS: multiple alignment of complete sequences information management system.

BACKGROUND: In the post-genomic era, systems-level studies are being performed that seek to explain complex biological systems by integrating diverse resources from fields such as genomics, proteomics or transcriptomics. New information management systems are now needed for the collection, validation and analysis of the vast amount of heterogeneous data available. Multiple alignments of complete sequences provide an ideal environment for the integration of this information in the context of the protein family. RESULTS: MACSIMS is a multiple alignment-based information management program that combines the advantages of both knowledge-based and ab initio sequence analysis methods. Structural and functional information is retrieved automatically from the public databases. In the multiple alignment, homologous regions are identified and the retrieved data is evaluated and propagated from known to unknown sequences with these reliable regions. In a large-scale evaluation, the specificity of the propagated sequence features is estimated to be >99%, i.e. very few false positive predictions are made. MACSIMS is then used to characterise mutations in a test set of 100 proteins that are known to be involved in human genetic diseases. The number of sequence features associated with these proteins was increased by 60%, compared to the features available in the public databases. An XML format output file allows automatic parsing of the MACSIM results, while a graphical display using the JalView program allows manual analysis. CONCLUSION: MACSIMS is a new information management system that incorporates detailed analyses of protein families at the structural, functional and evolutionary levels. MACSIMS thus provides a unique environment that facilitates knowledge extraction and the presentation of the most pertinent information to the biologist. A web server and the source code are available at http://bips.u-strasbg.fr/MACSIMS/.

Algorithms↗

Differential proteomics via probabilistic peptide identification scores.

Relative quantitation is key to enable differential proteomics and hence answer biological questions by comparing samples. Classical approaches involve stable isotope labeling with/without spiked standards. Although stable isotopes may lead to precise results, their application is not straightforward. In Proteomics, 2004, 4, 2333-2351, we proposed an approach where we summed peptide identification scores to derive a semiquantitative abundance indicator. In this study, we combine such an indicator with a statistical test to detect differentially expressed proteins. We demonstrate the effectiveness of this method by using mixtures of purified proteins and human plasma spiked with proteins at low-nanomolar concentrations. The impact of the number of repeated experiments is discussed, and we show that the statistical test we use performs well with two to three repetitions, whereas a classical t-test would require at least four repetitions to achieve the same performance. Typically, 2.5-5-fold changes are detected with 90-95% confidence in human plasma. The method is finally characterized by deriving estimates of its false positive and negative rates. This new characterization is valid for a wider class of methods such as spectrum sampling (Liu, H.; Sadygov, R. G.; Yates, J. R. III. Anal. Chem. 2004, 76, 4193-4201).

Animals↗

DNA polymorphism detector: an automated tool that searches for allelic matches in public databases for discrepancies found in clone or cDNA sequences.

SUMMARY: DNA polymorphism detector (DPD) is a new web application developed to help automate the process of cDNA clone validation. DPD identifies and highlights discrepancies between any cDNA clone sequence and its expected reference sequence. To determine if these differences correspond to natural genetic polymorphisms (versus artifacts introduced during clone production or evaluation), DPD uses the discrepancies, along with flanking sequences, to search GenBank for identical matching strings. If matching DNA sequences are found, DPD verifies that they are from the same gene. The application then reports the discrepancy as a polymorphism along with the corresponding GenBank reference information. AVAILABILITY: DPD is currently hosted by the Harvard Institute of Proteomics at http://www.hip.harvard.edu

Cloning, Molecular↗

Expanding the proteome two-dimensional gel electrophoresis reference map of human renal cortex by peptide mass fingerprinting.

Proteomics methodologies hold great promise in basic renal research and clinical nephrology. The classical approach for proteomic analysis couples two-dimensional gel electrophoresis (2-DE) with protein identification by mass spectrometry, to produce more global information regarding normal protein expression and alterations in different physiological and pathological states. In this report we have expanded the identification of proteins in the renal cortex, improving the previously published map to facilitate the study of different diseases affecting the human kidney. About 250 spots were analyzed by peptide mass fingerprinting, 89 proteins and 74 isoforms for some of them were identified and implemented in the normal human renal cortex 2-DE reference map. This more comprehensive view of the proteome of the human renal cortex could be of invaluable help to the differential proteomic display of urological diseases.

Databases, Protein↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

Reconstruction and functional characterization of the human mitochondrial metabolic network based on proteomic and biochemical data.

Diverse datasets including genomic, proteomic, isotopomer, and DNA sequence variation are becoming available for human mitochondria. Thus there is a need to integrate these data within an in silico modeling framework where mitochondrial biology and related disorders can be studied and analyzed. This paper reports a reconstruction and characterization of the human mitochondrial metabolic network based on proteomic and biochemical data. The 189 reactions included in this reconstruction are both elementally and charge-balanced and are assigned to their respective cellular compartments (mitochondrial, cytosol, or extracellular). The capabilities of the reconstructed network to fulfill three metabolic functions (ATP production, heme synthesis, and mixed phospholipid synthesis) were determined. Network-based analysis of the mitochondrial energy conversion process showed that the overall ATP yield per glucose is 31.5. Network flexibility, characterized by allowable variation in reaction fluxes, was evaluated using flux variability analysis and analysis of all of the possible optimal flux distributions. Results showed that the network has high flexibility for the biosynthesis of heme and phospholipids but modest flexibility for maximal ATP production. A subset of all of the optimal network flux distributions, computed with respect to the three metabolic functions individually, was found to be highly correlated, suggesting that this set may contain physiological meaningful fluxes. Examinations of optimal flux distributions also identified correlated reaction sets that form functional modules in the network.

Amino Acids↗

Chlamydomonas reinhardtii proteomics.

Proteomics, based on the expanding genomic resources, has begun to reveal new details of Chlamydomonas reinhardtii biology. In particular, analyses focusing on subproteomes have already provided new insight into the dynamics and composition of the photosynthetic apparatus, the chloroplast ribosome, the oxidative phosphorylation machinery of the mitochondria, and the flagellum. It assisted to discovered putative new components of the circadian clockwork as well as shed a light on thioredoxin protein-protein interactions. In the future, quantitative techniques may allow large scale comparison of protein expression levels. Advances in software algorithms will likely improve the use of genomic databases for mass spectrometry (MS) based protein identification and validation of gene models that have been predicted from the genomic DNA sequences. Although proteomics has only been recently applied for exploring C. reinhardtii biology, it will likely be utilized extensively in the near future due to the already existing genetic, genomic, and biochemical tools.

Animals↗

Gaining confidence in high-throughput protein interaction networks.

Although genome-scale technologies have benefited from statistical measures of data quality, extracting biologically relevant pathways from high-throughput proteomics data remains a challenge. Here we develop a quantitative method for evaluating proteomics data. We present a logistic regression approach that uses statistical and topological descriptors to predict the biological relevance of protein-protein interactions obtained from high-throughput screens for yeast. Other sources of information, including mRNA expression, genetic interactions and database annotations, are subsequently used to validate the model predictions without bias or cross-pollution. Novel topological statistics show hierarchical organization of the network of high-confidence interactions: protein complex interactions extend one to two links, and genetic interactions represent an even finer scale of organization. Knowledge of the maximum number of links that indicates a significant correlation between protein pairs (correlation distance) enables the integrated analysis of proteomics data with data from genetics and gene expression. The type of analysis presented will be essential for analyzing the growing amount of genomic and proteomics data in model organisms and humans.

Algorithms↗

Characterization of bovine seminal plasma by proteomics.

Previous investigations of bovine seminal plasma (BSP) have revealed the identities of the three major proteins, BSP-PDC109, BSP-A3 and BSP-30 kDa, which together constitute about half of the total protein, as well as about 30 of the minor proteins. Analyses of BSP by 2-DE have revealed about 250 protein spots, suggesting that much of the BSP proteome remains undescribed. In this study, BSP has been analyzed by 2-D LC-based and SDS-PAGE-based proteomic methods. Ninety-nine proteins were identified, including 49 minor proteins that have not previously been described in seminal plasma of any species.

Animals↗

Addressing the intrinsic disorder bottleneck in structural proteomics.

The Center for Eukaryotic Structural Genomics (CESG), as part of the Protein Structure Initiative (PSI), has established a high-throughput structure determination pipeline focused on eukaryotic proteins. NMR spectroscopy is an integral part of this pipeline, both as a method for structure determinations and as a means for screening proteins for stable structure. Because computational approaches have estimated that many eukaryotic proteins are highly disordered, about 1 year into the project, CESG began to use an algorithm (the Predictor of Naturally Disordered Regions, PONDR to avoid proteins that were likely to be disordered. We report a retrospective analysis of the effect of this filtering on the yield of viable structure determination candidates. In addition, we have used our current database of results on 70 protein targets from Arabidopsis thaliana and 1 from Caenorhabditis elegans, which were labeled uniformly with nitrogen-15 and screened for disorder by NMR spectroscopy, to compare the original algorithm with 13 other approaches for predicting disorder from sequence. Our study indicates that the efficiency of structural proteomics of eukaryotes can be improved significantly by removing targets predicted to be disordered by an algorithm chosen to provide optimal performance.

Algorithms↗

Trypsin cleaves exclusively C-terminal to arginine and lysine residues.

Almost all large-scale projects in mass spectrometry-based proteomics use trypsin to convert protein mixtures into more readily analyzable peptide populations. When searching peptide fragmentation spectra against sequence databases, potentially matching peptide sequences can be required to conform to tryptic specificity, namely, cleavage exclusively C-terminal to arginine or lysine. In many published reports, however, significant numbers of proteins are identified by non-tryptic peptides. Here we use the sub-parts per million mass accuracy of a new ion trap Fourier transform mass spectrometer to achieve more than a 100-fold increased confidence in peptide identification compared with typical ion trap experiments and show that trypsin cleaves solely C-terminal to arginine and lysine. We find that non-tryptic peptides occur only as the C-terminal peptides of proteins and as breakup products of fully tryptic peptides N-terminal to an internal proline. Simulating lower mass accuracy led to a large number of proteins erroneously identified with non-tryptic peptide hits. Our results indicate that such peptide hits in previous studies should be re-examined and that peptide identification should be based on strict trypsin specificity.

Animals↗

Toxicogenomics in drug discovery and development: mechanistic analysis of compound/class-dependent effects using the DrugMatrix database.

A range of genomics technologies are increasingly becoming integrated with existing scientific disciplines to broaden and strengthen existing capabilities and open new avenues of research in drug discovery and development. Examples of these new research fields are proteomics, pharmacogenomics, metabolomics and toxicogenomics. Here we review the application of toxicogenomics to improve the evaluation of drug safety, mechanism of action and toxicity in the drug discovery and development process.

Databases, Genetic↗

DbClustal: rapid and reliable global multiple alignments of protein sequences detected by database searches.

DbClustal addresses the important problem of the automatic multiple alignment of the top scoring full-length sequences detected by a database homology search. By combining the advantages of both local and global alignment algorithms into a single system, DbClustal is able to provide accurate global alignments of highly divergent, complex sequence sets. Local alignment information is incorporated into a ClustalW global alignment in the form of a list of anchor points between pairs of sequences. The method is demonstrated using anchors supplied by the Blast post-processing program, Ballast. The rapidity and reliability of DbClustal have been demonstrated using the recently annotated Pyrococcus abyssi proteome where the number of alignments with totally misaligned sequences was reduced from 20% to <2%. A web site has been implemented proposing BlastP database searches with automatic alignment of the top hits by DbClustal.

Algorithms↗

Gene expression profiling analysis in nephrology: towards molecular definition of renal disease.

The increase in progressive kidney disease, resulting in a constantly rising prevalence of endstage renal disease (ESRD), urgently warrants the development of more effective strategies to diagnose, prevent, and intervene in renal disease. Histological information obtained by renal biopsies (RBx) is a cornerstone of the current management of kidney disease. Renal tissue can provide critical information on the disease process not available by nontissue-based approaches. However, insight gained by conventional histopathology remains limited and additional strategies to define renal disease on a molecular level are required. The sequencing of the human genome, together with recent advances in genome-wide profiling techniques, has provided the framework for a comprehensive analysis of renal disease-associated transcriptional programs. In this review, strategies to apply these technological advances towards the analysis of RBx will be described, with special emphasis on their potential impact on clinical management, but also on their inherent limitations. Finally, an outlook towards the emerging proteomic studies of renal disease will be given.

Antigens, CD20↗