Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Proteomic informatics: in silico methods lead to data management challenges.

Proteomics, which identifies proteins and analyzes their function in cells, is foreseen as the next challenge in biomedicine as diseases in the body are most easily recognized through the function of their proteins. Achieving this recognition is more difficult than pure gene analysis: it is estimated that 35,000 genes are present in human DNA, encoding more than 1 million proteins. A myriad of in vitro and in silico technologies now exist for studying proteins and their biological function. This review focuses on the vast array of in silico proteomic analysis methods, highlights public and commercial repositories for this data, and discusses the challenges associated with resolving and integrating the knowledge originating from this data.

Amino Acid Sequence↗

A reference map and identification of porcine testis proteins using 2-DE and MS.

The development of the testis is essential for maturation of male mammals. A complete understanding of proteins expressed in the testis will provide biological information on many reproductive dysfunctions in males. The purposes of this study were to apply a proteomic approach to investigating protein composition and to establish a 2-D PAGE reference map for porcine testis proteins. MALDI-TOF MS was performed for protein identification. When 1 mg of total proteins was assayed by 2-D PAGE and stained with colloidal CBB, more than 400 proteins with a pI of pH 3-10 and M(r) of 10-200 kDa could be detected. Protein expression varied among individuals, with CV between 4.7 and 131.5%. A total of 447 protein spots were excised for identification, among which 337 spots were identified by searching the mass spectra against the NCBInr database. Identification of the remaining 110 spots was unsuccessful. A 2-D PAGE-based porcine testis protein database has been constructed on the basis of the results and will be published on the WWW. This database should be valuable for investigating the developmental biology and pathology of porcine testis.

Animals↗

Proteome analysis of liver cells expressing a full-length hepatitis C virus (HCV) replicon and biopsy specimens of posttransplantation liver from HCV-infected patients.

The development of a reproducible model system for the study of hepatitis C virus (HCV) infection has the potential to significantly enhance the study of virus-host interactions and provide future direction for modeling the pathogenesis of HCV. While there are studies describing global gene expression changes associated with HCV infection, changes in the proteome have not been characterized. We report the first large-scale proteome analysis of the highly permissive Huh-7.5 cell line containing a full-length HCV replicon. We detected >4,200 proteins in this cell line, including HCV replicon proteins, using multidimensional liquid chromatographic (LC) separations coupled to mass spectrometry. Consistent with the literature, a comparison of HCV replicon-positive and -negative Huh-7.5 cells identified expression changes of proteins involved in lipid metabolism. We extended these analyses to liver biopsy material from HCV-infected patients where a total of >1,500 proteins were detected from only 2 mug of liver biopsy protein digest using the Huh-7.5 protein database and the accurate mass and time tag strategy. These findings demonstrate the utility of multidimensional proteome analysis of the HCV replicon model system for assisting in the determination of proteins/pathways affected by HCV infection. Our ability to extend these analyses to the highly complex proteome of small liver biopsies with limiting protein yields offers the unique opportunity to begin evaluating the clinical significance of protein expression changes associated with HCV infection.

Amino Acid Sequence↗

Prediction of disulfide-bonded cysteines in proteomes with a hidden neural network.

A hidden neural network-based method is used to predict the bonding state of cysteines starting from the residue sequence of the protein chain. The method scores as high as 89% and 86% per cysteine residue and per protein, respectively, and in this overcomes other predictors of the same category. We then explore the efficacy of our predictor in computing the disulfide content of the whole proteome of Escherichia coli (K12 and O157), Aeropirum pernix, Thermotoga maritima, and Homo sapiens. We find that the percentage of extracellular disulfide containing proteins is higher than that of intracellular one, and that the human proteome is by far the one with the highest content of sulfur-sulfur linkages in proteins.

Cysteine↗

OrthoMCL-DB: querying a comprehensive multi-species collection of ortholog groups.

The OrthoMCL database (http://orthomcl.cbil.upenn.edu) houses ortholog group predictions for 55 species, including 16 bacterial and 4 archaeal genomes representing phylogenetically diverse lineages, and most currently available complete eukaryotic genomes: 24 unikonts (12 animals, 9 fungi, microsporidium, Dictyostelium, Entamoeba), 4 plants/algae and 7 apicomplexan parasites. OrthoMCL software was used to cluster proteins based on sequence similarity, using an all-against-all BLAST search of each species' proteome, followed by normalization of inter-species differences, and Markov clustering. A total of 511,797 proteins (81.6% of the total dataset) were clustered into 70,388 ortholog groups. The ortholog database may be queried based on protein or group accession numbers, keyword descriptions or BLAST similarity. Ortholog groups exhibiting specific phyletic patterns may also be identified, using either a graphical interface or a text-based Phyletic Pattern Expression grammar. Information for ortholog groups includes the phyletic profile, the list of member proteins and a multiple sequence alignment, a statistical summary and graphical view of similarities, and a graphical representation of domain architecture. OrthoMCL software, the entire FASTA dataset employed and clustering results are available for download. OrthoMCL-DB provides a centralized warehouse for orthology prediction among multiple species, and will be updated and expanded as additional genome sequence data become available.

Animals↗

Functional proteomics: The goalposts are moving.

Holistic understanding of protein function is a primary goal of the post-genome sequencing era. Functional genomic approaches are powerful and relatively straightforward but produce an incomplete picture at the protein level. Proteomics offers physiologically enriched insights to protein function, and ongoing advances are enabling proteome analyses to proceed with increased depth and efficiency. Exciting discoveries have emerged recently amidst growing awareness of the power of proteomics. However, while proven as a potent discovery tool, proteomics is under pressure to provide improved functional value particularly in concert with other investigative approaches. As reviewed here for ERp29, a recently discovered endoplasmic reticulum protein, the role of novel proteins can remain elusive even after substantial information has accrued. Thousands more proteins of uncertain function will be unveiled in the near future. Consequently, the goalposts are moving for proteomics both through increasing demand for high-value functional information and improving capacity to deliver.

Animals↗

MS1, MS2, and SQT-three unified, compact, and easily parsed file formats for the storage of shotgun proteomic spectra and identifications.

As the speed with which proteomic labs generate data increases along with the scale of projects they are undertaking, the resulting data storage and data processing problems will continue to challenge computational resources. This is especially true for shotgun proteomic techniques that can generate tens of thousands of spectra per instrument each day. One design factor leading to many of these problems is caused by storing spectra and the database identifications for a given spectrum as individual files. While these problems can be addressed by storing all of the spectra and search results in large relational databases, the infrastructure to implement such a strategy can be beyond the means of academic labs. We report here a series of unified text file formats for storing spectral data (MS1 and MS2) and search results (SQT) that are compact, easily parsed by both machine and humans, and yet flexible enough to be coupled with new algorithms and data-mining strategies.

Database Management Systems↗

[Introduction of proteomic approach to environmental medicine].

Recent progress in life science technology and the availability of much information on genes obtained by genome analysis has enabled us to analyze the changes of proteins on a large scale. Sets of proteins are called proteomes, and proteomics is the scientific field of proteome analysis including differential, post translational modification and interaction analyses. Various proteomic techniques, particularly two-dimensional gel electrophoresis (2-DE), mass spectrometry, protein chip methods, and surface plasmon resonance (SPR), are very useful for acquiring proteomes in cells, tissues and body fluid, and for analyzing interactions between a protein and other biofactors including proteins. A proteomic approach is also useful for determining biomarkers of diseases and key proteins involved in various stages of metabolism such as differentiation, cell cycle and apoptosis. Environmental pollutants including endocrine disruptors inhibit activities of various organs in wild animals and humans. Proteomic approaches could be very useful tools for elucidating the mechanisms of damage caused by environmental pollutants. In this review, we describe the application of a proteomic approach to the field of environmental medicine.

Animals↗

A collection of amino acid replacement matrices derived from clusters of orthologs.

Sequence divergence among orthologous proteins was characterized with 34 amino acid replacement matrices, sequence context analysis, and a phylogenetic tree. The model was trained on very large datasets of aligned protein sequences drawn from 15 organisms including protists, plants, Dictyostelium, fungi, and animals. Comparative tests with models currently used in phylogeny, i.e., with JTT+gamma+/-F and WAG+gamma+/-F, made on a test dataset of 380 multiple alignments containing protein sequences from all five of the major taxonomic groups mentioned, indicate that our model should be preferred over the JTT+gamma+/-F and WAG+gamma+/-F models on datasets similar to the test dataset. The strong performance of our model of orthologous protein sequence divergence can be attributed to its ability to better approximate amino acid equilibrium frequencies to compositions found in alignment columns.

Amino Acid Substitution↗

Integrated analysis of multiple data sources reveals modular structure of biological networks.

It has been a challenging task to integrate high-throughput data into investigations of the systematic and dynamic organization of biological networks. Here, we presented a simple hierarchical clustering algorithm that goes a long way to achieve this aim. Our method effectively reveals the modular structure of the yeast protein-protein interaction network and distinguishes protein complexes from functional modules by integrating high-throughput protein-protein interaction data with the added subcellular localization and expression profile data. Furthermore, we take advantage of the detected modules to provide a reliably functional context for the uncharacterized components within modules. On the other hand, the integration of various protein-protein association information makes our method robust to false-positives, especially for derived protein complexes. More importantly, this simple method can be extended naturally to other types of data fusion and provides a framework for the study of more comprehensive properties of the biological network and other forms of complex networks.

Algorithms↗

Polyethylene transformation by a psychrotolerant Rhodococcus strain assessed by transcriptomics and 13C-isotope tracing.

Polyethylene is increasingly accumulating in nature, including remote places like the Arctic. While abiotic processes fragment polyethylene in situ, biotic transformation by microorganisms is assumed to occur. However, the enzymes and pathways involved remain poorly characterized. In this study, we used an in-house biobank from cold environments to screen for potential bacteria capable of degrading polyethylene by screening the strains in silico using the database PlasticDB and in vivo using a fluorescence-based assay. Using transcriptomic and proteomic analyses to identify genes in promising candidate strains that encode extracellular enzymes potentially capable of degrading PE, we selected a Rhodococcus erythropolis strain and two of its enzymes: a hypothetical protein (Hypr1) and a lipase family protein (Lip2). Expressing the candidate genes heterologously in Escherichia coli resulted in positive results in the fluorescence-based assay for polyethylene transformation. Applying 13C-labelled polyethylene for assessing and estimating polyethylene transformation and carbon assimilation, we found that R. erythropolis and both untransformed and recombinant E. coli extracellularly transformed the initially added polyethylene after 70 days. In addition, untransformed E. coli and R. erythropolis converted small, but significant amounts of polyethylene-derived carbon to carbon dioxide. The 13C-label was also traced into the bacterial biomass of R. erythropolis. Overall, our results provide evidence for biotic transformation of untreated polyethylene and suggests a hypothetical protein and a lipase family protein as two novel enzyme candidates associated with PE transformation.

Rhodococcus↗

NovoBoard: A Comprehensive Framework for Evaluating the False Discovery Rate and Accuracy of De Novo Peptide Sequencing.

De novo peptide sequencing is one of the most fundamental research areas in mass spectrometry-based proteomics. Many methods have often been evaluated using a couple of simple metrics that do not fully reflect their overall performance. Moreover, there has not been an established method to estimate the false discovery rate (FDR) of de novo peptide-spectrum matches. Here we propose NovoBoard, a comprehensive framework to evaluate the performance of de novo peptide-sequencing methods. The framework consists of diverse benchmark datasets (including tryptic, nontryptic, immunopeptidomics, and different species) and a standard set of accuracy metrics to evaluate the fragment ions, amino acids, and peptides of the de novo results. More importantly, a new approach is designed to evaluate de novo peptide-sequencing methods on target-decoy spectra and to estimate and validate their FDRs. Our FDR estimation provides valuable information to assess the reliability of new peptides identified by de novo sequencing tools, especially when no ground-truth information is available to evaluate their accuracy. The FDR estimation can also be used to evaluate the capability of de novo peptide sequencing tools to distinguish between de novo peptide-spectrum matches and random matches. Our results thoroughly reveal the strengths and weaknesses of different de novo peptide-sequencing methods and how their performances depend on specific applications and the types of data.

Peptides↗

Omics in hereditary optic neuropathies: A systematic review of clinical studies with an integrated point of view.

Hereditary optic neuropathies are characterized by bilateral visual loss due to the degeneration of retinal ganglion cells, resulting in optic nerve degeneration and atrophy. Although the genetic origin of the main isolated and syndromic hereditary optic neuropathies has been characterized, the clinical phenotypes exhibit significant and poorly understood variability in both penetrance and expressivity. Additionally, the genetic and environmental factors that influence the onset of these optic neuropathies remain poorly understood, with limited biomarkers to predict disease progression or as readouts for therapeutic trials. Data-driven omics strategies allow deep phenotyping to improve our understanding of pathophysiological mechanisms and to search for new biomarkers and therapeutic targets. We explore whether the omics strategies applied to patients with hereditary optic neuropathies have provided such new insights. MEDLINE, Web of Science and EMBASE databases were screened for studies with terms relating to hereditary optic neuropathies, transcriptomics, epigenomics, proteomics, metabolomics and lipidomics in clinical studies exploring patients' samples. Out of 1244 references identified, 22 articles were included after double-masked data curation. These articles focused only on the 3 main forms of hereditary optic neuropathies, namely, OPA1-related dominant optic atrophy (n = 4), Leber hereditary optic neuropathy (n = 13), and Wolfram syndrome (n = 5). While the methodological designs and results of these studies were highly heterogeneous, they revealed molecular alterations that we have attempted to discuss at the integrated multi-omics level. This data integration highlighted several common pathophysiological mechanisms such as energetic impairment, endoplasmic reticulum stress, proteotoxic and oxidative stresses, lipid remodeling and altered amino acid and purine metabolisms, while suggesting potential new biomarkers and therapeutic targets. These findings underscore the potential of integrated multi-omics approaches to deepen our understanding of the phenotypic complexity of hereditary optic neuropathies and to support the development of innovative diagnostic and therapeutic strategies.

Humans↗

Data management solutions for protein therapeutic research and development.

Protein therapeutics, including monoclonal antibodies, are a growing focus of drug discovery research organizations. High-throughput screening of large libraries of protein variants is therefore becoming increasingly important in R&D. As a result, there is a need to link large numbers of variant protein sequences with chemical and biological assay data. This integration will allow more efficient data mining and facilitate decision-making regarding hit identification, lead optimization and drug development. In this paper, we present an implementation in which a widely used small-molecule high-throughput screening data management system has been adapted to meet the unique needs of protein drug discovery and development.

Antibodies, Monoclonal↗

Prediction of enzyme family classes.

Classes of newly found enzyme sequences are usually determined either by biochemical analysis of eukaryotic and prokaryotic genomes or by microarray chips. These experimental methods are both time-consuming and costly. With the explosion of protein sequences entering into databanks, it is highly desirable to explore the feasibility of selectively classifying newly found enzyme sequences into their respective enzyme classes by means of an automated method. This is indeed important because knowing which family or subfamily an enzyme belongs to may help deduce its catalytic mechanism and specificity, giving clues to the relevant biological function. In this study, a bioinformatical analysis was conducted for 2640 oxidoreductases classified into 16 subclasses according to the different types of substrates they act on during the catalytic process. Although it is an extremely complicated problem and might involve the knowledge of 3-dimensional structure as well as many other physical chemistry factors, some quite promising results have been obtained indicating that the family or subfamily of an enzyme is predictable to a considerable degree by means of sequence-based approach alone if a good training dataset can be established.

Amino Acid Sequence↗

Predicting eukaryotic protein subcellular location by fusing optimized evidence-theoretic K-Nearest Neighbor classifiers.

Facing the explosion of newly generated protein sequences in the post genomic era, we are challenged to develop an automated method for fast and reliably annotating their subcellular locations. Knowledge of subcellular locations of proteins can provide useful hints for revealing their functions and understanding how they interact with each other in cellular networking. Unfortunately, it is both expensive and time-consuming to determine the localization of an uncharacterized protein in a living cell purely based on experiments. To tackle the challenge, a novel hybridization classifier was developed by fusing many basic individual classifiers through a voting system. The "engine" of these basic classifiers was operated by the OET-KNN (Optimized Evidence-Theoretic K-Nearest Neighbor) rule. As a demonstration, predictions were performed with the fusion classifier for proteins among the following 16 localizations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cyanelle, (5) cytoplasm, (6) cytoskeleton, (7) endoplasmic reticulum, (8) extracell, (9) Golgi apparatus, (10) lysosome, (11) mitochondria, (12) nucleus, (13) peroxisome, (14) plasma membrane, (15) plastid, and (16) vacuole. To get rid of redundancy and homology bias, none of the proteins investigated here had >/=25% sequence identity to any other in a same subcellular location. The overall success rates thus obtained via the jack-knife cross-validation test and independent dataset test were 81.6% and 83.7%, respectively, which were 46 approximately 63% higher than those performed by the other existing methods on the same benchmark datasets. Also, it is clearly elucidated that the overwhelmingly high success rates obtained by the fusion classifier is by no means a trivial utilization of the GO annotations as prone to be misinterpreted because there is a huge number of proteins with given accession numbers and the corresponding GO numbers, but their subcellular locations are still unknown, and that the percentage of proteins with GO annotations indicating their subcellular components is even less than the percentage of proteins with known subcellular location annotation in the Swiss-Prot database. It is anticipated that the powerful fusion classifier may also become a very useful high throughput tool in characterizing other attributes of proteins according to their sequences, such as enzyme class, membrane protein type, and nuclear receptor subfamily, among many others. A web server, called "Euk-OET-PLoc", has been designed at http://202.120.37.186/bioinf/euk-oet for public to predict subcellular locations of eukaryotic proteins by the fusion OET-KNN classifier.

Amino Acids↗