Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Malaria: therapy, genes and vaccines.

Malaria kills over 3,000 children each day. Modern molecular and biochemical approaches are being used to help understand and control Plasmodium falciparum, the parasite that causes this deadly disease. New drugs are being invented for both chemoprophylaxis and therapeutic treatments and their use is discussed along side that of the more commonly used treatments. Classical genetic crosses coupled with molecular analysis of gene loci are use to explain the genetics behind the development of specific drug resistances that the parasites have naturally developed. Rapid advances in DNA sequencing techniques have allowed the compete sequencing of the P. falciparum and several other rodent malaria parasite genomes. Proteomics and computational analysis of these vast databanks are being used to model and investigate the three-dimensional structure of many key malaria proteins in an attempt to facilitate drug design. Recombinant protein expression in bacteria and yeast coupled with cGMP purification technologies and conditions have lead to the recent availability of several dozen malaria protein antigens for human-use Phase I and Phase II vaccine trials. Drug companies, private foundations, and key government agencies have contributed to the coordinated efforts needed to test these antigens, adjuvants and delivery methods in an effort to find an effective malaria vaccine that will prevent infection and disease.

Animals↗

The power and the limitations of cross-species protein identification by mass spectrometry-driven sequence similarity searches.

Mass spectrometry-driven BLAST (MS BLAST) is a database search protocol for identifying unknown proteins by sequence similarity to homologous proteins available in a database. MS BLAST utilizes redundant, degenerate, and partially inaccurate peptide sequence data obtained by de novo interpretation of tandem mass spectra and has become a powerful tool in functional proteomic research. Using computational modeling, we evaluated the potential of MS BLAST for proteome-wide identification of unknown proteins. We determined how the success rate of protein identification depends on the full-length sequence identity between the queried protein and its closest homologue in a database. We also estimated phylogenetic distances between organisms under study and related reference organisms with completely sequenced genomes that allow substantial coverage of unknown proteomes.

Animals↗

Finding function: evaluation methods for functional genomic data.

BACKGROUND: Accurate evaluation of the quality of genomic or proteomic data and computational methods is vital to our ability to use them for formulating novel biological hypotheses and directing further experiments. There is currently no standard approach to evaluation in functional genomics. Our analysis of existing approaches shows that they are inconsistent and contain substantial functional biases that render the resulting evaluations misleading both quantitatively and qualitatively. These problems make it essentially impossible to compare computational methods or large-scale experimental datasets and also result in conclusions that generalize poorly in most biological applications. RESULTS: We reveal issues with current evaluation methods here and suggest new approaches to evaluation that facilitate accurate and representative characterization of genomic methods and data. Specifically, we describe a functional genomics gold standard based on curation by expert biologists and demonstrate its use as an effective means of evaluation of genomic approaches. Our evaluation framework and gold standard are freely available to the community through our website. CONCLUSION: Proper methods for evaluating genomic data and computational approaches will determine how much we, as a community, are able to learn from the wealth of available data. We propose one possible solution to this problem here but emphasize that this topic warrants broader community discussion.

Algorithms↗

Understanding the yeast proteome: a bioinformatics perspective.

Rapid development of genomic and proteomic methodologies has provided a wealth of data for deciphering the biomolecular circuitry of a living cell. The main areas of computational research of proteomes outlined in this review are: understanding the system, its features and parameters to help plan the experiments; data integration, to help produce more reliable data sets; visualization and other forms of data representation to simplify interpretation; modeling of the functional regulation; and systems biology. With false-positive rates reaching 50% even in the more reliable data sets, handling the experimental error remains one of the most challenging tasks. Integrative approaches, incorporating results of various genome- and proteome-wide experiments, allow for minimizing the error and bring with them significant predictive power.

Computational Biology↗

PCAS--a precomputed proteome annotation database resource.

BACKGROUND: Many model proteomes or "complete" sets of proteins of given organisms are now publicly available. Much effort has been invested in computational annotation of those "draft" proteomes. Motif or domain based algorithms play a pivotal role in functional classification of proteins. Employing most available computational algorithms, mainly motif or domain recognition algorithms, we set up to develop an online proteome annotation system with integrated proteome annotation data to complement existing resources. RESULTS: We report here the development of PCAS (ProteinCentric Annotation System) as an online resource of pre-computed proteome annotation data. We applied most available motif or domain databases and their analysis methods, including hmmpfam search of HMMs in Pfam, SMART and TIGRFAM, RPS-PSIBLAST search of PSSMs in CDD, pfscan of PROSITE patterns and profiles, as well as PSI-BLAST search of SUPERFAMILY PSSMs. In addition, signal peptide and TM are predicted using SignalP and TMHMM respectively. We mapped SUPERFAMILY and COGs to InterPro, so the motif or domain databases are integrated through InterPro. PCAS displays table summaries of pre-computed data and a graphical presentation of motifs or domains relative to the protein. As of now, PCAS contains human IPI, mouse IPI, and rat IPI, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae, and S. pombe proteome.PCAS is available at http://pak.cbi.pku.edu.cn/proteome/gca.php CONCLUSION: PCAS gives better annotation coverage for model proteomes by employing a wider collection of available algorithms. Besides presenting the most confident annotation data, PCAS also allows customized query so users can inspect statistically less significant boundary information as well. Therefore, besides providing general annotation information, PCAS could be used as a discovery platform. We plan to update PCAS twice a year. We will upgrade PCAS when new proteome annotation algorithms identified.

Algorithms↗

Informatics solutions for high-throughput proteomics.

The success of mass-spectrometry-based proteomics as a method for analyzing proteins in biological samples is accompanied by challenges owning to demands for increased throughput. These challenges arise from the vast volume of data generated by proteomics experiments combined with the heterogeneity in data formats, processing methods, software tools and databases that are involved in the translation of spectral data into relevant and actionable information for scientists. Informatics aims to provide answers to these challenges by transferring existing solutions from information management to proteomics and/or by generating novel computational methods for automation of proteomics data processing.

Databases, Protein↗

Induction of apoptosis in mouse liver by microcystin-LR: a combined transcriptomic, proteomic, and simulation strategy.

Microcystins (MCs) are a family of cyclic heptapeptide hepatotoxins produced by freshwater species of cyanobacteria that have been implicated in the development of liver cancer, necrosis, and even deadly intrahepatic bleeding. MC-LR, the most toxic MC variant, is also the most commonly encountered in a contaminated aquatic system. This study presents the first data in the toxicological research of MCs that combines the use of standard apoptotic assays with transcriptomics, proteomic technologies, and computer simulations. By using histochemistry, DNA fragmentation assays, and flow cytometry analysis, we determined that MC-LR causes rapid, dose-dependent apoptosis in mouse liver when BALB/c mice are treated with MC-LR for 24 h at doses of either 50, 60, or 70 microg/kg of body weight. We then used gene expression profiling to demonstrate differential expressions (>2-fold) of 61 apoptosis-related genes in cells treated with MC-LR. Further proteomic analysis identified a total of 383 proteins of which 35 proteins were up-regulated and 30 proteins were down-regulated more than 2.5-fold when compared with controls. Combining computer simulations with the transcriptomic and proteomic data, we found that low doses (50 microg/kg) of MC-LR lead to apoptosis primarily through the BID-BAX-BCL-2 pathway, whereas high doses of MC-LR (70 microg/kg) caused apoptosis via a reactive oxygen species pathway. These results indicated that MC-LR exposure can cause apoptosis in mouse liver and revealed two independent pathways playing a major regulatory role in MC-LR-induced apoptosis, thereby contributing to a better understanding of the hepatotoxicity and the tumor-promoting mechanisms of MCs.

Animals↗

Computational prediction of proteotypic peptides for quantitative proteomics.

Mass spectrometry-based quantitative proteomics has become an important component of biological and clinical research. Although such analyses typically assume that a protein's peptide fragments are observed with equal likelihood, only a few so-called 'proteotypic' peptides are repeatedly and consistently identified for any given protein present in a mixture. Using >600,000 peptide identifications generated by four proteomic platforms, we empirically identified >16,000 proteotypic peptides for 4,030 distinct yeast proteins. Characteristic physicochemical properties of these peptides were used to develop a computational tool that can predict proteotypic peptides for any protein from any organism, for a given platform, with >85% cumulative accuracy. Possible applications of proteotypic peptides include validation of protein identifications, absolute quantification of proteins, annotation of coding sequences in genomes, and characterization of the physical principles governing key elements of mass spectrometric workflows (e.g., digestion, chromatography, ionization and fragmentation).

Algorithms↗

ORCO: Ollivier-Ricci Curvature-Omics-an unsupervised method for analyzing robustness in biological systems.

MOTIVATION: Although recent advanced sequencing technologies have improved the resolution of genomic and proteomic data to better characterize molecular phenotypes, efficient computational tools to analyze and interpret large-scale omic data are still needed. RESULTS: To address this, we have developed a network-based bioinformatic tool called Ollivier-Ricci curvature for omics (ORCO). ORCO incorporates omics data and a network describing biological relationships between the genes or proteins and computes Ollivier-Ricci curvature (ORC) values for individual interactions. ORC is an edge-based measure that assesses network robustness. It captures functional cooperation in gene signaling using a consistent information-passing measure, which can help investigators identify therapeutic targets and key regulatory modules in biological systems. ORC has identified novel insights in multiple cancer types using genomic data and in neurodevelopmental disorders using brain imaging data. This tool is applicable to any data that can be represented as a network. AVAILABILITY AND IMPLEMENTATION: ORCO is an open-source Python package and is publicly available on GitHub at https://github.com/aksimhal/ORC-Omics.

Software↗

Identification of the most abundant secreted proteins from the salivary glands of the sand fly Lutzomyia longipalpis, vector of Leishmania chagasi.

Using massive cDNA sequencing, proteomics and customized computational biology approaches, we have isolated and identified the most abundant secreted proteins from the salivary glands of the sand fly Lutzomyia longipalpis. Out of 550 randomly isolated clones from a full-length salivary gland cDNA library, we found 143 clusters or families of related proteins. Out of these 143 families, 35 were predicted to be secreted proteins. We confirmed, by Edman degradation of Lu. longipalpis salivary proteins, the presence of 17 proteins from this group. Full-length sequence for 35 cDNA messages for secretory proteins is reported, including an RGD-containing peptide, three members of the yellow-related family of proteins, maxadilan, a PpSP15-related protein, six members of a family of putative anticoagulants, an antigen 5-related protein, a D7-related protein, a cDNA belonging to the Cimex apyrase family of proteins, a protein homologous to a silk protein with amino acid repeats resembling extracellular matrix proteins, a 5'-nucleotidase, a peptidase, a palmitoyl-hydrolase, an endonuclease, nine novel peptides and four different groups of proteins with no homologies to any protein deposited in accessible databases. Sixteen of these proteins appear to be unique to sand flies. With this approach, we have tripled the number of isolated secretory proteins from this sand fly. Because of the relationship between the vertebrate host immune response to salivary proteins and protection to parasite infection, these proteins are promising markers for vector exposure and attractive targets for vaccine development to control Leishmania chagasi infection.

Amino Acid Sequence↗

Clinical proteomics and mass spectrometry profiling for cancer detection.

A key challenge in the clinical proteomics of cancer is the identification of biomarkers that would enable early detection, diagnosis and monitoring of disease progression to improve long-term survival of patients. Recent advances in proteomic instrumentation and computational methodologies offer a unique chance to rapidly identify these new candidate markers or pattern of markers. The combination of retentate affinity chromatography and mass spectrometry is one of the most interesting new approaches for cancer diagnostics using proteomic profiling. This review presents two technologies in this field, surface-enhanced laser desorption/ionization time-of-flight and Clinprot, and aims to summarize the results of studies obtained with the first of them for the early diagnosis of human cancer. Despite promising results, the use of the proteomic profiling as a diagnostic tool brought some controversies and technical problems, and still requires some efforts to be standardized and validated.

Biomarkers, Tumor↗

Systems biology in the cell nucleus.

The mammalian nucleus is arguably the most complex cellular organelle. It houses the vast majority of an organism's genetic material and is the site of all major genome regulatory processes. Reductionist approaches have been spectacularly successful at dissecting at the molecular level many of the key processes that occur within the nucleus, particularly gene expression. At the same time, the limitations of analyzing single nuclear processes in spatial and temporal isolation and the validity of generalizing observations of single gene loci are becoming evident. The next level of understanding of genome function is to integrate our knowledge of their sequences and the molecular mechanisms involved in nuclear processes with our insights into the spatial and temporal organization of the nucleus and to elucidate the interplay between protein and gene networks in regulatory circuits. To do so, catalogues of genomes and proteomes as well as a precise understanding of the behavior of molecules in living cells are required. Converging technological developments in genomics, proteomics, dynamics and computation are now leading towards such an integrated biological understanding of genome biology and nuclear function.

Animals↗

[Clinical proteomics: towards early detection of cancers].

A key challenge in clinical proteomic of cancer is the identification of biomarkers that would allow early detection, diagnosis and monitor progression of the disease to improve long-term survival of patients. Recent advances in proteomic instrumentation and computational methodologies offer unique chance to rapidly identify these new candidate markers or pattern of markers. The combination of retentate affinity chromatography and surfaced-enhanced laser desorption/ionization time-of-flight (SELDI-TOF) mass spectrometry is one of the most interesting new approaches for cancer diagnostic using proteomic profiling. This review aims to summarize the results of studies that have used this new technology method for the early diagnosis of human cancer. Despite promising results, the use of the proteomic profiling as a diagnostic tool brought some controversies and technical problems and still requires some efforts to be standardised and validated.

Biomarkers, Tumor↗

Prediction of drug metabolism: the case of cytochrome P450 2D6.

Cytochromes P450 (Cyt P450s) constitute the most important biotransformation enzymes involved in the biotransformation of drugs and other xenobiotics. Because drug metabolism by Cyt P450s plays such an important role in the disposition and in the pharmacological and toxicological effects of drugs, early consideration of ADME-properties is increasingly seen as essential for the discovery and the development of new drugs and drug candidates. The primary aim of this paper is to present various computational approaches used to rationalize and predict the activity and substrate selectivity of Cyt P450s, as well as the possibilities and limitations of these approaches, now and in the future. Attention is also paid to the experimental validation of these approaches by using high-throughput screening (HTS) of affinities to drug-drug interactions at the level of Cyt P450-isoenzymes. Since human Cyt P450 2D6 is one of the most important drug metabolizing enzymes and since in this regard much pioneering work has been done with this Cyt P450, Cyt P450 2D6 is chosen as a model for this discussion. Apart from early mechanism-based ab initio calculations on substrates of Cyt P450 2D6, pharmacophore modeling of ligands (i.e. both substrates and inhibitors) of Cyt P450 2D6 and protein homology modeling have been used successfully for the rationalisation and prediction of metabolite formation by this Cyt P450 isoenzyme. Significant protein structure-related species differences have been reported recently. It is concluded that not one computational approach is capable of rationalizing and reliably predicting metabolite formation by Cyt P450 2D6, but that it is rather the combination of the various complimentary approaches. It is moreover concluded, that experimental validation of the computational models and predictions is often still lacking. With the advent of novel, easily and well applicable in vitro based high throughput assays for ligand binding and turnover this limitation could be overcome soon, however. When effective links with other new and recent developments, such as bioinformatics, neural network computing, genomics and proteomics can be created, in silico rationalisation and prediction of drug metabolism by Cyt P450s is likely to become one of the key technologies in early drug discovery and development processes.

Animals↗

Synthetic modular systems--reverse engineering of signal transduction.

During the last decades, biology has decomposed cellular systems into genetic, functional and molecular networks. It has become evident that these networks consist of components with specific functions (e.g., proteins and genes). This has generated a considerable amount of knowledge and hypotheses concerning cellular organization. The idea discussed here is to test the extent of this knowledge by reconstructing, or reverse engineering, new synthetic biological systems from known components. We will discuss how integration of computational methods with proteomics and engineering concepts might lead us to a deeper and more abstract understanding of signal transduction systems. Designing and successfully introducing synthetic proteins into cellular pathways would provide us with a powerful research tool with many applications, such as development of biosensors, protein drugs and rewiring of biological pathways.

Animals↗

Bioinformatics in mass spectrometry data analysis for proteomics studies.

Mass spectrometry is a technique widely employed for the identification and characterization of proteins. The role of bioinformatics is fundamental for the elaboration of mass spectrometry data due to the amount of data that this technique can produce. To process data efficiently, new software packages and algorithms are continuously being developed to improve protein identification and characterization in terms of high-throughput and statistical accuracy. However, many limitations exist concerning bioinformatics spectral data elaboration. This review aims to critically cover the recent and future developments of new bioinformatics approaches in mass spectrometry data analysis for proteomics studies.

Computational Biology↗