Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Bioinformatics of large-scale protein interaction networks.

We survey recent techniques for construction and prediction of large-scale protein interaction networks, focusing on computational processing steps. Special emphasis is placed on critical assessment of data completeness and reliability of the various approaches. Once built, protein interaction networks can be used for functional annotation or to generate higher-level biological hypotheses on pathways.

Bacterial Proteins↗

[Protein of Escherichia coli interacting specifically with human low density lipoproteins].

Escherichia coli 48 kDa protein interacting specifically with human low-density lipoproteins is described. The dissociation constant of this highly specific interaction was found to be equal to 4 mkg LDL per 1 ml or 7.3 x 10 M, which is comparable with the dissociation constant of the complex formed by LDL and human LDL receptor. A protocol for purifying the E. Coli binding protein was developed and antibodies against this purified protein were raised. The absence of sequences with homology to the ligand-binding repeats of the human LDL receptor in E. Coli proteome was shown by computer analysis of E. Coli genome. A conclusion was made that binding of the human LDL with specific E. Coli protein is thus mediated by other sequences and by another mechanism different from that, which occurs in human cells during the interaction of lipoproteins with their specific receptor. The establishment of specific interaction between E. Coli protein and human LDL can turn out to be useful in the future for purifying lipoproteins of a specific class and for administering plasmapheresis in patients with severe hyperlipoproteinemia.

Animals↗

Challenges for the identification of biological systems from in vivo time series data.

Modern methods of high-throughput molecular biology render it possible to generate time series of metabolite concentrations and the expression of genes and proteins in vivo. These time profiles contain valuable information about the structure and dynamics of the underlying biological system. This information is implicit and its extraction is a challenging but ultimately very rewarding task for the mathematical modeler. Using a well-suited modeling framework, such as Biochemical Systems Theory (BST), it is possible to formulate the extraction of information as an inverse problem that in principle may be solved with a genetic algorithm or nonlinear regression. However, two types of issues associated with this inverse problem make the extraction task difficult. One type pertains to the algorithmic difficulties encountered in nonlinear regressions with moderate and large systems. The other type is of an entirely different nature. It is a consequence of assumptions that are often taken for granted in the design and analysis of mathematical models of biological systems and that need to be revisited in the context of inverse analyses. The article describes the extraction process and some of its challenges and proposes partial solutions.

Linear Models↗

Proteomics combined with single-cell sequencing reveals key genes and computational lead compound related to ligamentum flavum hypertrophy, lactate metabolism and lactate modification.

Ligamentum flavum hypertrophy (LFH) is a hallmark pathological feature of lumbar spinal stenosis; however, its underlying molecular mechanisms remain incompletely understood. Lactate metabolism and related lactylation modifications have emerged as critical links between cellular metabolism and epigenetic regulation, with established roles in various fibrotic and inflammatory diseases. Nevertheless, the specific contribution of lactylation to LFH pathogenesis remains unexplored. In this study, we integrated proteomic profiling of ligamentum flavum tissues with single-cell transcriptomic data to identify differentially expressed proteins associated with LFH. Cross-referencing these genes with genes involved in lactate metabolism and lactylation yielded 16 candidate genes. Through functional enrichment analysis, protein-protein interaction network construction, and GraphBAN model prediction, we identified five hub genes (NDUFS2, HMOX1, SPR, FABP5, and PFKP) and two potential lead compounds (ZINC000014879975 and ZINC000242437513). Molecular docking analysis confirmed favorable binding affinities between these compounds, suggesting that they may serve as potential lead compounds worthy of further experimental investigation. Single-cell analysis further revealed that macrophages occupy a central position in the LFH microenvironment, resulting in pronounced metabolic reprogramming and remodeling of intercellular communication networks, particularly via the MIF-CD74/CD44 axis, under pathological conditions.

Proteomics↗

Functional genomics and proteomics in the clinical neurosciences: data mining and bioinformatics.

The goal of this chapter is to introduce some of the available computational methods for expression analysis. Genomic and proteomic experimental techniques are briefly discussed to help the reader understand these methods and results better in context with the biological significance. Furthermore, a case study is presented that will illustrate the use of these analytical methods to extract significant biomarkers from high-throughput microarray data. Genomic and proteomic data analysis is essential for understanding the underlying factors that are involved in human disease. Currently, such experimental data are generally obtained by high-throughput microarray or mass spectrometry technologies among others. The sheer amount of raw data obtained using these methods warrants specialized computational methods for data analysis. Biomarker discovery for neurological diagnosis and prognosis is one such example. By extracting significant genomic and proteomic biomarkers in controlled experiments, we come closer to understanding how biological mechanisms contribute to neural degenerative diseases such as Alzheimers' and how drug treatments interact with the nervous system. In the biomarker discovery process, there are several computational methods that must be carefully considered to accurately analyze genomic or proteomic data. These methods include quality control, clustering, classification, feature ranking, and validation. Data quality control and normalization methods reduce technical variability and ensure that discovered biomarkers are statistically significant. Preprocessing steps must be carefully selected since they may adversely affect the results of the following expression analysis steps, which generally fall into two categories: unsupervised and supervised. Unsupervised or clustering methods can be used to group similar genomic or proteomic profiles and therefore can elucidate relationships within sample groups. These methods can also assign biomarkers to sub-groups based on their expression profiles across patient samples. Although clustering is useful for exploratory analysis, it is limited due to its inability to incorporate expert knowledge. On the other hand, classification and feature ranking are supervised, knowledge-based machine learning methods that estimate the distribution of biological expression data and, in doing so, can extract important information about these experiments. Classification is closely coupled with feature ranking, which is essentially a data reduction method that uses classification error estimation or other statistical tests to score features. Biomarkers can subsequently be extracted by eliminating insignificantly ranked features. These analytical methods may be equally applied to genetic and proteomic data. However, because of both biological differences between the data sources and technical differences between the experimental methods used to obtain these data, it is important to have a firm understanding of the data sources and experimental methods. At the same time, regardless of the data quality, it is inevitable that some discovered biomarkers are false positives. Thus, it is important to validate discovered biomarkers. The validation process may be slow; yet, the overall biomarker discovery process is significantly accelerated due to initial feature ranking and data reduction steps. Information obtained from the validation process may also be used to refine data analysis procedures for future iteration. Biomarker validation may be performed in a number of ways - bench-side in traditional labs, web-based electronic resources such as gene ontology and literature databases, and clinical trials.

Animals↗

Web-accessible proteome databases for microbial research.

The analysis of proteomes of biological organisms represents a major challenge of the post-genome era. Classical proteomics combines two-dimensional electrophoresis (2-DE) and mass spectrometry (MS) for the identification of proteins. Novel technologies such as isotope coded affinity tag (ICAT)-liquid chromatography/mass spectrometry (LC/MS) open new insights into protein alterations. The vast amount and diverse types of proteomic data require adequate web-accessible computational and database technologies for storage, integration, dissemination, analysis and visualization. A proteome database system (http://www.mpiib-berlin.mpg.de/2D-PAGE) for microbial research has been constructed which integrates 2-DE/MS, ICAT-LC/MS and functional classification data of proteins with genomic, metabolic and other biological knowledge sources. The two-dimensional polyacrylamide gel electrophoresis database delivers experimental data on microbial proteins including mass spectra for the validation of protein identification. The ICAT-LC/MS database comprises experimental data for protein alterations of mycobacterial strains BCG vs. H37Rv. By formulating complex queries within a functional protein classification database "FUNC_CLASS" for Mycobacterium tuberculosis and Helicobacter pylori the researcher can gather precise information on genes, proteins, protein classes and metabolic pathways. The use of the R language in the database architecture allows high-level data analysis and visualization to be performed "on-the-fly". The database system is centrally administrated, and investigators without specific bioinformatic competence in database construction can submit their data. The database system also serves as a template for a prototype of a European Proteome Database of Pathogenic Bacteria. Currently, the database system includes proteome information for six strains of microorganisms.

Bacterial Proteins↗

Toward computer-based cleavage site prediction of cysteine endopeptidases.

Identification of relevant substrates is essential for elucidation of in vivo functions of peptidases. The recent availability of the complete genome sequences of many eukaryotic organisms holds the promise of identifying specific peptidase substrates by systematic proteome analyses in combination with computer-based screening of genome databases. Currently available proteomics and bioinformatics tools are not sufficient for reliable endopeptidase substrate predictions. To address these shortcomings the bioinformatics tool 'PEPS' (Prediction of Endopeptidase Substrates) has been developed and is presented here. PEPS uses individual rule-based endopeptidase cleavage site scoring matrices (CSSM). The efficiency of PEPS in predicting putative caspase 3, cathepsin B and cathepsin L cleavage sites is demonstrated in comparison to established algorithms. Mortalin, a member of the heat shock protein family HSP70, was identified by PEPS as a putative cathepsin L substrate. Comparative proteome analyses of cathepsin L-deficient and wild-type mouse fibroblasts showed that mortalin is enriched in the absence of cathepsin L. These results indicate that CSSM/PEPS can correctly predict relevant peptidase substrates.

Animals↗

Prediction of disulfide-bonded cysteines in proteomes with a hidden neural network.

A hidden neural network-based method is used to predict the bonding state of cysteines starting from the residue sequence of the protein chain. The method scores as high as 89% and 86% per cysteine residue and per protein, respectively, and in this overcomes other predictors of the same category. We then explore the efficacy of our predictor in computing the disulfide content of the whole proteome of Escherichia coli (K12 and O157), Aeropirum pernix, Thermotoga maritima, and Homo sapiens. We find that the percentage of extracellular disulfide containing proteins is higher than that of intracellular one, and that the human proteome is by far the one with the highest content of sulfur-sulfur linkages in proteins.

Cysteine↗

MPAC: a computational framework for inferring pathway activities from multi-omic data.

MOTIVATION: Fully capturing cellular state requires examining genomic, epigenomic, transcriptomic, proteomic, and other assays for a biological sample and comprehensive computational modeling to reason with the complex and sometimes conflicting measurements. Modeling these so-called multi-omic data is especially beneficial in disease analysis, where observations across omic data types may reveal unexpected patient groupings and inform clinical outcomes and treatments. RESULTS: We present Multi-omic Pathway Analysis of Cells (MPAC), a computational framework that interprets multi-omic data through prior knowledge from biological pathways. MPAC leverages network relationships encoded in pathways through a factor graph to infer consensus activity levels for proteins and associated pathway entities from multi-omic data, runs permutation testing to eliminate spurious activity predictions, and groups biological samples by pathway activities to allow identifying and prioritizing proteins with potential clinical relevance, e.g. associated with patient prognosis. Using DNA copy number alteration and RNA-seq data from head and neck squamous cell carcinoma patients from The Cancer Genome Atlas as an example, we demonstrate that MPAC predicts a patient subgroup related to immune responses not identified by analysis with either input omic data type alone. Key proteins identified via this subgroup have pathway activities related to clinical outcome as well as immune cell composition. Our MPAC R package enables similar multi-omic analyses on new datasets. AVAILABILITY AND IMPLEMENTATION: The MPAC package is available at Bioconductor https://bioconductor.org/packages/MPAC.

Humans↗

A statistical framework to discover true associations from multiprotein complex pull-down proteomics data sets.

Experimental processes to collect and process proteomics data are increasingly complex, and the computational methods to assess the quality and significance of these data remain unsophisticated. These challenges have led to many biological oversights and computational misconceptions. We developed an empirical Bayes model to analyze multiprotein complex (MPC) proteomics data derived from peptide mass spectrometry detections of purified protein complex pull-down experiments. Using our model and two yeast proteomics data sets, we estimated that there should be an average of about 20 true associations per MPC, almost 10 times as high as was previously estimated. For data sets generated to mimic a real proteome, our model achieved on average 80% sensitivity in detecting true associations, as compared with the 3% sensitivity in previous work, while maintaining a comparable false discovery rate of 0.3%. Cross-examination of our results with protein complexes confirmed by various experimental techniques demonstrates that many true associations that cannot be identified by previous approach are identified by our method.

Algorithms↗

Identifying cytotoxic T cell epitopes from genomic and proteomic information: "The human MHC project.".

Complete genomes of many species including pathogenic microorganisms are rapidly becoming available and with them the encoded proteins, or proteomes. Proteomes are extremely diverse and constitute unique imprints of the originating organisms allowing positive identification and accurate discrimination, even at the peptide level. It is not surprising that peptides are key targets of the immune system. It follows that proteomes can be translated into immunogens once it is known how the immune system generates and handles peptides. Recent advances have identified many of the basic principles involved. The single most selective event is that of peptide binding to MHC, making it particularly important to establish accurate descriptions and predictions of peptide binding for the most common MHC variants. These predictions should be integrated with those of other steps involved in antigen processing, as these become available. The ability to translate the accumulating primary sequence databases in terms of immune recognition should enable scientists and clinicians to analyze any protein of interest for the presence of potentially immunogenic epitopes. The computational tools to scan entire proteomes should also be developed, as this would enable a rational approach to vaccine development and immunotherapy. Thus, candidate vaccine epitopes might be predicted from the various microbial genome projects, tumor vaccine candidates from mRNA expression profiling of tumors ("transcriptomes") and auto-antigens from the human genome.

Antigen Presentation↗

Feature-based reappraisal of the Bacillus subtilis exoproteome.

Proteomics-based verification of computer-assisted predictions on bacterial protein export have indicated that problems occur with the distinction between (Sec-type) signal peptides that govern protein secretion, and lipoprotein signal peptides or amino-terminal membrane anchors that cause protein retention in the membrane. Therefore, the main aim of this study was to investigate whether feature-based predictions by the SecretomeP (SecP) algorithm will aid the proteomics-based analysis of protein export in Bacillus subtilis. The SecP algorithm is trained to recognize features such as secondary structure and disordered regions, which are generally present in secreted proteins. The results showed that membrane-retained proteins receive, in general, high SecP scores, similar to the scores of secretory proteins. Importantly, the SecP algorithm aided in the re-evaluation of a class of previously identified proteins that remain attached to the membrane despite the presence of an apparent Sec-type signal peptide. These so-called 'Sec-attached' proteins receive on average a lower SecP score, and several of these proteins could be unmasked as transmembrane proteins by combined SecP and signal peptide analyses. Finally, the present study suggests that feature-based outlier analysis may provide leads towards the discovery of novel special-purpose pathways for bacterial protein export.

Algorithms↗

Proteomic data exchange and storage: using Proteios.

Proteios (http://www.proteios.org) is an initiative for the development of a comprehensive open source system for storage, organization, analysis, and annotation of proteomics experiments. The Proteios platform is based on existing principles for proteomics data publishing and data exchange.

Computational Biology↗

On graphical and numerical characterization of proteomics maps.

We outlined a mathematical approach suitable for characterization of experimental data given by 2-D densitograms. In particular we consider numerical characterization of proteomics maps. The basis of our approach is to order "spots" of a 2-D map and assign them unique labels (that in general will depend on the criteria used for ordering). In this way a map is "translated" into a sequence. In the next step one associates with the generated sequence a geometrical path and views such a path as a mathematical object that needs characterization. We have ordered spots representing proteins in 2-D gel plates according to their relative intensities which results in a zigzag path that produces a complicated "fingerprint" pattern. Mathematical characterization of zigzag pattern follows similar mathematical characterizations of embedded patterns based on matrices, the elements of which are given as quotients of Euclidean distance between spots and the distance along the zigzag path. The leading eigenvalue of constructed matrices is taken to represent characterization of the original 2-D map. Comparison of different 2-D maps (simulated by using random generator) allows one to construct partial order, which although qualitative in nature gives some insight into perturbation induced by foreign agents to the proteome of the control cell.

Computer Graphics↗

Comparative proteome analysis of Brucella melitensis vaccine strain Rev 1 and a virulent strain, 16M.

The genus Brucella consists of bacterial pathogens that cause brucellosis, a major zoonotic disease characterized by undulant fever and neurological disorders in humans. Among the different Brucella species, Brucella melitensis is considered the most virulent. Despite successful use in animals, the vaccine strains remain infectious for humans. To understand the mechanism of virulence in B. melitensis, the proteome of vaccine strain Rev 1 was analyzed by two-dimensional gel electrophoresis and compared to that of virulent strain 16M. The two strains were grown under identical laboratory conditions. Computer-assisted analysis of the two B. melitensis proteomes revealed proteins expressed in either 16M or Rev 1, as well as up- or down-regulation of proteins specific for each of these strains. These proteins were identified by peptide mass fingerprinting. It was found that certain metabolic pathways may be deregulated in Rev 1. Expression of an immunogenic 31-kDa outer membrane protein, proteins utilized for iron acquisition, and those that play a role in sugar binding, lipid degradation, and amino acid binding was altered in Rev 1.

Bacterial Proteins↗