Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Protein structural class identification directly from NMR spectra using averaged chemical shifts.

Knowledge of the three-dimensional structure of proteins is integral to understanding their functions, and a necessity in the era of proteomics. A wide range of computational methods is employed to estimate the secondary, tertiary, and quaternary structures of proteins. Comprehensive experimental methods, on the other hand, are limited to nuclear magnetic resonance (NMR) and X-ray crystallography. The full characterization of individual structures, using either of these techniques, is extremely time intensive. The demands of high throughput proteomics necessitate the development of new, faster experimental methods for providing structural information. As a first step toward such a method, we explore the possibility of determining the structural classes of proteins directly from their NMR spectra, prior to resonance assignment, using averaged chemical shifts. This is achieved by correlating NMR-based information with empirical structure-based information available in widely used electronic databases. The results are analyzed statistically for their significance. The robustness of the method as a structure predictor is probed by applying it to a set of proteins of unknown structure. Our results show that this NMR-based method can be used as a low-resolution tool for protein structural class identification.

Algorithms↗

Discover true association rates in multi-protein complex proteomics data sets.

Experimental processes to collect and process proteomics data are increasingly complex, while the computational methods to assess the quality and significance of these data remain unsophisticated. These challenges have led to many biological oversights and computational misconceptions. We developed a complete empirical Bayes model to analyze multi-protein complex (MPC) proteomics data derived from peptide mass spectrometry detections of purified protein complex pull-down experiments. Our model considers not only bait-prey associations, but also prey-prey associations missed in previous work. Using our model and a yeast MPC proteomics data set, we estimated that there should be an average of 28 true associations per MPC, almost ten times as high as was previously estimated. For data sets generated to mimic a real proteome, our model achieved on average 80% sensitivity in detecting true associations, as compared with the 3% sensitivity in previous work, while maintaining a comparable false discovery rate of 0.3%.

Algorithms↗

ImmunoTar-integrative prioritization of cell surface targets for cancer immunotherapy.

MOTIVATION: Cancer remains a leading cause of mortality globally. Recent improvements in survival have been facilitated by the development of targeted and less toxic immunotherapies, such as chimeric antigen receptor (CAR)-T cells and antibody-drug conjugates (ADCs). These therapies, effective in treating both pediatric and adult patients with solid and hematological malignancies, rely on the identification of cancer-specific surface protein targets. While technologies like RNA sequencing and proteomics exist to survey these targets, identifying optimal targets for immunotherapies remains a challenge in the field. RESULTS: To address this challenge, we developed ImmunoTar, a novel computational tool designed to systematically prioritize candidate immunotherapeutic targets. ImmunoTar integrates user-provided RNA-sequencing or proteomics data with quantitative features from multiple public databases, selected based on predefined criteria, to generate a score representing the gene's suitability as an immunotherapeutic target. We validated ImmunoTar using three distinct cancer datasets, demonstrating its effectiveness in identifying both known and novel targets across various cancer phenotypes. By compiling diverse data into a unified platform, ImmunoTar enables comprehensive evaluation of surface proteins, streamlining target identification and empowering researchers to efficiently allocate resources, thereby accelerating the development of effective cancer immunotherapies. AVAILABILITY AND IMPLEMENTATION: Code and data to run and test ImmunoTar are available at https://github.com/sacanlab/immunotar.

Humans↗

Using standard positions and image fusion to create proteome maps from collections of two-dimensional gel electrophoresis images.

Databases for two-dimensional protein gels pose new challenges in extracting meaningful information from large numbers of experiments. In order to create expression profiles, positions of corresponding protein spots across all gel images have to be established. In larger gel sets errors may accumulate rapidly during this spot matching process, effectively limiting the number of samples available for data mining. Here we present a novel approach for organizing spot data based on the concept of a standard position for a protein species. Standard positions are meaningful average positions that are determined using all occurrences of a protein species. They can be extended to spots that are not annotated via interpolation. The standard position of a spot can serve as a unifying index across all gels in a database, thus allowing creation and analysis of expression profiles that span the whole collection. The standard position gives a much more accurate estimation of a spot's position on a gel than can be obtained using theoretical isoelectric point and molecular weight. Positional indexing is a complement to a priori identifications (e.g. by mass spectrometry or Edman degradation). Moreover it can be used in advance to select spots that are worth identifying because they show relevant expression profiles. Furthermore, we show how to combine all spots that occur on any of the gels into one synthetic but nevertheless realistic-looking image. This composite image is produced such that all spots have their standard positions. It can serve as a proteome reference map for an organism. As an application, we have computed a reference map from 23 gel images of Bacillus subtilis, using an enhanced prerelease version of the gel analysis software Delta2D (DECODON, Greifswald, Germany).

Bacterial Proteins↗

The secrets of a functional synapse--from a computational and experimental viewpoint.

BACKGROUND: Neuronal communication is tightly regulated in time and in space. The neuronal transmission takes place in the nerve terminal, at a specialized structure called the synapse. Following neuronal activation, an electrical signal triggers neurotransmitter (NT) release at the active zone. The process starts by the signal reaching the synapse followed by a fusion of the synaptic vesicle and diffusion of the released NT in the synaptic cleft; the NT then binds to the appropriate receptor, and as a result, a potential change at the target cell membrane is induced. The entire process lasts for only a fraction of a millisecond. An essential property of the synapse is its capacity to undergo biochemical and morphological changes, a phenomenon that is referred to as synaptic plasticity. RESULTS: In this survey, we consider the mammalian brain synapse as our model. We take a cell biological and a molecular perspective to present fundamental properties of the synapse:(i) the accurate and efficient delivery of organelles and material to and from the synapse; (ii) the coordination of gene expression that underlies a particular NT phenotype; (iii) the induction of local protein expression in a subset of stimulated synapses. We describe the computational facet and the formulation of the problem for each of these topics. CONCLUSION: Predicting the behavior of a synapse under changing conditions must incorporate genomics and proteomics information with new approaches in computational biology.

Animals↗

An improved prediction of chloroplast proteins reveals diversities and commonalities in the chloroplast proteomes of Arabidopsis and rice.

Proteins that form part of the chloroplast proteome can be identified by computational prediction of the N-terminal presequences (chloroplast transit peptides, cTPs) of their cytoplasmic precursor proteins. The accuracy of four different cTP predictors has been evaluated on a test set of 4500 proteins whose subcellular localization is known, and was found to be substantially lower than previously reported. A combination of cTP prediction programs was superior to any one of the predictors alone. This combination was employed to estimate the size and composition of the chloroplast proteomes of Arabidopsis and rice, and about 2100 (Arabidopsis thaliana) and 4800 (Oryza sativa) different chloroplast proteins with a cTP are predicted to be encoded by their nuclear genomes. A subset of around 900 chloroplast proteins, predominantly derived from the cyanobacterial endosymbiont and with functions mostly related to metabolism, energy and transcription, is shared by the two species. This points to the existence of both conserved nucleus-encoded chloroplast proteins that are predominantly of prokaryotic origin, and a large fraction of taxon-specific chloroplast-targeted proteins, in flowering plants.

Arabidopsis↗

Evolution of drug metabolism: hitchhiking the technology bandwagon.

1. The application of a range of established and emerging technologies and experimental approaches has allowed investigation of cytochrome P450 (CYP) and uridine diphosphate-glucuronosyltransferase (UGT) at the functional, structural and molecular levels to address questions of therapeutic relevance, particularly the wide interindividual variability in metabolic clearance characteristic of drugs and chemicals metabolized by these enzymes. 2. Studies in vivo initially identified the various factors that contribute to interindividual variability. Subsequently, human liver microsomal kinetic approaches, together with the cloning and functional characterization of recombinant CYP and UGT isoforms, led to the development of in vitro strategies that allowed the qualitative prediction of those factors likely to alter the metabolic clearance of a given compound in vivo. More recently, computer (in silico) modelling has been used to complement the laboratory based procedures. 3. The application of molecular biological approaches additionally allowed identification of the mutations responsible for CYP and UGT genetic polymorphism and, in some instances, the domains and individual amino acids that confer isoform substrate and inhibitor selectivities. Homology models, developed using X-ray crystallographic data as the template, potentially enable prediction of the functional consequences of altered CYP structure. 4. The rapid advances occurring in genomics, proteomics, gene expression analysis and computer modelling will allow further unravelling of the complexities of drug metabolism and improved prospects for the individualization of drug therapy.

Animals↗

Artificial neural network analysis for evaluation of peptide MS/MS spectra in proteomics.

The aim of the work was to explore usefulness of artificial neural network (ANN) analysis for the evaluation of proteomics data. The analysis was applied to the data generated by the widely used protein identification program Sequest, completed with several structural parameters readily calculated from peptide molecular formulas. Proteins from yeast cells were identified based on the MS/MS spectra of peptides. The constructed ANN was demonstrated to classify automatically as either "good" or "bad" the peptide MS/MS spectra otherwise classified manually. An appropriately trained ANN proves to be a high-throughput tool facilitating examination of Sequest's results. ANNs are recommended as a means of automatic processing of large amounts of MS/MS data, which normally must be considered in the analysis of complex mixtures of proteins in proteomics.

Artificial Intelligence↗

On 3-D graphical representation of proteomics maps and their numerical characterization.

We consider numerical characterization of proteomics maps by representing a map as a three-dimensional graphical object based on x, y coordinates of the spots and using their relative abundance as the z coordinate. In our representation the protein spots are first ordered based on their relative abundance and labeled accordingly. In the next step a 3-D path is constructed connecting spots having adjacent labels. Finally a matrix is constructed by assigning to each pairs of labels (i, j) matrix element, the numerical value of which is based on the quotients of the Euclidean distance and the distance along the 3-D zigzag between the two points. The approach has been illustrated on a fragment of a proteomics map and compared with 2-D graphical representation of proteomics maps.

Computer Graphics↗

MitoP2, an integrated database on mitochondrial proteins in yeast and man.

The aim of the MitoP2 database (http://ihg.gsf.de/mitop2) is to provide a comprehensive list of mitochondrial proteins of yeast and man. Based on the current literature we created an annotated reference set of yeast and human proteins. In addition, data sets relevant to the study of the mitochondrial proteome are integrated and accessible via search tools and links. They include computational predictions of signalling sequences, and summarize results from proteome mapping, mutant screening, expression profiling, protein-protein interaction and cellular sublocalization studies. For each individual approach, specificity and sensitivity for allocating mitochondrial proteins was calculated. By providing the evidence for mitochondrial candidate proteins the MitoP2 database lends itself to the genetic characterization of human mitochondriopathies.

Computational Biology↗

The NOESY jigsaw: automated protein secondary structure and main-chain assignment from sparse, unassigned NMR data.

High-throughput, data-directed computational protocols for Structural Genomics (or Proteomics) are required in order to evaluate the protein products of genes for structure and function at rates comparable to current gene-sequencing technology. This paper presents the JIGSAW algorithm, a novel high-throughput, automated approach to protein structure characterization with nuclear magnetic resonance (NMR). JIGSAW applies graph algorithms and probabilistic reasoning techniques, enforcing first-principles consistency rules in order to overcome a 5-10% signal-to-noise ratio. It consists of two main components: (1) graph-based secondary structure pattern identification in unassigned heteronuclear NMR data, and (2) assignment of spectral peaks by probabilistic alignment of identified secondary structure elements against the primary sequence. Deferring assignment eliminates the bottleneck faced by traditional approaches, which begin by correlating peaks among dozens of experiments. JIGSAW utilizes only four experiments, none of which requires 13C-labeled protein, thus dramatically reducing both the amount and expense of wet lab molecular biology and the total spectrometer time. Results for three test proteins demonstrate that JIGSAW correctly identifies 79-100% of alpha-helical and 46-65% of beta-sheet NOE connectivities and correctly aligns 33-100% of secondary structure elements. JIGSAW is very fast, running in minutes on a Pentium-class Linux workstation. This approach yields quick and reasonably accurate (as opposed to the traditional slow and extremely accurate) structure calculations. It could be useful for quick structural assays to speed data to the biologist early in an investigation and could in principle be applied in an automation-like fashion to a large fraction of the proteome.

Algorithms↗

Classification of bacterial species from proteomic data using combinatorial approaches incorporating artificial neural networks, cluster analysis and principal components analysis.

MOTIVATION: Robust computer algorithms are required to interpret the vast amounts of proteomic data currently being produced and to generate generalized models which are applicable to 'real world' scenarios. One such scenario is the classification of bacterial species. These vary immensely, some remaining remarkably stable whereas others are extremely labile showing rapid mutation and change. Such variation makes clinical diagnosis difficult and pathogens may be easily misidentified. RESULTS: We applied artificial neural networks (Neuroshell 2) in parallel with cluster analysis and principal components analysis to surface enhanced laser desorption/ionization (SELDI)-TOF mass spectrometry data with the aim of accurately identifying the bacterium Neisseria meningitidis from species within this genus and other closely related taxa. A subset of ions were identified that allowed for the consistent identification of species, classifying >97% of a separate validation subset of samples into their respective groups. AVAILABILITY: Neuroshell 2 is commercially available from Ward Systems.

Algorithms↗

MFAML: a standard data structure for representing and exchanging metabolic flux models.

SUMMARY: MFAML is a standard data structure designed for the formal representation and effective exchange of metabolic flux models. It allows for the explicit description of stationary states of a metabolic system by defining environmental/genetic conditions of the system, e.g. flux measurements, balancing constraints and physiological objectives as well as basic information on metabolites and reactions. In addition, a library of MFAML comprising a model parser and a converter provides an open framework for establishing the pipeline from metabolic modeling to metabolic flux analysis. AVAILABILITY: MFAML (version 1) is fully described and available at http://mbel.kaist.ac.kr/mfaml/.

Computer Simulation↗

Does the proteome encode organellar pH?

Inherent to the proteome itself, may be information that enables proteins to buffer pH at a level that promotes their own function within a specialized compartment. We observe that the distribution of computed isoelectric points in the yeast proteome matches experimentally derived organellar pH estimates across distinct subcellular compartments. This raises an interesting evolutionary question: did the pI of proteins and the pH of organelles co-evolve to optimize function?

Animals↗

Methods for peptide identification by spectral comparison.

BACKGROUND: Tandem mass spectrometry followed by database search is currently the predominant technology for peptide sequencing in shotgun proteomics experiments. Most methods compare experimentally observed spectra to the theoretical spectra predicted from the sequences in protein databases. There is a growing interest, however, in comparing unknown experimental spectra to a library of previously identified spectra. This approach has the advantage of taking into account instrument-dependent factors and peptide-specific differences in fragmentation probabilities. It is also computationally more efficient for high-throughput proteomics studies. RESULTS: This paper investigates computational issues related to this spectral comparison approach. Different methods have been empirically evaluated over several large sets of spectra. First, we illustrate that the peak intensities follow a Poisson distribution. This implies that applying a square root transform will optimally stabilize the peak intensity variance. Our results show that the square root did indeed outperform other transforms, resulting in improved accuracy of spectral matching. Second, different measures of spectral similarity were compared, and the results illustrated that the correlation coefficient was most robust. Finally, we examine how to assemble multiple spectra associated with the same peptide to generate a synthetic reference spectrum. Ensemble averaging is shown to provide the best combination of accuracy and efficiency. CONCLUSION: Our results demonstrate that when combined, these methods can boost the sensitivity and specificity of spectral comparison. Therefore they are capable of enhancing and complementing existing tools for consistent and accurate peptide identification.

Journal Article↗