Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Computational proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Bioinformatics and proteomics approaches for aging research.

Aging is a natural phenomenon that affects the entire physiology of an organism. Elucidating the molecular mechanisms underlying this complex process remains a major challenge today. Humans make poor models for research into aging because of their long life span. Thus, most of the current knowledge is through studies conducted in lower organisms. Large differences in life spans make it difficult to extrapolate the results of experiments carried out in model organisms to humans. Recent advances in genomic and proteomic technologies now permit generation of data pertaining to aging on a large-scale. In addition, several web-based community resources and databases are available that provide easy access to the available data. Use of bioinformatics and systems biology type of approaches provide a framework to start dissecting this complex biological phenomenon. Here, we discuss various genomic, transcriptomic and proteomic approaches that have the potential to provide a comprehensive mechanistic insight into the aging process.

Aging↗

Shotgun proteomics of cyanobacteria--applications of experimental and data-mining techniques.

Cyanobacteria are photosynthetic bacteria notable for their ability to produce hydrogen and a variety of interesting secondary metabolites. As a result of the growing number of completed cyanobacterial genome projects, the development of post-genomics analysis for this important group has been accelerating. DNA microarrays and classical two-dimensional gel electrophoresis (2DE) were the first technologies applied in such analyses. In many other systems, 'shotgun' proteomics employing multi-dimensional liquid chromatography and tandem mass spectrometry has proven to be a powerful tool. However, this approach has been relatively under-utilized in cyanobacteria. This study assesses progress in cyanobacterial shotgun proteomics to date, and adds a new perspective by developing a protocol for the shotgun proteomic analysis of the filamentous cyanobacterium Anabaena variabilis ATCC 29413, a model for N(2) fixation. Using approaches for enhanced protein extraction, 646 proteins were identified, which is more than double the previous results obtained using 2DE. Notably, the improved extraction method and shotgun approach resulted in a significantly higher representation of basic and hydrophobic proteins. The use of protein bioinformatics tools to further mine these shotgun data is illustrated through the application of PSORTb for localization, the grand average hydropathy (GRAVY) index for hydrophobicity, LipoP for lipoproteins and the exponentially modified protein abundance index (emPAI) for abundance. The results are compared with the most well-studied cyanobacterium, Synechocystis sp. PCC 6803. Some general issues in shotgun proteome identification and quantification are then addressed.

Computational Biology↗

Measures of diversity for populations and distances between individuals with highly reorganizable genomes.

In this paper we address the problem of defining a measure of diversity for a population of individuals whose genome can be subjected to major reorganizations during the evolutionary process. To this end, we introduce a measure of diversity for populations of strings of variable length defined on a finite alphabet, and from this measure we derive a semi-metric distance between pairs of strings. The definitions are based on counting the number of substrings of the strings, considered first separately and then collectively. This approach is related to the concept of linguistic complexity, whose definition we generalize from single strings to populations. Using the substring count approach we also define a new kind of Tanimoto distance between strings. We show how to extend the approach to representations that are not based on strings and, in particular, to the tree-based representations used in the field of genetic programming. We describe how suffix trees can allow these measures and distances to be implemented with a computational cost that is linear in both space and time relative to the length of the strings and the size of the population. The definitions were devised to assess the diversity of populations having genomes of variable length and variable structure during evolutionary computation runs, but applications in quantitative genomics, proteomics, and pattern recognition can be also envisaged.

Computational Biology↗

Software-induced variance in two-dimensional gel electrophoresis image analysis.

Experimental variability in 2-DE is well documented, but little attention has been paid to variability arising from postexperimental quantitative analyses using various 2-DE software packages. The performance of two 2-DE analysis software programs, Phoretix 2D Expression v2004 (Expression) and PDQuest 7.2 (PDQuest), was evaluated in this study. All available background subtraction and smoothing algorithms were tested using both data generated from one single 2-DE gel image, thus excluding experimental variance, and with authentic sets of replicate gels (n = 5). A slight shift of the image boundaries (the "cropping area") caused both programs to induce variance in protein spot quantification of otherwise identical gel images. The resulting variance for PDQuest (CV(mean) = 8%) was approximately twice that for Expression (CV(mean) = 4%). In authentic sets of replicate 2-DE gels (n = 5), the experimental variance confounded the software-induced variance to some extent. However, Expression still outperformed PDQuest, which exhibited software-induced variance as high as 25% of the total observed variance. Surprisingly, the complete omission of background subtraction algorithms resulted in the least amount of software-based variance. These data indicate that 2-DE gel analysis software constitutes a significant source of the variance observed in quantitative proteomics, and that the use of background subtraction algorithms can further increase the variance.

Algorithms↗

Identification of selenomethionine in selenized yeast using two-dimensional liquid chromatography-mass spectrometry based proteomic analysis.

Selenium-enriched yeast has been commonly used as a nutritional supplement. Here we describe a protocol used to investigate the metabolic fate of inorganic selenium in yeast. We provide definitive, mass spectrometry based evidence for the non-specific incorporation of selenomethionine in the yeast proteome involving the replacement of about 30% of all methionine with selenomethionine.

Dietary Supplements↗

Conserved sequences of prokaryotic proteomes and their compositional age.

A full repertoire of octapeptides which are present in at least 30 bacterial proteomes of total 131 currently available is computationally derived and filtered. An original search technique is used that, in terms of computational time and memory, is similar to the Suffix tree method. The presence of a given sequence in a large number of proteomes qualifies it as a conserved sequence. The larger the number of proteomes where it is found, the higher is the conservation. The concept of compositional age of the amino acid sequences ("compositional clock") is introduced for the first time. The compositional age is calculated on the basis of the consensus temporal order of appearance of amino acids in early evolution. The correlation between the compositional age and the sequence conservation is established.

Amino Acid Sequence↗

Genome informatics: current status and future prospects.

This article reviews recent advances in genomics and informatics relevant to cardiovascular research. In particular, we review the status of (1) whole genome sequencing efforts in human, mouse, rat, zebrafish, and dog; (2) the development of data mining and analysis tools; (3) the launching of the National Heart, Lung, and Blood Institute Programs for Genomics Applications and Proteomics Initiative; (4) efforts to characterize the cardiac transcriptome and proteome; and (5) the current status of computational modeling of the cardiac myocyte. In each instance, we provide links to relevant sources of information on the World Wide Web and critical appraisals of the promises and the challenges of an expanding and diverse information landscape.

Animals↗

Fifteenth annual Pezcoller symposium: molecular in vivo visualization of cancer cells.

Recent advances in the understanding of the molecular phenomena underlying neoplastic development are leading to the recognition of metabolic features characteristic of the cancer cell. It has become possible to visualize the presence of these cells in vivo and to follow their progression toward increasing anaplastic behavior through measurements of molecular markers that can be achieved by physical or biochemical means.

Animals↗

SNA--a toolbox for the stoichiometric analysis of metabolic networks.

BACKGROUND: Despite recent algorithmic and conceptual progress, the stoichiometric network analysis of large metabolic models remains a computationally challenging problem. RESULTS: SNA is a interactive, high performance toolbox for analysing the possible steady state behaviour of metabolic networks by computing the generating and elementary vectors of their flux and conversions cones. It also supports analysing the steady states by linear programming. The toolbox is implemented mainly in Mathematica and returns numerically exact results. It is available under an open source license from: http://bioinformatics.org/project/?group_id=546. CONCLUSION: Thanks to its performance and modular design, SNA is demonstrably useful in analysing genome scale metabolic networks. Further, the integration into Mathematica provides a very flexible environment for the subsequent analysis and interpretation of the results.

Algorithms↗

Transforming omics data into context: bioinformatics on genomics and proteomics raw data.

Differential gene expression analysis and proteomics have exerted significant impact on the elucidation of concerted cellular processes, as simultaneous measurement of hundreds to thousands of individual objects on the level of RNA and protein ensembles became technically feasible. The availability of such data sets has promised a profound understanding of phenomena on an aggregate level, expressed as the phenotypic response (observables) of cells, e.g., in the presence of drugs, or characterization of cells and tissue displaying distinct patho-physiological states. However, the step of transforming these data into context, i.e., linking distinct expression or abundance patterns with phenotypic observables - and furthermore enabling a sound biological interpretation on the level of reaction networks and concerted pathways, is still a major shortcoming. This finding is certainly based on the enormous complexity embedded in cellular reaction networks, but a variety of computational approaches have been developed over the last few years to overcome these issues. This review provides an overview on computational procedures for analysis of genomic and proteomic data introducing a sequential analysis workflow: Explorative statistics for deriving a first, from the purely statistical viewpoint, relevant candidate gene/protein list, followed by co-regulation and network analysis to biologically expand this core list toward functional networks and pathways. The review on these procedures is complemented by example applications tailored at identification of disease-associated proteins. Optimization of computational procedures involved, in conjunction with the continuous increase in additional biological data, clearly has the potential of boosting our understanding of processes on a cell-wide level.

Animals↗

Next-generation brain proteomics: Integrating single-cell, spatial, and multi-omics for clinical biomarker discovery.

The mammalian brain's functional complexity arises from the sophisticated architecture of neurons and glia. This network is essentially defined by its dynamic proteome, which reveals the functional execution underlying neural computation and disease. This review integrates the technological leap in neuroproteomics. It has moved beyond bulk tissue proteome cataloguing to high-sensitivity single-cell and spatial resolution. We detail how next-generation platforms, such as TIMS-PASEF and Orbitrap-Astral, have enabled deeper and faster phenotypic profiling of limited brain samples. However, the proteome coverage remains constrained by dynamic range, sample loss, ionisation bias and incomplete detection of low-abundance regulatory proteins. We further examine how such studies have revealed the proteomic remodelling that drives lineage specification and synaptic plasticity by linking temporal protein expression waves to biological function. Crucially, we delineate the clinical translational trajectory, illustrating how aberrant signatures are verified in cerebrospinal fluid (CSF) and validated in plasma to support precision medicine. Finally, we argue for the necessity of "fused" multi-omics integration and Artificial Intelligence (AI) to decode the non-linear molecular logic of brain pathology.

Humans↗

Sensitive quantitative predictions of peptide-MHC binding by a 'Query by Committee' artificial neural network approach.

We have generated Artificial Neural Networks (ANN) capable of performing sensitive, quantitative predictions of peptide binding to the MHC class I molecule, HLA-A*0204. We have shown that such quantitative ANN are superior to conventional classification ANN, that have been trained to predict binding vs non-binding peptides. Furthermore, quantitative ANN allowed a straightforward application of a 'Query by Committee' (QBC) principle whereby particularly information-rich peptides could be identified and subsequently tested experimentally. Iterative training based on QBC-selected peptides considerably increased the sensitivity without compromising the efficiency of the prediction. This suggests a general, rational and unbiased approach to the development of high quality predictions of epitopes restricted to this and other HLA molecules. Due to their quantitative nature, such predictions will cover a wide range of MHC-binding affinities of immunological interest, and they can be readily integrated with predictions of other events involved in generating immunogenic epitopes. These predictions have the capacity to perform rapid proteome-wide searches for epitopes. Finally, it is an example of an iterative feedback loop whereby advanced, computational bioinformatics optimize experimental strategy, and vice versa.

HLA-A Antigens↗

Report on the 1st International Conference of the Hellenic Proteomics Society in Athens, Greece.

The First International Conference of the Hellenic Proteomics Society took place at the Foundation for Biomedical Research of the Academy of Athens, Athens, Greece, 22-25 May 2005. Scientists from about 20 countries attended this conference where proceedings in proteomics methodologies, advances in mass spectrometry, proteomics applications, disease diagnostics and bioinformatics were presented. The relatively small size of the meeting gave the opportunity for the attendees to interact and discuss projects and collaborations. The event was completely financed by companies which exhibited their products and services related to proteomics research.

Computational Biology↗

APBioNet: the Asia-Pacific regional consortium for bioinformatics.

Bioinformatics and computational biology, along with the related fields of genomics, proteomics, functional genomics and systems biology are new wave scientific disciplines that harness composite computational power across networks to advance biological knowledge at the most basic level and to direct traditional laboratory-based research efforts in the biomedical sciences. 'Fostering the growth of bioinformatics and allied disciplines in the Asia-Pacific region' is the motto of the first regional bioinformatics society, the Asia-Pacific Bioinformatics Network (APBioNet). APBioNet addresses the issues of hardware, software, databases and networks pertaining to bioinformatics, with the additional layer of pertinent education, training and research. Recent milestones achieved include hosting an international bioinformatics symposium in Asia and setting up large-scale regional grid-computing projects.

Asia↗

LOCATE: a mouse protein subcellular localization database.

We present here LOCATE, a curated, web-accessible database that houses data describing the membrane organization and subcellular localization of proteins from the FANTOM3 Isoform Protein Sequence set. Membrane organization is predicted by the high-throughput, computational pipeline MemO. The subcellular locations of selected proteins from this set were determined by a high-throughput, immunofluorescence-based assay and by manually reviewing >1700 peer-reviewed publications. LOCATE represents the first effort to catalogue the experimentally verified subcellular location and membrane organization of mammalian proteins using a high-throughput approach and provides localization data for approximately 40% of the mouse proteome. It is available at http://locate.imb.uq.edu.au.

Animals↗

Assessing the precision of high-throughput computational and laboratory approaches for the genome-wide identification of protein subcellular localization in bacteria.

BACKGROUND: Identification of a bacterial protein's subcellular localization (SCL) is important for genome annotation, function prediction and drug or vaccine target identification. Subcellular fractionation techniques combined with recent proteomics technology permits the identification of large numbers of proteins from distinct bacterial compartments. However, the fractionation of a complex structure like the cell into several subcellular compartments is not a trivial task. Contamination from other compartments may occur, and some proteins may reside in multiple localizations. New computational methods have been reported over the past few years that now permit much more accurate, genome-wide analysis of the SCL of protein sequences deduced from genomes. There is a need to compare such computational methods with laboratory proteomics approaches to identify the most effective current approach for genome-wide localization characterization and annotation. RESULTS: In this study, ten subcellular proteome analyses of bacterial compartments were reviewed. PSORTb version 2.0 was used to computationally predict the localization of proteins reported in these publications, and these computational predictions were then compared to the localizations determined by the proteomics study. By using a combined approach, we were able to identify a number of contaminants and proteins with dual localizations, and were able to more accurately identify membrane subproteomes. Our results allowed us to estimate the precision level of laboratory subproteome studies and we show here that, on average, recent high-precision computational methods such as PSORTb now have a lower error rate than laboratory methods. CONCLUSION: We have performed the first focused comparison of genome-wide proteomic and computational methods for subcellular localization identification, and show that computational methods have now attained a level of precision that is exceeding that of high-throughput laboratory approaches. We note that analysis of all cellular fractions collectively is required to effectively provide localization information from laboratory studies, and we propose an overall approach to genome-wide subcellular localization characterization that capitalizes on the complementary nature of current laboratory and computational methods.

Bacterial Proteins↗