Challenges and prospects of plant proteomics.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
UNLABELLED: BIOCHAM (the BIOCHemical Abstract Machine) is a software environment for modeling biochemical systems. It is based on two aspects: (1) the analysis and simulation of boolean, kinetic and stochastic models and (2) the formalization of biological properties in temporal logic. BIOCHAM provides tools and languages for describing protein networks with a simple and straightforward syntax, and for integrating biological properties into the model. It then becomes possible to analyze, query, verify and maintain the model with respect to those properties. For kinetic models, BIOCHAM can search for appropriate parameter values in order to reproduce a specific behavior observed in experiments and formalized in temporal logic. Coupled with other methods such as bifurcation diagrams, this search assists the modeler/biologist in the modeling process. AVAILABILITY: BIOCHAM (v. 2.5) is a free software available for download, with example models, at http://contraintes.inria.fr/BIOCHAM/.
Pathway Analyst (Path-A) is a publicly available web server (http://path-a.cs.ualberta.ca) that predicts metabolic pathways. It takes a FASTA format file containing a set of query protein sequences from a single organism (a partial or complete proteome) and identifies those sequences that are likely to participate in any of its supported metabolic pathways (currently 10). Path-A uses a number of machine-learning and sequence analysis techniques (e.g. SVM, BLAST and HMM) to predict pathways. Each machine-learned classifier exploits similarity between sequences in the pathways of its model organisms and sequences in the query set. It predicts the pathways that are present in the query organism and annotates each predicted reaction and catalyst, using the appropriate sequences from the query set. Path-A also provides a browsable and searchable database of the pathways for the model organisms that are used to make its predictions. Path-A's predictor sets (using different classifier technologies) have been evaluated using standard cross-validation techniques on a dataset of 10 metabolic pathways across 13 model organisms--a total of 125 organism-specific pathways. The most accurate classifier technology obtained a mean precision of 78.3% and a mean recall of 92.6% in predicting all catalyst proteins, of all reactions, in all pathways present in the dataset. Although Path-A currently only supports metabolic pathways, the underlying prediction techniques are general enough for other types of pathways. Consequently, it is our intent to extend Path-A to predict other types of pathways, including signalling pathways.
A great challenge in the proteomics and structural genomics era is to predict protein structure and function, including identification of those proteins that are partially or wholly unstructured. Disordered regions in proteins often contain short linear peptide motifs (e.g., SH3 ligands and targeting signals) that are important for protein function. We present here DisEMBL, a computational tool for prediction of disordered/unstructured regions within a protein sequence. As no clear definition of disorder exists, we have developed parameters based on several alternative definitions and introduced a new one based on the concept of "hot loops," i.e., coils with high temperature factors. Avoiding potentially disordered segments in protein expression constructs can increase expression, foldability, and stability of the expressed protein. DisEMBL is thus useful for target selection and the design of constructs as needed for many biochemical studies, particularly structural biology and structural genomics projects. The tool is freely available via a web interface (http://dis.embl.de) and can be downloaded for use in large-scale studies.
BACKGROUND: The increasing number of protein sequences and 3D structure obtained from genomic initiatives is leading many of us to focus on proteomics, and to dedicate our experimental and computational efforts on the creation and analysis of information derived from 3D structure. In particular, the high-throughput generation of protein-protein interaction data from a few organisms makes such an approach very important towards understanding the molecular recognition that make-up the entire protein-protein interaction network. Since the generation of sequences, and experimental protein-protein interactions increases faster than the 3D structure determination of protein complexes, there is tremendous interest in developing in silico methods that generate such structure for prediction and classification purposes. In this study we focused on classifying protein family members based on their protein-protein interaction distinctiveness. Structure-based classification of protein-protein interfaces has been described initially by Ponstingl et al. 1 and more recently by Valdar et al. 2 and Mintseris et al. 3, from complex structures that have been solved experimentally. However, little has been done on protein classification based on the prediction of protein-protein complexes obtained from homology modeling and docking simulation. RESULTS: We have developed an in silico classification system entitled HODOCO (Homology modeling, Docking and Classification Oracle), in which protein Residue Potential Interaction Profiles (RPIPS) are used to summarize protein-protein interaction characteristics. This system applied to a dataset of 64 proteins of the death domain superfamily was used to classify each member into its proper subfamily. Two classification methods were attempted, heuristic and support vector machine learning. Both methods were tested with a 5-fold cross-validation. The heuristic approach yielded a 61% average accuracy, while the machine learning approach yielded an 89% average accuracy. CONCLUSION: We have confirmed the reliability and potential value of classifying proteins via their predicted interactions. Our results are in the same range of accuracy as other studies that classify protein-protein interactions from 3D complex structure obtained experimentally. While our classification scheme does not take directly into account sequence information our results are in agreement with functional and sequence based classification of death domain family members.
A report on the Fourth Annual HUPO World Congress (HUPO2005) 'From Defining the Proteome to Understanding Function', Munich, Germany, 28 August-1 September 2005.
Reference maps of the cytosolic, cell surface and extracellular proteome fractions of the amino acid-producing soil bacterium Corynebacterium efficiens YS-314 were established. The analysis window covers a pI range from 3 to 7 along with a molecular mass range from 10 to 130 kDa. After second-dimensional separation on SDS-PAGE and Coomassie staining, computational analysis detected 635 protein spots in the cytosolic proteome fraction, whereas 76 and 102 spots were detected in the cell surface and extracellular proteomes, respectively. By means of MALDI-TOF-MS and tryptic peptide mass fingerprinting, 164 cytosolic proteins, 49 proteins of the cell surface and 89 extracellular protein spots were identified, representing in total 177 different proteins. Additionally, reference maps of the three cellular proteome fractions of the close phylogenetic relative Corynebacterium glutamicum ATCC 13032 were generated and used for comparative proteomics. Classification according to the Clusters of Orthologous Groups of proteins scheme and abundance analysis of the identified proteins revealed species-specific differences. The high abundance of molecular chaperones and amino acid biosynthesis enzymes in C. efficiens points to environmental adaptations of this recently discovered amino acid-producing bacterium.
There has been substantial evidence for more than three decades that the major psychiatric illnesses such as schizophrenia, bipolar disorder, autism, and alcoholism have a strong genetic basis. During the past 15 years considerable effort has been expended in trying to establish the genetic loci associated with susceptibility to these and other mental disorders using principally linkage analysis. Despite this, only a handful of specific genes have been identified, and it is now generally recognized that further advances along these lines will require the analysis of literally hundreds of affected individuals and their families. Fortunately, the emergence in the past three years of a number of new approaches and more effective tools has given new hope to those engaged in the search for the underlying genetic and environmental factors involved in causing these illnesses, which collectively are among the most serious in all societies. Chief among these new tools is the availability of the entire human genome sequence and the prospect that within the next several years the entire complement of human genes will be known and the functions of most of their protein products elucidated. In the meantime the search for susceptibility loci is being facilitated by the availability of single nucleotide polymorphisms (SNPs) and by the beginning of haplotype mapping, which tracks the distribution of clusters of SNPs that segregate as a group. Together with high throughput DNA sequencing, microarrays for whole genome scanning, advances in proteomics, and the development of more sophisticated computer programs for analyzing sequence and association data, these advances hold promise of greatly accelerating the search for the genetic basis of most mental illnesses while, at the same time, providing molecular targets for the development of new and more effective therapies.
Now that the human genome is completed, the characterization of the proteins encoded by the sequence remains a challenging task. The study of the complete protein complement of the genome, the "proteome," referred to as proteomics, will be essential if new therapeutic drugs and new disease biomarkers for early diagnosis are to be developed. Research efforts are already underway to develop the technology necessary to compare the specific protein profiles of diseased versus nondiseased states. These technologies provide a wealth of information and rapidly generate large quantities of data. Processing the large amounts of data will lead to useful predictive mathematical descriptions of biological systems which will permit rapid identification of novel therapeutic targets and identification of metabolic disorders. Here, we present an overview of the current status and future research approaches in defining the cancer cell's proteome in combination with different bioinformatics and computational biology tools toward a better understanding of health and disease.
BACKGROUND: Although endurance exercise benefits liver health, sex-specific adaptive trajectories remain unclear. This study mapped dynamic liver adaptation in males and females during prolonged training and identified underlying molecular programs. METHODS: Using publicly available time-resolved liver multi-omics data generated by the Molecular Transducers of Physical Activity Consortium (MoTrPAC), we established a computational pipeline for differential analysis of transcriptomic, proteomic, phosphoproteomic, and metabolomic data with FDR correction, followed by FGSEA pathway enrichment. Kinase activities were inferred through ortholog mapping and PhosphoSitePlus. Cross-omics co-expression networks were constructed using WGCNA and topological overlap to link omics features with physiological phenotypes. For experimental validation, liver tissues were collected from endurance-trained Sprague-Dawley rats, and key nodes were confirmed by Western blotting, qRT-PCR, and immunofluorescence/immunohistochemical staining. Public scRNA-seq data were further integrated to map multi-omics signals to single-cell resolution and assess functional changes in specific cell types. RESULTS: The hepatic response to exercise stress was stage-specific, shifting from early transcriptional activation to later proteomic and metabolic remodeling. Multi-omics integration revealed distinct sex-associated adaptive trajectories: males were more strongly associated with energy metabolism, redox-related programs, and amino acid/organic acid catabolism, whereas females showed prominent membrane lipid remodeling, proteostasis -related programs, and mitochondrial/ribosomal translational features. Single-cell analysis showed that tissue remodeling occurred without major lineage turnover, instead involving altered communication among pre-existing cell communities. Validation of PPP1R3G identified a protein-dominant exercise-responsive marker, supporting the contribution of post-transcriptional or protein-level regulation. CONCLUSIONS: Hepatic adaptation to endurance stress follows a cross-omics evolutionary pattern with sex-specific reprogramming of energy supply and homeostatic maintenance. This time-resolved framework clarifies how exercise improves liver function and supports sex-oriented metabolic interventions and therapeutic target discovery.
Prophylactic vaccination against tuberculosis with BCG gained much of the credit for the decline of TB in Europe. However, with TB resurgent in many parts of the world, better vaccines are urgently needed. To improve on BCG, a rapid, rational approach to vaccine discovery is needed. Fortunately, advances in the fields of molecular biology and computer science have spawned new disciplines: Genomics, Proteomics and Transcriptomics are transforming the ways in which candidate vaccine antigens are discovered. In this review, we discuss how these new approaches have accelerated the pace of antigen discovery and vaccine development, and highlight some of the most promising new candidate vaccines and vaccine targets.
With the large amount of genomics and proteomics data that we are confronted with, computational support for the elucidation of protein function becomes more and more pressing. Many different kinds of biological data harbour signals of protein function, but these signals are often concealed. Computational methods that use protein sequence and structure data can be used for discovering these signals. They provide information that can substantially speed up experimental function elucidation. In this review we concentrate on such methods.
Recent developments in proteomics and genomics provide huge quantities of data to analyze. Automatic interpretation of mass spectrometry data has become essential for high-throughput processes aiming to study complete proteomes. There exist two main sources of mass spectrometric data: peptide mass fingerprint and fragmentation spectra, both of which require specific bioinformatic algorithms. We present a survey of these algorithms and discuss the efficiency of the different approaches and the possible improvements that may lead to a complete automatic high-throughput identification process.
In this paper, based on the approach by combining the "functional domain composition" [K.C. Chou, Y. D. Cai, J. Biol. Chem. 277 (2002) 45765] and the pseudo-amino acid composition [K.C. Chou, Proteins Struct. Funct. Genet. 43 (2001) 246; Correction Proteins Struct. Funct. Genet. 2044 (2001) 2060], the Nearest Neighbour Algorithm (NNA) was developed for predicting the protein subcellular location. Very high success rates were observed, suggesting that such a hybrid approach may become a useful high-throughput tool in the area of bioinformatics and proteomics.
Explore the source record for details and available documents.
MOTIVATION: Membrane proteins are an abundant and functionally relevant subset of proteins that putatively include from about 15 up to 30% of the proteome of organisms fully sequenced. These estimates are mainly computed on the basis of sequence comparison and membrane protein prediction. It is therefore urgent to develop methods capable of selecting membrane proteins especially in the case of outer membrane proteins, barely taken into consideration when proteome wide analysis is performed. This will also help protein annotation when no homologous sequence is found in the database. Outer membrane proteins solved so far at atomic resolution interact with the external membrane of bacteria with a characteristic beta barrel structure comprising different even numbers of beta strands (beta barrel membrane proteins). In this they differ from the membrane proteins of the cytoplasmic membrane endowed with alpha helix bundles (all alpha membrane proteins) and need specialised predictors. RESULTS: We develop a HMM model, which can predict the topology of beta barrel membrane proteins using, as input, evolutionary information. The model is cyclic with 6 types of states: two for the beta strand transmembrane core, one for the beta strand cap on either side of the membrane, one for the inner loop, one for the outer loop and one for the globular domain state in the middle of each loop. The development of a specific input for HMM based on multiple sequence alignment is novel. The accuracy per residue of the model is 83% when a jack knife procedure is adopted. With a model optimisation method using a dynamic programming algorithm seven topological models out of the twelve proteins included in the testing set are also correctly predicted. When used as a discriminator, the model is rather selective. At a fixed probability value, it retains 84% of a non-redundant set comprising 145 sequences of well-annotated outer membrane proteins. Concomitantly, it correctly rejects 90% of a set of globular proteins including about 1200 chains with low sequence identity (<30%) and 90% of a set of all alpha membrane proteins, including 188 chains.
MOTIVATION: Simulation and modeling is becoming a standard approach to understand complex biochemical processes. Therefore, there is a big need for software tools that allow access to diverse simulation and modeling methods as well as support for the usage of these methods. RESULTS: Here, we present COPASI, a platform-independent and user-friendly biochemical simulator that offers several unique features. We discuss numerical issues with these features; in particular, the criteria to switch between stochastic and deterministic simulation methods, hybrid deterministic-stochastic methods, and the importance of random number generator numerical resolution in stochastic simulation. AVAILABILITY: The complete software is available in binary (executable) for MS Windows, OS X, Linux (Intel) and Sun Solaris (SPARC), as well as the full source code under an open source license from http://www.copasi.org.
Explore the source record for details and available documents.