Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

Integration from proteins to organs: the Physiome Project.

The Physiome Project will provide a framework for modelling the human body, using computational methods that incorporate biochemical, biophysical and anatomical information on cells, tissues and organs. The main project goals are to use computational modelling to analyse integrative biological function and to provide a system for hypothesis testing.

Computer Simulation↗

The prokaryotic selenoproteome.

In the genetic code, the UGA codon has a dual function as it encodes selenocysteine (Sec) and serves as a stop signal. However, only the translation terminator function is used in gene annotation programs, resulting in misannotation of selenoprotein genes. Here, we applied two independent bioinformatics approaches to characterize a selenoprotein set in prokaryotic genomes. One method searched for selenoprotein genes by identifying RNA stem-loop structures, selenocysteine insertion sequence elements; the second approach identified Sec/Cys pairs in homologous sequences. These analyses identified all or almost all selenoproteins in completely sequenced bacterial and archaeal genomes and provided a view on the distribution and composition of prokaryotic selenoproteomes. In addition, lineage-specific and core selenoproteins were detected, which provided insights into the mechanisms of selenoprotein evolution. Characterization of selenoproteomes allows interpretation of other UGA codons in completed genomes of prokaryotes as terminators, addressing the UGA dual-function problem.

Amino Acid Sequence↗

Challenges to be faced in the reconstruction of metabolic networks from public databases.

In the post-genomic era, the biochemical information for individual compounds, enzymes, reactions to be found within named organisms has become readily available. The well-known KEGG and BioCyc databases provide a comprehensive catalogue for this information and have thereby substantially aided the scientific community. Using these databases, the complement of enzymes present in a given organism can be determined and, in principle, used to reconstruct the metabolic network. However, such reconstructed networks contain numerous properties contradicting biological expectation. The metabolic networks for a number of organisms are reconstructed from KEGG and BioCyc databases, and features of these networks are related to properties of their originating database.

Algorithms↗

The Pathway Tools software.

MOTIVATION: Bioinformatics requires reusable software tools for creating model-organism databases (MODs). RESULTS: The Pathway Tools is a reusable, production-quality software environment for creating a type of MOD called a Pathway/Genome Database (PGDB). A PGDB such as EcoCyc (see http://ecocyc.org) integrates our evolving understanding of the genes, proteins, metabolic network, and genetic network of an organism. This paper provides an overview of the four main components of the Pathway Tools: The PathoLogic component supports creation of new PGDBs from the annotated genome of an organism. The Pathway/Genome Navigator provides query, visualization, and Web-publishing services for PGDBs. The Pathway/Genome Editors support interactive updating of PGDBs. The Pathway Tools ontology defines the schema of PGDBs. The Pathway Tools makes use of the Ocelot object database system for data management services for PGDBs. The Pathway Tools has been used to build PGDBs for 13 organisms within SRI and by external users.

Abstracting and Indexing↗

HPID: the Human Protein Interaction Database.

UNLABELLED: The Human Protein Interaction Database (http://www.hpid.org) was designed (1) to provide human protein interaction information pre-computed from existing structural and experimental data, (2) to predict potential interactions between proteins submitted by users and (3) to provide a depository for new human protein interaction data from users. Two types of interaction are available from the pre-computed data: (1) interactions at the protein superfamily level and (2) those transferred from the interactions of yeast proteins. Interactions at the superfamily level were obtained by locating known structural interactions of the PDB in the SCOP domains and identifying homologs of the domains in the human proteins. Interactions transferred from yeast proteins were obtained by identifying homologs of the yeast proteins in the human proteins. For each human protein in the database and each query submitted by users, the protein superfamilies and yeast proteins assigned to the protein are shown, along with their interacting partners. We have also developed a set of web-based programs so that users can visualize and analyze protein interaction networks in order to explore the networks further. AVAILABILITY: http://www.hpid.org.

Algorithms↗

Inferring quantitative models of regulatory networks from expression data.

MOTIVATION: Genetic networks regulate key processes in living cells. Various methods have been suggested to reconstruct network architecture from gene expression data. However, most approaches are based on qualitative models that provide only rough approximations of the underlying events, and lack the quantitative aspects that are critical for understanding the proper function of biomolecular systems. RESULTS: We present fine-grained dynamical models of gene transcription and develop methods for reconstructing them from gene expression data within the framework of a generative probabilistic model. Unlike previous works, we employ quantitative transcription rates, and simultaneously estimate both the kinetic parameters that govern these rates, and the activity levels of unobserved regulators that control them. We apply our approach to expression datasets from yeast and show that we can learn the unknown regulator activity profiles, as well as the binding affinity parameters. We also introduce a novel structure learning algorithm, and demonstrate its power to accurately reconstruct the regulatory network from those datasets.

Binding Sites↗

Predicting protein-protein interaction by searching evolutionary tree automorphism space.

MOTIVATION: Uncovering the protein-protein interaction network is a fundamental step in the quest to understand the molecular machinery of a cell. This motivates the search for efficient computational methods for predicting such interactions. Among the available predictors are those that are based on the co-evolution hypothesis "evolutionary trees of protein families (that are known to interact) are expected to have similar topologies". Many of these methods are limited by the fact that they can handle only a small number of protein sequences. Also, details on evolutionary tree topology are missing as they use similarity matrices in lieu of the trees. RESULTS: We introduce MORPH, a new algorithm for predicting protein interaction partners between members of two protein families that are known to interact. Our approach can also be seen as a new method for searching the best superposition of the corresponding evolutionary trees based on tree automorphism group. We discuss relevant facts related to the predictability of protein-protein interaction based on their co-evolution. When compared with related computational approaches, our method reduces the search space by approximately 3 x 10(5)-fold and at the same time increases the accuracy of predicting correct binding partners.

Algorithms↗

Computational identification of human mitochondrial proteins based on homology to yeast mitochondrially targeted proteins.

MOTIVATION: Patients with defects of the mitochondrial respiratory chain due to mutations in nuclear genes are often undiagnosable due to the lack of information about the role of these genes. We therefore sought to produce a novel dataset of human nuclear-encoded mitochondrial proteins. RESULTS: We have used the web-based computer program Mitoprot to predict which proteins in the Saccharomyces cerevisiae genome are targeted to mitochondria. We then used this protein dataset to identify the homologous human proteins in the Unigene database using TBLASTN from NCBI. Human proteins with an Expectation value <10(-5) and an Identity >30% were accepted as true homologues of the yeast proteins. These human proteins were then reanalyzed with Mitoprot. The final set of proteins comprises a dataset of 361 human mitochondrially targeted proteins with homology to all S.cerevisiae mitochondrially targeted proteins. One hundred twenty eight of these proteins are novel and are of unknown function. SUPPLEMENTARY INFORMATION: Supplementary tables will be available from http://www.sickkids.ca/Robinsonlab/

Conserved Sequence↗

TopNet: a tool for comparing biological sub-networks, correlating protein properties with topological statistics.

Biological networks are a topic of great current interest, particularly with the publication of a number of large genome-wide interaction datasets. They are globally characterized by a variety of graph-theoretic statistics, such as the degree distribution, clustering coefficient, characteristic path length and diameter. Moreover, real protein networks are quite complex and can often be divided into many sub-networks through systematic selection of different nodes and edges. For instance, proteins can be sub-divided by expression level, length, amino-acid composition, solubility, secondary structure and function. A challenging research question is to compare the topologies of sub- networks, looking for global differences associated with different types of proteins. TopNet is an automated web tool designed to address this question, calculating and comparing topological characteristics for different sub-networks derived from any given protein network. It provides reasonable solutions to the calculation of network statistics for sub-networks embedded within a larger network and gives simplified views of a sub-network of interest, allowing one to navigate through it. After constructing TopNet, we applied it to the interaction networks and protein classes currently available for yeast. We were able to find a number of potential biological correlations. In particular, we found that soluble proteins had more interactions than membrane proteins. Moreover, amongst soluble proteins, those that were highly expressed, had many polar amino acids, and had many alpha helices, tended to have the most interaction partners. Interestingly, TopNet also turned up some systematic biases in the current yeast interaction network: on average, proteins with a known functional classification had many more interaction partners than those without. This phenomenon may reflect the incompleteness of the experimentally determined yeast interaction network.

Algorithms↗

CYCLONET--an integrated database on cell cycle regulation and carcinogenesis.

Computational modelling of mammalian cell cycle regulation is a challenging task, which requires comprehensive knowledge on many interrelated processes in the cell. We have developed a web-based integrated database on cell cycle regulation in mammals in normal and pathological states (Cyclonet database). It integrates data obtained by 'omics' sciences and chemoinformatics on the basis of systems biology approach. Cyclonet is a specialized resource, which enables researchers working in the field of anticancer drug discovery to analyze the wealth of currently available information in a systematic way. Cyclonet contains information on relevant genes and molecules; diagrams and models of cell cycle regulation and results of their simulation; microarray data on cell cycle and on various types of cancer, information on drug targets and their ligands, as well as extensive bibliography on modelling of cell cycle and cancer-related gene expression data. The Cyclonet database is also accessible through the BioUML workbench, which allows flexible querying, analyzing and editing the data by means of visual modelling. Cyclonet aims to predict promising anticancer targets and their agents by application of Prediction of Activity Spectra for Substances. The Cyclonet database is available at http://cyclonet.biouml.org.

Animals↗

Signatures of domain shuffling in the human genome.

To elucidate the role of exon shuffling in shaping the complexity of the human genome/proteome, we have systematically analyzed intron phase distributions in the coding sequence of human protein domains. We found that introns at the boundaries of domains show high excess of symmetrical phase combinations (i.e., 0-0, 1-1, and 2-2), whereas nonboundary introns show no excess symmetry. This suggests that exon shuffling has primarily involved rearrangement of structural and functional domains as a whole. Furthermore, we found that domains flanked by phase 1 introns have dramatically expanded in the human genome due to domain shuffling and that 1-1 symmetrical domains and domain families are nonrandomly distributed with respect to their age. The predominance and extracellular location of 1-1 symmetrical domains among domains specific to metazoans suggests that they are associated with the rise of multicellularity. On the other hand, 0-0 symmetrical domains tend to be over-represented among ancient protein domains that are shared between the eukaryotic and prokaryotic kingdoms, which is compatible with the suggestion of primordial domain shuffling in the progenote. To see whether the human data reflect general genomic patterns of metazoans, similar analyses were done for the nematode Caenorhabditis elegans. Although the C. elegans data generally concur with the human patterns, we identified fewer intron-bounded domains in this organism, consistent with the lower complexity of C. elegans genes. [The following individuals kindly provided reagents, samples, or unpublished information as indicated in the paper: Z. Gu and R. Stevens.]

Animals↗

Toward supportive data collection tools for plant metabolomics.

Over recent years, a number of initiatives have proposed standard reporting guidelines for functional genomics experiments. Associated with these are data models that may be used as the basis of the design of software tools that store and transmit experiment data in standard formats. Central to the success of such data handling tools is their usability. Successful data handling tools are expected to yield benefits in time saving and in quality assurance. Here, we describe the collection of datasets that conform to the recently proposed data model for plant metabolomics known as ArMet (architecture for metabolomics) and illustrate a number of approaches to robust data collection that have been developed in collaboration between software engineers and biologists. These examples also serve to validate ArMet from the data collection perspective by demonstrating that a range of software tools, supporting data recording and data upload to central databases, can be built using the data model as the basis of their design.

Arabidopsis↗

Genome-wide analysis of SPAK/OSR1 binding motifs.

Based on the alignment of 12 sequences of protein motifs that interact with the kinases SPAK (Ste20-related proline alanine-rich kinase) and OSR1 (oxidative stress response 1), we performed genome-wide searches of the sequence [S/G/V]RFx[V/I]xx[V/I/T/S]xx, where x represents any amino acid. The "Mus musculus" search resulted in the identification of 131 mouse proteins containing 137 SPAK/OSR1 putative binding motifs. Similar numbers were found for human, zebrafish, fruit fly, and worm. A little more than half of the mouse proteins containing SPAK/OSR1 binding domains (53%) were also identified in the human search, whereas approximately 17-18% of these common hits were identified in the zebrafish search. The mouse proteins could be divided into two broad categories: 2/3 had an identified function, whereas 1/3 were either predicted or of unknown function. The known proteins were grouped as transport proteins, other membrane proteins, kinases, phosphatases, cytoskeletal, ribosomal, nuclear, enzymes, and others. Analysis of the location of the SPAK/OSR1 binding motif within the protein sequence revealed distribution throughout the entire length, but with preference to the extreme amino- or carboxyl termini for a large number of proteins. Analysis of the amino acid composition of the motifs revealed a preponderance of serine residues at positions 5, 6, 7, and 8. In summary, our new search found and thus confirms the 12 proteins previously shown to interact with the kinases and identifies 119 potential new targets for SPAK and OSR1 in the mouse proteome.

Amino Acid Motifs↗

INTEGRATOR: interactive graphical search of large protein interactomes over the Web.

BACKGROUND: The rapid growth of protein interactome data has elevated the necessity and importance of network analysis tools. However, unlike pure text data, network search spaces are of exponential complexity. This poses special challenges for storing, searching, and navigating this data efficiently. Moreover, development of effective web interfaces has been difficult. RESULTS: We present Integrator, a web-integrated graphical search tool for protein-protein interaction networks across 50+ genomes. CONCLUSION: Integrator provides single and multiple protein searches of the Bioverse database containing experimentally-derived and predicted protein-protein interactions. The interface provides animated local network views, rapid subgraph manipulation, and cross-referencing of functional annotations. Integrator is available at http://bioverse.compbio.washington.edu/integrator.

Algorithms↗

A case study in pathway knowledgebase verification.

BACKGROUND: Biological databases and pathway knowledge-bases are proliferating rapidly. We are developing software tools for computer-aided hypothesis design and evaluation, and we would like our tools to take advantage of the information stored in these repositories. But before we can reliably use a pathway knowledge-base as a data source, we need to proofread it to ensure that it can fully support computer-aided information integration and inference. RESULTS: We design a series of logical tests to detect potential problems we might encounter using a particular knowledge-base, the Reactome database, with a particular computer-aided hypothesis evaluation tool, HyBrow. We develop an explicit formal language from the language implicit in the Reactome data format and specify a logic to evaluate models expressed using this language. We use the formalism of finite model theory in this work. We then use this logic to formulate tests for desirable properties (such as completeness, consistency, and well-formedness) for pathways stored in Reactome. We apply these tests to the publicly available Reactome releases (releases 10 through 14) and compare the results, which highlight Reactome's steady improvement in terms of decreasing inconsistencies. We also investigate and discuss Reactome's potential for supporting computer-aided inference tools. CONCLUSION: The case study described in this work demonstrates that it is possible to use our model theory based approach to identify problems one might encounter using a knowledge-base to support hypothesis evaluation tools. The methodology we use is general and is in no way restricted to the specific knowledge-base employed in this case study. Future application of this methodology will enable us to compare pathway resources with respect to the generic properties such resources will need to possess if they are to support automated reasoning.

Algorithms↗