Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Integrated analysis of the genome and the transcriptome by FANTOM.

The key to reliable annotation of a mammalian genome is broad characterisation of the transcriptional output, the transcriptome. FANTOM, the functional annotation of mouse cDNA, is a large-scale analysis of both the genome and the transcriptome of the mouse. In the early days of this work, the transcripts were characterised using our sophisticated methods. After the timely release of the first draft of mouse genome sequences, interesting information was obtained by its integration with these one-by-one annotations. Moreover, each transcript included its expression profile. Here, the two integrated annotation methods used by FANTOM are reviewed: one-by-one and categorised. One-by-one annotation refers to naming carried out based on well-known transcripts or its fragments using the top-down-style pipeline developed mostly by the FANTOM project. Categorised annotation, which refers to transcript grouping, not only helps naming of unknown transcripts, but will be the most utilised method for integration of the genome and the transcriptome from now on.

Abstracting and Indexing↗

Representing, storing and accessing molecular interaction data: a review of models and tools.

One important aim within systems biology is to integrate disparate pieces of information, leading to discovery of higher-level knowledge about important functionality within living organisms. This makes standards for representation of data and technology for exchange and integration of data important key points for development within the area. In this article, we focus on the recent developments within the field. We compare the recent updates to the three standard representations for exchange of data SBML, PSI MI and BioPAX. In addition, we give an overview of available tools for these three standards and a discussion on how these developments support possibilities for data exchange and integration.

Computational Biology↗

The genomic threading database.

UNLABELLED: The Genomic Threading Database currently contains structural annotations for the genomes of over 100 recently sequenced organisms. Annotations are carried out by using our modified GenTHREADER software and through implementing grid technology. AVAILABILITY: http://bioinf.cs.ucl.ac.uk/GTD

Database Management Systems↗

Beyond the clause: extraction of phosphorylation information from medline abstracts.

MOTIVATION: Phosphorylation is an important biochemical reaction that plays a critical role in signal transduction pathways and cell-cycle processes. A text mining system to extract the phosphorylation relation from the literature is reported. The focus of this paper is on the new methods developed and implemented to connect and merge pieces of information about phosphorylation mentioned in different sentences in the text. The effectiveness and accuracy of the system as a whole as well as that of the methods for extraction beyond a clause/sentence is evaluated using an independently annotated dataset, the Phospho.ELM database. The new methods developed to merge pieces of information from different sentences are shown to be effective in significantly raising the recall without much difference in precision.

Artificial Intelligence↗

A query language for biological networks.

MOTIVATION: Many areas of modern biology are concerned with the management, storage, visualization, comparison and analysis of networks, but no appropriate query language for such complex data structures yet exists. RESULTS: We have designed and implemented the pathway query language (PQL) for querying large protein interaction or pathway databases. PQL is based on a simple graph data model with extensions reflecting properties of biological objects. Queries match subgraphs in the database based on node properties and paths between nodes. The syntax is easy to learn for anybody familiar with SQL. As an important feature, a query may require a certain structure in the database to exist but return a different subgraph. We have tested PQL queries on networks of up to 16,000 nodes and found it to scale very well. AVAILABILITY: The code is available on request from the author.

Computational Biology↗

Automated genome annotation and pathway identification using the KEGG Orthology (KO) as a controlled vocabulary.

MOTIVATION: High-throughput technologies such as DNA sequencing and microarrays have created the need for automated annotation of large sets of genes, including whole genomes, and automated identification of pathways. Ontologies, such as the popular Gene Ontology (GO), provide a common controlled vocabulary for these types of automated analysis. Yet, while GO offers tremendous value, it also has certain limitations such as the lack of direct association with pathways. RESULTS: We demonstrated the use of the KEGG Orthology (KO), part of the KEGG suite of resources, as an alternative controlled vocabulary for automated annotation and pathway identification. We developed a KO-Based Annotation System (KOBAS) that can automatically annotate a set of sequences with KO terms and identify both the most frequent and the statistically significantly enriched pathways. Results from both whole genome and microarray gene cluster annotations with KOBAS are comparable and complementary to known annotations. KOBAS is a freely available stand-alone Python program that can contribute significantly to genome annotation and microarray analysis.

Artificial Intelligence↗

MEPD: a resource for medaka gene expression patterns.

The Medaka Expression Pattern Database (MEPD) is a database for gene expression patterns determined by in situ hybridization in the small freshwater fish medaka (Oryzias latipes). Data have been collected from various research groups and MEPD is developing into a central expression pattern depository within the medaka community. Gene expression patterns are described by images and terms of a detailed medaka anatomy ontology of over 4000 terms, which we have developed for this purpose and submitted to Open Biological Ontologies. Sequences have been annotated via BLAST match results and using Gene Ontology terms. These new features will facilitate data analyses using bioinformatics approaches and allow cross-species comparisons of gene expression patterns. Presently, MEPD has 19,757 entries, for 1024 of them the expression pattern has been determined.

Animals↗

Increasing confidence of protein interactomes using network topological metrics.

MOTIVATION: Experimental limitations in high-throughput protein-protein interaction detection methods have resulted in low quality interaction datasets that contained sizable fractions of false positives and false negatives. Small-scale, focused experiments are then needed to complement the high-throughput methods to extract true protein interactions. However, the naturally vast interactomes would require much more scalable approaches. RESULTS: We describe a novel method called IRAP* as a computational complement for repurification of the highly erroneous experimentally derived protein interactomes. Our method involves an iterative process of removing interactions that are confidently identified as false positives and adding interactions detected as false negatives into the interactomes. Identification of both false positives and false negatives are performed in IRAP* using interaction confidence measures based on network topological metrics. Potential false positives are identified amongst the detected interactions as those with very low computed confidence values, while potential false negatives are discovered as the undetected interactions with high computed confidence values. Our results from applying IRAP* on large-scale interaction datasets generated by the popular yeast-two-hybrid assays for yeast, fruit fly and worm showed that the computationally repurified interaction datasets contained potentially lower fractions of false positive and false negative errors based on functional homogeneity. AVAILABILITY: The confidence indices for PPIs in yeast, fruit fly and worm as computed by our method can be found at our website http://www.comp.nus.edu.sg/~chenjin/fpfn.

Animals↗

Analysis and prediction of mammalian protein glycation.

Glycation is a nonenzymatic process in which proteins react with reducing sugar molecules and thereby impair the function and change the characteristics of the proteins. Glycation is involved in diabetes and aging where the accumulation of glycation products causes side effects. In this study, we statistically investigate the glycation of epsilon amino groups of lysines and also train a sequence-based predictor. The statistical analysis suggests that acidic amino acids, mainly glutamate, and lysine residues catalyze the glycation of nearby lysines. The catalytic acidic amino acids are found mainly C-terminally from the glycation site, whereas the basic lysine residues are found mainly N-terminally. The predictor was made by combining 60 artificial neural networks in a balloting procedure. The cross-validated Matthews correlation coefficient for the predictor is 0.58, which is quite impressive given the relatively small amount of experimental data available. The method is made available at www.cbs.dtu.dk/services/NetGlycate-1.0.

Animals↗

SMART 4.0: towards genomic data integration.

SMART (Simple Modular Architecture Research Tool) is a web tool (http://smart.embl.de/) for the identification and annotation of protein domains, and provides a platform for the comparative study of complex domain architectures in genes and proteins. The January 2004 release of SMART contains 685 protein domains. New developments in SMART are centred on the integration of data from completed metazoan genomes. SMART now uses predicted proteins from complete genomes in its source sequence databases, and integrates these with predictions of orthology. New visualization tools have been developed to allow analysis of gene intron-exon structure within the context of protein domain structure, and to align these displays to provide schematic comparisons of orthologous genes, or multiple transcripts from the same gene. Other improvements include the ability to query SMART by Gene Ontology terms, improved structure database searching and batch retrieval of multiple entries.

Algorithms↗

THGS: a web-based database of Transmembrane Helices in Genome Sequences.

Transmembrane Helices in Genome Sequences (THGS) is an interactive web-based database, developed to search the transmembrane helices in the user-interested gene sequences available in the Genome Database (GDB). The proposed database has provision to search sequence motifs in transmembrane and globular proteins. In addition, the motif can be searched in the other sequence databases (Swiss-Prot and PIR) or in the macromolecular structure database, Protein Data Bank (PDB). Further, the 3D structure of the corresponding queried motif, if it is available in the solved protein structures deposited in the Protein Data Bank, can also be visualized using the widely used graphics package RASMOL. All the sequence databases used in the present work are updated frequently and hence the results produced are up to date. The database THGS is freely available via the world wide web and can be accessed at http:// pranag.physics.iisc.ernet.in/thgs/ or http://144.16. 71.10/thgs/.

Animals↗

The effect of experimental resolution on the performance of knowledge-based discriminatory functions for protein structure selection.

The key to an accurate method of protein structure prediction is the development of an effective discriminatory function. Knowledge-based discriminatory functions extract parameters from statistical analysis of experimentally determined protein structures. We assess how the quality of the protein structures used for compiling statistics affects the performance of a residue-specific all-atom probability discriminatory function (RAPDF). We find that the discriminatory power correlates with the quality of the structural dataset on which the RAPDF is parameterized in a statistically significant manner. The overrepresentation of unfavorable contacts in the low-resolution and NMR structures contributes to the major errors in the compilation of the conditional probabilities. Such errors weaken the discriminatory power of the function, especially when decoy conformations also contain considerable numbers of unfavorable contacts. This indicates that using high-resolution structural datasets after filtering out unfavorable contacts can improve the performance of knowledge-based discriminatory functions.

Computational Biology↗

Structural genomics: computational methods for structure analysis.

The success of structural genomics initiatives requires the development and application of tools for structure analysis, prediction, and annotation. In this paper we review recent developments in these areas; specifically structure alignment, the detection of remote homologs and analogs, homology modeling and the use of structures to predict function. We also discuss various rationales for structural genomics initiatives. These include the structure-based clustering of sequence space and genome-wide function assignment. It is also argued that structural genomics can be integrated into more traditional biological research if specific biological questions are included in target selection strategies.

Amino Acid Motifs↗

Rationally selected basis proteins: a new approach to selecting proteins for spectroscopic secondary structure analysis.

Protein basis sets have been extensively used as reference data for the determination of protein structure with optical methods such as circular dichroism and infrared spectroscopies. We have taken a new approach to basis protein selection by utilizing three crystal structure classification databases: CATH, SCOP, and PDB_SELECT. Through the use of the information available in these and other online resources, we identified 115 commercially available proteins as potential basis set candidates. By carefully screening the quality of the crystal structures and commercial protein preparations, we obtained a final set of 50 rationally selected proteins (RaSP50) that has been optimized for use in spectroscopic protein structure determination studies. These proteins span the full range of known protein folds as well as alpha-helix and beta-sheet contents, and they represent a more comprehensive variety of fold types than any previous reference set. This report includes a detailed presentation of the reasoning behind the rational protein selection process, a description of the properties of the RaSP50 set, and a discussion of the types of structural and spectral variations that are represented in the set.

Circular Dichroism↗

Identification and characterization of subfamily-specific signatures in a large protein superfamily by a hidden Markov model approach.

BACKGROUND: Most profile and motif databases strive to classify protein sequences into a broad spectrum of protein families. The next step of such database studies should include the development of classification systems capable of distinguishing between subfamilies within a structurally and functionally diverse superfamily. This would be helpful in elucidating sequence-structure-function relationships of proteins. RESULTS: Here, we present a method to diagnose sequences into subfamilies by employing hidden Markov models (HMMs) to find windows of residues that are distinct among subfamilies (called signatures). The method starts with a multiple sequence alignment (MSA) of the subfamily. Then, we build a HMM database representing all sliding windows of the MSA of a fixed size. Finally, we construct a HMM histogram of the matches of each sliding window in the entire superfamily. To illustrate the efficacy of the method, we have applied the analysis to find subfamily signatures in two well-studied superfamilies: the cadherin and the EF-hand protein superfamilies. As a corollary, the HMM histograms of the analyzed subfamilies revealed information about their Ca2+ binding sites and loops. CONCLUSIONS: The method is used to create HMM databases to diagnose subfamilies of protein superfamilies that complement broad profile and motif databases such as BLOCKS, PROSITE, Pfam, SMART, PRINTS and InterPro.

Binding Sites↗

An SVM-based system for predicting protein subnuclear localizations.

BACKGROUND: The large gap between the number of protein sequences in databases and the number of functionally characterized proteins calls for the development of a fast computational tool for the prediction of subnuclear and subcellular localizations generally applicable to protein sequences. The information on localization may reveal the molecular function of novel proteins, in addition to providing insight on the biological pathways in which they function. The bulk of past work has been focused on protein subcellular localizations. Furthermore, no specific tool has been dedicated to prediction at the subnuclear level, despite its high importance. In order to design a suitable predictive system, the extraction of subtle sequence signals that can discriminate among proteins with different subnuclear localizations is the key. RESULTS: New kernel functions used in a support vector machine (SVM) learning model are introduced for the measurement of sequence similarity. The k-peptide vectors are first mapped by a matrix of high-scored pairs of k-peptides which are measured by BLOSUM62 scores. The kernels, measuring the similarity for sequences, are then defined on the mapped vectors. By combining these new encoding methods, a multi-class classification system for the prediction of protein subnuclear localizations is established for the first time. The performance of the system is evaluated with a set of proteins collected in the Nuclear Protein Database (NPD). The overall accuracy of prediction for 6 localizations is about 50% (vs. random prediction 16.7%) for single localization proteins in the leave-one-out cross-validation; and 65% for an independent set of multi-localization proteins. This integrated system can be accessed at http://array.bioengr.uic.edu/subnuclear.htm. CONCLUSION: The integrated system benefits from the combination of predictions from several SVMs based on selected encoding methods. Finally, the predictive power of the system is expected to improve as more proteins with known subnuclear localizations become available.

Algorithms↗

MeMo: a hybrid SQL/XML approach to metabolomic data management for functional genomics.

BACKGROUND: The genome sequencing projects have shown our limited knowledge regarding gene function, e.g. S. cerevisiae has 5-6,000 genes of which nearly 1,000 have an uncertain function. Their gross influence on the behaviour of the cell can be observed using large-scale metabolomic studies. The metabolomic data produced need to be structured and annotated in a machine-usable form to facilitate the exploration of the hidden links between the genes and their functions. DESCRIPTION: MeMo is a formal model for representing metabolomic data and the associated metadata. Two predominant platforms (SQL and XML) are used to encode the model. MeMo has been implemented as a relational database using a hybrid approach combining the advantages of the two technologies. It represents a practical solution for handling the sheer volume and complexity of the metabolomic data effectively and efficiently. The MeMo model and the associated software are available at http://dbkgroup.org/memo/. CONCLUSION: The maturity of relational database technology is used to support efficient data processing. The scalability and self-descriptiveness of XML are used to simplify the relational schema and facilitate the extensibility of the model necessitated by the creation of new experimental techniques. Special consideration is given to data integration issues as part of the systems biology agenda. MeMo has been physically integrated and cross-linked to related metabolomic and genomic databases. Semantic integration with other relevant databases has been supported through ontological annotation. Compatibility with other data formats is supported by automatic conversion.

Computer Simulation↗