Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

MIPS: a database for genomes and protein sequences.

The Munich Information Center for Protein Sequences (MIPS-GSF), Martinsried, near Munich, Germany, continues its longstanding tradition to develop and maintain high quality curated genome databases. In addition, efforts have been intensified to cover the wealth of complete genome sequences in a systematic, comprehensive form. Bioinformatics, supporting national as well as European sequencing and functional analysis projects, has resulted in several up-to-date genome-oriented databases. This report describes growing databases reflecting the progress of sequencing the Arabidopsis thaliana (MATDB) and Neurospora crassa genomes (MNCDB), the yeast genome database (MYGD) extended by functional analysis data, the database of annotated human EST-clusters (HIB) and the database of the complete cDNA sequences from the DHGP (German Human Genome Project). It also contains information on the up-to-date database of complete genomes (PEDANT), the classification of protein sequences (ProtFam) and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database. These databases can be accessed through the MIPS WWW server (http://www. mips.biochem.mpg.de).

Arabidopsis↗

Structure and action of urocanase.

Urocanase (EC 4.2.1.49) from Pseudomonas putida was crystallized after removing one of the seven free thiol groups. The crystal structure was solved by multiwavelength anomalous diffraction (MAD) using a seleno-methionine derivative and then refined at 1.14 A resolution. The enzyme is a symmetric homodimer of 2 x 557 amino acid residues with tightly bound NAD+ cofactors. Each subunit consists of a typical NAD-binding domain inserted into a larger core domain that forms the dimer interface. The core domain has a novel chain fold and accommodates the substrate urocanate in a surface depression. The NAD domain sits like a lid on the core domain depression and points with the nicotinamide group to the substrate. Substrate, nicotinamide and five water molecules are completely sequestered in a cavity. Most likely, one of these water molecules hydrates the substrate during catalysis. This cavity has to open for substrate passage, which probably means lifting the NAD domain. The observed atomic arrangement at the active center gives rise to a detailed proposal for the catalytic mechanism that is consistent with published chemical data. As expected, the variability of the residues involved is low, as derived from a family of 58 proteins annotated as urocanases in the data banks. However, one well-embedded member of this family showed a significant deviation at the active center indicating an incorrect annotation.

Amino Acid Sequence↗

The Genome Sequence DataBase version 1.0 (GSDB): from low pass sequences to complete genomes.

The Genome Sequence DataBase (GSDB) has completed its conversion to an improved relational database. The new database, GSDB 1.0, is fully operational and publicly available. Data contributions, including both original sequence submissions and community annotation, are being accomplished through the use of a graphical client-server interface tool, the GSDB Annotator, and via GIO (GSDB Input/Output) files. Data retrieval services are being provided through a new Web Query Tool and direct SQL. All methods of data contribution and data retrieval fully support the new data types that have been incorporated into GSDB, including discontiguous sequences, multiple sequence alignments, and community annotation.

Animals↗

High-quality peptide evidence for annotating non-canonical open reading frames as human proteins.

A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.

GENCODE↗

Adaptive algorithm of automated annotation.

MOTIVATION: It is common knowledge that the avalanche of data arriving from the sequencing projects cannot be annotated either experimentally or manually by experts. The need for a reliable and convenient tool for automated sequence annotation is broadly recognized. RESULTS: Here, we describe the Adaptive Algorithm of Automated Annotation (A(4)) based on a statistical approach to this problem. The mathematical model relates a set of homologous sequences and descriptions of their functional properties, and calculates the probabilities of transferring a sequence description onto its homologue. The proposed model is adaptive, its parameters (distribution characteristics, transference probabilities, thresholds, etc.) are dynamic, i.e. are generated individually for the sequences and various functional properties (words of the description). The proposed technique significantly outperforms the widely used test for frequency threshold, which is a special case of our model realized for the simplest set of parameters. The prediction technique has been realized as a computer program and tested on a random sequence sampling from SWISS-PROT. AVAILABILITY: The automated annotation program based on the proposed algorithm is available through the Web browser at http://www.genebee.msu.su/services/annot/basic.html.

Algorithms↗

Hippocrates: an integrated platform for telemedicine applications.

This paper describes 'Hippocrates', an integrated platform for telemedicine applications. Hippocrates allows computer supported co-operative work based on patient data folders consisting of selected diagnostic images, annotation text, patient history and other information. All data transferred is encrypted to ensure confidentiality and integrity. It operates on a local level over TCP/IP LAN environment and on a remote level over public ISDN lines.

Computer Security↗

GenomeRNAi: a database for cell-based RNAi phenotypes.

RNA interference (RNAi) has emerged as a powerful tool to generate loss-of-function phenotypes in a variety of organisms. Combined with the sequence information of almost completely annotated genomes, RNAi technologies have opened new avenues to conduct systematic genetic screens for every annotated gene in the genome. As increasing large datasets of RNAi-induced phenotypes become available, an important challenge remains the systematic integration and annotation of functional information. Genome-wide RNAi screens have been performed both in Caenorhabditis elegans and Drosophila for a variety of phenotypes and several RNAi libraries have become available to assess phenotypes for almost every gene in the genome. These screens were performed using different types of assays from visible phenotypes to focused transcriptional readouts and provide a rich data source for functional annotation across different species. The GenomeRNAi database provides access to published RNAi phenotypes obtained from cell-based screens and maps them to their genomic locus, including possible non-specific regions. The database also gives access to sequence information of RNAi probes used in various screens. It can be searched by phenotype, by gene, by RNAi probe or by sequence and is accessible at http://rnai.dkfz.de.

Animals↗

XcisClique: analysis of regulatory bicliques.

BACKGROUND: Modeling of cis-elements or regulatory motifs in promoter (upstream) regions of genes is a challenging computational problem. In this work, set of regulatory motifs simultaneously present in the promoters of a set of genes is modeled as a biclique in a suitably defined bipartite graph. A biologically meaningful co-occurrence of multiple cis-elements in a gene promoter is assessed by the combined analysis of genomic and gene expression data. Greater statistical significance is associated with a set of genes that shares a common set of regulatory motifs, while simultaneously exhibiting highly correlated gene expression under given experimental conditions. METHODS: XcisClique, the system developed in this work, is a comprehensive infrastructure that associates annotated genome and gene expression data, models known cis-elements as regular expressions, identifies maximal bicliques in a bipartite gene-motif graph; and ranks bicliques based on their computed statistical significance. Significance is a function of the probability of occurrence of those motifs in a biclique (a hypergeometric distribution), and on the new sum of absolute values statistic (SAV) that uses Spearman correlations of gene expression vectors. SAV is a statistic well-suited for this purpose as described in the discussion. RESULTS: XcisClique identifies new motif and gene combinations that might indicate as yet unidentified involvement of sets of genes in biological functions and processes. It currently supports Arabidopsis thaliana and can be adapted to other organisms, assuming the existence of annotated genomic sequences, suitable gene expression data, and identified regulatory motifs. A subset of Xcis Clique functionalities, including the motif visualization component MotifSee, source code, and supplementary material are available at https://bioinformatics.cs.vt.edu/xcisclique/.

Algorithms↗

Prediction of protein function using protein-protein interaction data.

Assigning functions to novel proteins is one of the most important problems in the post-genomic era. Several approaches have been applied to this problem, including analyzing gene expression patterns, phylogenetic profiles, protein fusions and protein-protein interactions. We develop a novel approach that applies the theory of Markov random fields to infer a protein's functions using protein-protein interaction data and the functional annotations of its interaction protein partners. For each function of interest and a protein, we predict the probability that the protein has that function using Bayesian approaches. Unlike in other available approaches for protein annotation where a protein has or does not have a function of interest, we give a probability for having the function. This probability indicates how confident we are about the prediction. We apply our method to predict cellular functions (43 categories including a category "others") for yeast proteins defined in the Yeast Proteome Database (YPD), using the protein-protein interaction data from the Munich Information Center for Protein Sequences (MIPS, http://mips.gsf.de). We show that our approach outperforms other available methods for function prediction based on protein interaction data.

Amino Acid Sequence↗

Comparative EST analyses in plant systems.

Expressed sequence tag (EST) data are a major contributor to the known plant sequence space. Organization of the data into non-redundant clusters representing tentative unique genes provides snapshots of the gene repertoires of a species. This chapter reviews availability of sequences and sequence analysis results and describes several resources and tools that should facilitate broad-based utilization of EST data for gene structure annotation, gene discovery, and comparative genomics.

Alternative Splicing↗

AGML Central: web based gel proteomic infrastructure.

SUMMARY: AGML Central is a web-based open-source public infrastructure for dissemination of two-dimensional Gel Electrophoresis (2-DE) proteomics data in AGML format (Annotated Gel Markup Language). It includes a growing collection of converters from proprietary formats such as those produced by PDQUEST (BioRad), PHORETIX 2-D (Nonlinear Dynamics) and Melanie (GenBio SA). The resulting unifying AGML formatted entry, with or without the raw gel images, is optionally stored in a database for future reference. AGML Central was developed to provide a common platform for data dissemination and development of 2-DE data analysis tools. This resource responds to an increasing use of AGML for 2-DE public source data representation which requires automated tools for conversion from proprietary formats. Conversion and short-term storage is made publicly available, permanent storage requires prior registering. A JAVA applet visualizer was developed to visualize the AGML data with cross-reference links. In order to facilitate automated access a SOAP web service is also included in the AGML Central infrastructure. AVAILABILITY: http://bioinformatics.musc.edu/agmlcentral.

Database Management Systems↗

A SNP-centric database for the investigation of the human genome.

BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are an increasingly important tool for genetic and biomedical research. Although current genomic databases contain information on several million SNPs and are growing at a very fast rate, the true value of a SNP in this context is a function of the quality of the annotations that characterize it. Retrieving and analyzing such data for a large number of SNPs often represents a major bottleneck in the design of large-scale association studies. DESCRIPTION: SNPper is a web-based application designed to facilitate the retrieval and use of human SNPs for high-throughput research purposes. It provides a rich local database generated by combining SNP data with the Human Genome sequence and with several other data sources, and offers the user a variety of querying, visualization and data export tools. In this paper we describe the structure and organization of the SNPper database, we review the available data export and visualization options, and we describe how the architecture of SNPper and its specialized data structures support high-volume SNP analysis. CONCLUSIONS: The rich annotation database and the powerful data manipulation and presentation facilities it offers make SNPper a very useful online resource for SNP research. Its success proves the great need for integrated and interoperable resources in the field of computational biology, and shows how such systems may play a critical role in supporting the large-scale computational analysis of our genome.

Databases, Genetic↗

Arabidopsis thaliana proteomics: from proteome to genome.

Proteomics has become an important approach for investigating cellular processes and network functions. Significant improvements have been made during the last few years in technologies for high-throughput proteomics, both at the level of data analysis software and mass spectrometry hardware. As proteomics technologies advance and become more widely accessible, efforts of cataloguing and quantifying full proteomes are underway to complement other genomics approaches, such as RNA and metabolite profiling. Of particular interest is the application of proteome data to improve genome annotation and to include information on post-translational protein modifications with the annotation of the corresponding gene. This type of analysis requires a paradigm shift because amino acid sequences must be assigned to peptides without relying on existing protein databases. In this review, advances and current limitations of full proteome analysis are briefly highlighted using the model plant Arabidopsis thaliana as an example. Strategies to identify peptides are also discussed on the basis of MS/MS data in a protein database-independent approach.

Arabidopsis↗

IML: An image markup language.

Image Markup Language is an extensible markup language (XML) schema used to describe both image metadata and annotations. It describes both data pertaining to an entire image, and data that are tied to specific regions or features of the image. Developed for a specific domain in Medical Education, this pa-per describes extensions to take advantage of the Dublin Core metadata standard, and of an XML schema for vector graphics representation. We have developed a prototype system of open source tools implementing an authoring system, a client system, and an image annotation database which can be queried though the Web.

Diagnostic Imaging↗

Gene ontology application to genomic functional annotation, statistical analysis and knowledge mining.

While a massive amount of biomolecular information is increasingly accumulating in different databanks, on the other hand high-throughput technologies are generating a great quantity of data that need to be annotated with the genomic information available, and interpreted. To this aim, the use of specific ontologies can greatly help either in integrating different information stored within heterogeneous databanks, or in identifying and clustering sequence data sharing common characteristics. In the molecular biology domain, the Gene Ontology (GO) is the most developed and widely used ontology. To demonstrate its great utility in the annotation and biological interpretation of gene sets obtained by means of high-throughput experiments, we implemented the web application here described. It enables functional annotations of a given gene set on a genomic scale and across different species. Within our application the annotations provided by the GO vocabulary allow either to easily bind several information from different resources, or to cluster annotated genes according to their biological characteristics. Through the GO structure it is also possible to represent biological concepts with different specificity levels, from very general to very precise concepts. Furthermore, the statistical evaluation of the categorizations provided by the GO annotations enables to highlight the most significant biological characteristics of a gene set, and therefore to mine knowledge from data. Our created tool meets the need to manage a vast quantity of biological data with a simple user interface adapt also for users with limited informatics knowledge, leading them to evaluate the functional significance of experiment's results with graphical views and statistical indexes in a well-known web browser user interface.

Genomics↗

Prediction of protein function using protein-protein interaction data.

Assigning functions to novel proteins is one of the most important problems in the postgenomic era. Several approaches have been applied to this problem, including the analysis of gene expression patterns, phylogenetic profiles, protein fusions, and protein-protein interactions. In this paper, we develop a novel approach that employs the theory of Markov random fields to infer a protein's functions using protein-protein interaction data and the functional annotations of protein's interaction partners. For each function of interest and protein, we predict the probability that the protein has such function using Bayesian approaches. Unlike other available approaches for protein annotation in which a protein has or does not have a function of interest, we give a probability for having the function. This probability indicates how confident we are about the prediction. We employ our method to predict protein functions based on "biochemical function," "subcellular location," and "cellular role" for yeast proteins defined in the Yeast Proteome Database (YPD, www.incyte.com), using the protein-protein interaction data from the Munich Information Center for Protein Sequences (MIPS, mips.gsf.de). We show that our approach outperforms other available methods for function prediction based on protein interaction data. The supplementary data is available at www-hto.usc.edu/~msms/ProteinFunction.

Bayes Theorem↗

SPINE bioinformatics and data-management aspects of high-throughput structural biology.

SPINE (Structural Proteomics In Europe) was established in 2002 as an integrated research project to develop new methods and technologies for high-throughput structural biology. Development areas were broken down into workpackages and this article gives an overview of ongoing activity in the bioinformatics workpackage. Developments cover target selection, target registration, wet and dry laboratory data management and structure annotation as they pertain to high-throughput studies. Some individual projects and developments are discussed in detail, while those that are covered elsewhere in this issue are treated more briefly. In particular, this overview focuses on the infrastructure of the software that allows the experimentalist to move projects through different areas that are crucial to high-throughput studies, leading to the collation of large data sets which are managed and eventually archived and/or deposited.

Computational Biology↗

An integrated analysis and database system for full-length cDNA.

Annotation and database system of full-length cDNA sequences was developed. As the components of the system, ORF annotation system, functional annotation system based on database search results, mapping annotation system, and integrated retrieval and display system were developed. In the ORF annotation system integrated analyses using conventional tools are performed and useful retrieval interface using motif list are introduced. In the functional annotation system based on database search results, a new method that characterizes a given unknown cDNA was developed by using a profile of similarity level over words appearing in sequence database entries. In the mapping annotation system, we linked by similarity searches full-length cDNA sequences with database DNA sequences that are already mapped on chromosomes. By using these links, full-length cDNAs can be retrieved by the retrieval condition of physical mapping information. Genetic disease information mapped on the physical mapping site can also be displayed by this system. Furthermore, we constructed an integrated database system for these analyzed data, and thus enabled annotation and selection of full-length cDNAs from points of both gene function and mapping information.

Chromosome Mapping↗