Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

GIMS: an integrated data storage and analysis environment for genomic and functional data.

Effective analyses in functional genomics require access to many kinds of biological data. For example, the analysis of upregulated genes in a microarray experiment might be aided by information concerning protein interactions or proteins' cellular locations. However, such information is often stored in different formats at different sites, in ways that may not be amenable to integrated analysis. The Genome Information Management System (GIMS) is an object database that integrates genomic data with data on the transcriptome, protein-protein interactions, metabolic pathways and annotations, such as gene ontology terms and identifiers. The resulting system supports the running of analyses over this integrated data resource, and provides comprehensive facilities for handling and interrelating the results of these analyses. GIMS has been used to store Saccharomyces cerevisiae data, and we demonstrate how the integrated storage of diverse types of data can be beneficial for analysis, using combinations of complex queries. As an example, we describe how GIMS has been used to analyse a collection of aryl alcohol dehydrogenase gene deletion mutants. The GIMS database can be accessed remotely using a Java application that can be downloaded from http://img.cs.man.ac.uk/gims.

Computational Biology↗

The TIGR gene indices: reconstruction and representation of expressed gene sequences.

Expressed sequence tags (ESTs) have provided a first glimpse of the collection of transcribed sequences in a variety of organisms. However, a careful analysis of this sequence data can provide significant additional functional, structural and evolutionary information. Our analysis of the public EST sequences, available through the TIGR Gene Indices (TGI; http://www.tigr.org/tdb/tdb.html ), is an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed for selected organisms by first clustering, then assembling EST and annotated gene sequences from GenBank. This process produces a set of unique, high-fidelity virtual transcripts, or tentative consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, and to provide links between orthologous and paralogous genes.

Base Sequence↗

Predicted role for the archease protein family based on structural and sequence analysis of TM1083 and MTH1598, two proteins structurally characterized through structural genomics efforts.

Recently, the structures of two proteins belonging to the archease family, TM1083 from Thermotoga maritima and MTH1598 from Methanobacterium thermoautotrophicum, have been solved independently by two Protein Structure Initiative structural genomics pilot centers using X-ray crystallography and NMR, respectively. The archease protein family is a good example of one of the paradoxes of structural genomics: Approximately one third of protein structures produced by structural genomics centers have no known function and are still annotated as "hypothetical proteins" in the Protein Data Bank. In the case of archeases, despite the existence of two protein structures and abundant sequence information, there is still no function assigned to this protein family. Here, our group predicts, based on structural similarity, sequence conservation, and gene context analyses, that members of this protein family might function as chaperones or modulators of proteins involved in DNA/RNA processing. The conservation of genomic context for this protein family is constant from Archaea and Bacteria to humans, and suggests that unannotated open reading frames contiguous to them could be novel RNA/DNA binding proteins.

Amino Acid Motifs↗

Extracting knowledge from dynamics in gene expression.

Most investigations of coordinated gene expression have focused on identifying correlated expression patterns between genes by examining their normalized static expression levels. In this study, we focus on the dynamics of gene expression by seeking to identify correlated patterns of changes in genetic expression level. In doing so, we build upon methods developed in clinical informatics to detect temporal trends of laboratory and other clinical data. We construct relevance networks from Saccharomyces cerevisiae gene-expression dynamics data and find genes with related functional annotations grouped together. While some of these associations are also found using a standard expression level analysis, many are identified exclusively through the dynamic analysis. These results strongly suggest that the analysis of gene expression dynamics is a necessary and important tool for studying regulatory and other functional relationships among genes. The source code developed for this investigation is freely available to all non-commercial investigators by contacting the authors.

Cluster Analysis↗

Sequence organization and matrix attachment regions of the human serine protease inhibitor gene cluster at 14q32.1.

The human serine protease inhibitor (serpin) gene cluster at 14q32.1 is a useful model system for studying the regulation of gene activity and chromatin structure. We demonstrated previously that the six known serpin genes in this region were organized into two subclusters of three genes each that occupied approximately 370 kb of DNA. To more fully understand the genomic organization of this region, we annotated a 1-Mb sequence contig from data from the Genoscope sequencing consortium (http://www.genoscope.cns.fr/ ). We report that 11 different serpin genes reside within the 14q32.1 cluster, including two novel alpha1-antiproteinase-like gene sequences, a kallistatin-like sequence, and two recently identified serpins that had not been mapped previously to 14q32.1. The genomic regions proximal and distal to the serpin cluster contain a variety of unrelated gene sequences of diverse function. To gain insight into the chromatin organization of the region, sequences with putative nuclear matrix-binding potential were identified by using the MAR-Wiz algorithm, and these MAR-Wiz candidate sequences were tested for nuclear matrix-binding activity in vitro. Several differences between the MAR-Wiz predictions and the results of biochemical tests were observed. The genomic organization of the serpin gene cluster is discussed.

Base Composition↗

Descartes' fly: the geometry of genomic annotation.

The completion of the Drosophila melanogaster genome marks another significant milestone in the growth of sequence information. But it also contributes to the ever-widening gap between sequence information and biological knowledge. One important approach to reducing this gap is theoretical inference through computational technologies. Many computer programs have been designed to annotate genomic sequence information with biologically relevant information. Here, I suggest that all of these methods have a common structure in which the sequence fragments are "coordinated" by some method of description such as Hidden Markov models. The key to the algorithms lies in constructing the most efficient set of coordinates that allow extrapolation and interpolation from existing knowledge. Efficient extrapolation and interpolation are produced if the sequence fragments acquire a natural geometrical structure in the coordinated description. Finding such a coordinate frame is an inductive problem with no algorithmic solution. The greater part of the problem of genomic annotation lies in biological modeling of the data rather than in algorithmic improvements.

Animals↗

Early human development and the chief sources of information on staged human embryos.

In a brief historical survey, the importance of Wilhelm His, senior, to human embryology is emphasized. He provided the impetus to Mall to establish the Carnegie Embryological Collection, which serves as a 'Bureau of Standards' for early human development. The Carnegie system of 23 stages for the embryonic period proper (first 8 postovulatory wk) is outlined, and some common misusages are noted. Finally, because of the difficulty in tracking down data based on staged human embryos, an annotated list of more than 40 key references is provided.

Embryology↗

Neuroendocrinology of protochordates: insights from Ciona genomics.

The genome for two species of Ciona is available making these tunicates excellent models for studies on the evolution of the chordates. In this review most of the data is from Ciona intestinalis, as the annotation of the C. savignyi genome is not yet available. The phylogenetic position of tunicates at the origin of the chordates and the nature of the genome before expansion in vertebrates allows tunicates to be used as a touchstone for understanding genes that either preceded or arose in vertebrates. A comparison of Ciona, a sea squirt, to other model organisms such as a nematode, fruit fly, zebrafish, frog, chicken and mouse shows that Ciona has many useful traits including accessibility for embryological, lineage tracing, forward genetics, and loss- or gain-of-function experiments. For neuroendocrine studies, these traits are important for determining gene function, whereas the availability of the genome is critical for identification of ligands, receptors, transcription factors and signaling pathways. Four major neurohormones and their receptors have been identified by cloning and to some extent by function in Ciona: gonadotropin-releasing hormone, insulin, insulin-like growth factor, and cionin, a member of the CCK/gastrin family. The simplicity of tunicates should be an advantage in searching for novel functions for these hormones. Other neuroendocrine components that have been annotated in the genome are a multitude of receptors, which are available for cloning, expression and functional studies.

Animals↗

Genome-wide expression dynamics during mouse embryonic development reveal similarities to Drosophila development.

Gene transcription mediates many vital aspects of mammalian embryonic development. A comprehensive characterization and analysis of the dynamics of gene transcription in the embryo is therefore likely to provide significant insights into the basic mechanisms of this process. We used microarrays to map transcription in the mouse embryo in the important period from embryonic day 8 (e8.0) to postnatal day 1 (p1) during which the bulk of the differentiation and development of organ systems takes place. Analysis of these expression profiles revealed distinct patterns of gene expression which correlate with the differentiation of organs including the nervous system, liver, skin, lungs, and digestive system, among others. Statistical analysis of the data based on Gene Ontology (GO) group annotation showed that specific temporal sequence patterns in gene class utilization across development are very similar to patterns seen during the embryonic development of Drosophila, suggesting conservation of the temporal progression of these processes across 550 million years of evolution. The temporal profiles of gene expression and activation of processes revealed here provide intriguing insights into the mechanisms of mammalian development, embryogenesis, and organogenesis, as well as into the evolution of developmental processes.

Animals↗

Large-scale analysis of the human and mouse transcriptomes.

High-throughput gene expression profiling has become an important tool for investigating transcriptional activity in a variety of biological samples. To date, the vast majority of these experiments have focused on specific biological processes and perturbations. Here, we have generated and analyzed gene expression from a set of samples spanning a broad range of biological conditions. Specifically, we profiled gene expression from 91 human and mouse samples across a diverse array of tissues, organs, and cell lines. Because these samples predominantly come from the normal physiological state in the human and mouse, this dataset represents a preliminary, but substantial, description of the normal mammalian transcriptome. We have used this dataset to illustrate methods of mining these data, and to reveal insights into molecular and physiological gene function, mechanisms of transcriptional regulation, disease etiology, and comparative genomics. Finally, to allow the scientific community to use this resource, we have built a free and publicly accessible website (http://expression.gnf.org) that integrates data visualization and curation of current gene annotations.

Animals↗

The putative malate/lactate dehydrogenase from Pseudomonas putida is an NADPH-dependent delta1-piperideine-2-carboxylate/delta1-pyrroline-2-carboxylate reductase involved in the catabolism of D-lysine and D-proline.

A Pseudomonas putida ATCC12633 gene, dpkA, encoding a putative protein annotated as malate/L-lactate dehydrogenase in various sequence data bases was disrupted by homologous recombination. The resultant dpkA(-) mutant was deprived of the ability to use D-lysine and also D-proline as a sole carbon source. The dpkA gene was cloned and overexpressed in Escherichia coli, and the gene product was characterized. The enzyme showed neither malate dehydrogenase nor lactate dehydrogenase activity but catalyzed the NADPH-dependent reduction of such cyclic imines as Delta(1)-piperideine-2-carboxylate and Delta(1)-pyrroline-2-carboxylate to form L-pipecolate and L-proline, respectively. NADH also served as a hydrogen donor for both substrates, although the reaction rates were less than 1% of those with NADPH. The reverse reactions were also catalyzed by the enzyme but at much lower rates. Thus, the enzyme has dual metabolic functions, and we named the enzyme Delta(1)-piperideine-2-carboxylate/Delta(1)-pyrroline-2-carboxylate reductase, the first member of a novel subclass in a large family of NAD(P)-dependent oxidoreductases.

Bacterial Proteins↗

Compass of 47,787 cattle ESTs.

Comparative Mapping by Annotation and Sequence Similarity (COMPASS) has been demonstrated to be an effective approach for predicting the chromosome location of expressed sequence tags (ESTs) and other sequence-based markers on the basis of comparative mapping information. Herein, we describe the development and use of a computer program to execute the COMPASS strategy en masse. The program was used to identify orthologs and predict map locations of 47,787 cattle ESTs. Among these 47,787 ESTs, 30,097 had significant matches with sequences in the human UniGene database and 21,311 were annotated with human GB4 radiation hybrid mapping data. These sequences are contained within 9,956 and 6,295 individual human UniGene clusters, respectively. The putative human orthologs and predicted cattle chromosome locations of the 21,311 cattle ESTs with GB4 mapping data are provided in this report as a resource for the research community.

Animals↗

Characterization of a normalized cDNA library from bovine intestinal muscle and epithelial tissues.

Tissue-specific cDNA library sequences (expressed sequence tags, or EST) yield a detailed snapshot of gene expression and are useful in developing second-generation molecular resources (i.e., microarrays) for gene expression profiling. The objective of this study was to develop and characterize an intestine-specific cDNA library to examine the transcriptome of the bovine gut and identify expressed genes that influence ruminant nutrition and health. We describe BARC-8BOV, a normalized cDNA library developed from mRNA isolated from four distinct intestinal locations (duodenal, jejunal and ileal small intestine, colon) of Holstein dairy cattle resulting in 19,110 5'-EST deposited into the NCBI GenBank EST database. Assembly and clustering of these 19,110 clone sequences yielded 11,208 unique elements (3,419 contigs and 7,789 singletons) with an average length of 695 base pairs. Analysis strongly suggests normalization and tissue pooling were effective at increasing the discovery rate of new bovine sequence. A total of 1,123 sequence elements not previously identified in cattle, but with similarity to known genes in other animal species, were identified and shown to be involved in numerous critical biological processes. An additional 745 transcripts were not previously represented as EST in nucleotide or protein databases, and further analysis of these could lead to the identification of gut-specific transcript variants of known genes or potentially the discovery of novel bovine genes. Of the 11,208 assembled sequences, 11,034, or 98.4%, match sequences present in the bovine DNA trace archive at NCBI, and add to a bovine EST database previously lacking significant gut tissue representation. Ultimately, these data will also contribute in efforts to annotate the bovine genome.

Animals↗

Simultaneous modelling of metabolic, genetic and product-interaction networks.

The creation of cell models from annotated genome information, as well as additional data from other databases, requires both a format and medium for its distribution. Standards are described for the representation of the data in the form of Document Type Definitions (DTDs) for XML files. Separate DTDs are detailed for genetic, metabolic and gene product-interaction networks, which can be used to hold information on individual subsystems, or which may be combined to create a whole cell DTD. In the execution of this work, a fifth DTD was also created for a metabolite thesaurus, which allows incorporation of metabolite synonyms and generic nomenclature data into the models. A gene-regulation classification scheme was also created, to facilitate incorporation of gene regulatory information in an efficient manner. The work is described with particular reference to the metabolic network of Escherichia coli, which contains 808 individual enzymes. The assignment of confidence levels to these data, through the use of Gene Ontology evidence codes, is highlighted. In silico investigations may now be performed using the mathematical simulation workbench, DBsolve, which incorporates the facility to introduce data directly from XML.

Computational Biology↗

A System for Automated Bacterial (genome) Integrated Annotation--SABIA.

UNLABELLED: A web-based software suite, SABIA (System for Automated Bacterial Integrated Annotation), is described that provides a comprehensive computational support for the assembly and annotation of whole bacterial genomes from the data derived from sequencing projects. AVAILABILITY: Both SABIA and supplementary materials are available at http://www.sabia.lncc.br

Algorithms↗

The Enhanced Microbial Genomes Library.

Since the obtention of the complete sequence of Haemophilus influenzae Rd in 1995, the number of bacterial genomes entirely sequenced has regularly increased. A problem is that the quality of the annotations of these very large sequences is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of that size. In this context, we have decided to build the Enhanced Microbial Genomes Library (EMGLib) in which these two problems are alleviated. This library contains all the complete genomes from bacteria already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on World Wide Web servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib/emglib. html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

EMGLib: the enhanced microbial genomes library (update 2000).

As the number of complete microbial genomes publicly available is still growing, the problem of annotation quality in these very large sequences remains unsolved. Indeed, the number of annotations associated with complete genomes is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of such size. In this context, the Enhanced Microbial Genomes Library (EMGLib) was developed to try to alleviate these problems. This library contains all the complete genomes from prokaryotes (bacteria and archaea) already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on WWW servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib.html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

EyeSite: a semi-automated database of protein families in the eye.

The EyeSite is a web-based database of protein families for proteins that function in the eye and their homologous sequences. The resource clusters proteins at different levels of homology in order to facilitate functional annotation of sequences and modelling of proteins from structural homologues. Eye proteins are organized into the tissue types in which they function and are clustered into homologous families using a novel protocol employing the TribeMCL algorithm. Homologous families are further subdivided into sequence clusters for which multiple sequence alignments are generated. Structural annotations from the CATH domain database are provided for nearly 90% of the sequences, and protein family annotations from the Pfam database for approximately 86%. Homology models have also been generated where appropriate. The EyeSite is stored in a relational database and is extensively linked to other online bioinformatics resources to help relate allelic variants, annotations and clinical details to the derived data in the database. The EyeSite is available for online search, sequence information and model retrieval at http://eyesite.cryst.bbk.ac.uk/.

Amino Acid Sequence↗