Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

The Mammalian Phenotype Ontology as a tool for annotating, analyzing and comparing phenotypic information.

The Mammalian Phenotype (MP) Ontology enables robust annotation of mammalian phenotypes in the context of mutations, quantitative trait loci and strains that are used as models of human biology and disease. The MP Ontology supports different levels and richness of phenotypic knowledge and flexible annotations to individual genotypes. It continues to develop dynamically via collaborative input from research groups, mutagenesis consortia, and biological domain experts. The MP Ontology is currently used by the Mouse Genome Database and Rat Genome Database to represent phenotypic data.

Animals↗

The UCSC Genome Browser Database.

The University of California Santa Cruz (UCSC) Genome Browser Database is an up to date source for genome sequence data integrated with a large collection of related annotations. The database is optimized to support fast interactive performance with the web-based UCSC Genome Browser, a tool built on top of the database for rapid visualization and querying of the data at many levels. The annotations for a given genome are displayed in the browser as a series of tracks aligned with the genomic sequence. Sequence data and annotations may also be viewed in a text-based tabular format or downloaded as tab-delimited flat files. The Genome Browser Database, browsing tools and downloadable data files can all be found on the UCSC Genome Bioinformatics website (http://genome.ucsc.edu), which also contains links to documentation and related technical information.

Animals↗

Creating the gene ontology resource: design and implementation.

The exponential growth in the volume of accessible biological information has generated a confusion of voices surrounding the annotation of molecular information about genes and their products. The Gene Ontology (GO) project seeks to provide a set of structured vocabularies for specific biological domains that can be used to describe gene products in any organism. This work includes building three extensive ontologies to describe molecular function, biological process, and cellular component, and providing a community database resource that supports the use of these ontologies. The GO Consortium was initiated by scientists associated with three model organism databases: SGD, the Saccharomyces Genome database; FlyBase, the Drosophila genome database; and MGD/GXD, the Mouse Genome Informatics databases. Additional model organism database groups are joining the project. Each of these model organism information systems is annotating genes and gene products using GO vocabulary terms and incorporating these annotations into their respective model organism databases. Each database contributes its annotation files to a shared GO data resource accessible to the public at http://www.geneontology.org/. The GO site can be used by the community both to recover the GO vocabularies and to access the annotated gene product data sets from the model organism databases. The GO Consortium supports the development of the GO database resource and provides tools enabling curators and researchers to query and manipulate the vocabularies. We believe that the shared development of this molecular annotation resource will contribute to the unification of biological information.

Animals↗

The Genome Sequence DataBase version 1.0 (GSDB): from low pass sequences to complete genomes.

The Genome Sequence DataBase (GSDB) has completed its conversion to an improved relational database. The new database, GSDB 1.0, is fully operational and publicly available. Data contributions, including both original sequence submissions and community annotation, are being accomplished through the use of a graphical client-server interface tool, the GSDB Annotator, and via GIO (GSDB Input/Output) files. Data retrieval services are being provided through a new Web Query Tool and direct SQL. All methods of data contribution and data retrieval fully support the new data types that have been incorporated into GSDB, including discontiguous sequences, multiple sequence alignments, and community annotation.

Animals↗

The UCSC Genome Browser Database: update 2006.

The University of California Santa Cruz Genome Browser Database (GBD) contains sequence and annotation data for the genomes of about a dozen vertebrate species and several major model organisms. Genome annotations typically include assembly data, sequence composition, genes and gene predictions, mRNA and expressed sequence tag evidence, comparative genomics, regulation, expression and variation data. The database is optimized to support fast interactive performance with web tools that provide powerful visualization and querying capabilities for mining the data. The Genome Browser displays a wide variety of annotations at all scales from single nucleotide level up to a full chromosome. The Table Browser provides direct access to the database tables and sequence data, enabling complex queries on genome-wide datasets. The Proteome Browser graphically displays protein properties. The Gene Sorter allows filtering and comparison of genes by several metrics including expression data and several gene properties. BLAT and In Silico PCR search for sequences in entire genomes in seconds. These tools are highly integrated and provide many hyperlinks to other databases and websites. The GBD, browsing tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Amino Acid Sequence↗

The UCSC genome browser database: update 2007.

The University of California, Santa Cruz Genome Browser Database contains, as of September 2006, sequence and annotation data for the genomes of 13 vertebrate and 19 invertebrate species. The Genome Browser displays a wide variety of annotations at all scales from the single nucleotide level up to a full chromosome and includes assembly data, genes and gene predictions, mRNA and EST alignments, and comparative genomics, regulation, expression and variation data. The database is optimized for fast interactive performance with web tools that provide powerful visualization and querying capabilities for mining the data. In the past year, 22 new assemblies and several new sets of human variation annotation have been released. New features include VisiGene, a fully integrated in situ hybridization image browser; phyloGif, for drawing evolutionary tree diagrams; a redesigned Custom Track feature; an expanded SNP annotation track; and many new display options. The Genome Browser, other tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Animals↗

Synteny-defined candidate genes for congenital and idiopathic scoliosis.

Idiopathic scoliosis (IS) is a common but poorly understood syndrome. Congenital scoliosis (CS) is less common but comparably unexplored. Previous studies suggest that each has a significant genetic component. However, the occurrence of scoliosis in the presence of other hereditary connective tissue syndromes raises the possibility that IS and CS are in fact a heterogeneous group of disorders with varied pathogenetic mechanisms. Mouse mutations have proven informative in identifying genes that are important in the development of the musculoskeletal system and provided important mechanistic insights regarding their roles in human disease. We sought to identify candidate genes for human IS and CS by reviewing mouse mutations with phenotypes affecting the axial skeleton. We performed a systematic review using the Mouse Genome Database (MGD), the Genome Database (GDB), and the Online Mendelian Inheritance in Man (OMIM) world-wide-web sites with additional searches performed based on the results of this initial search. We identified approximately 400 mouse mutations, reviewed approximately 250 of these for vertebral phenotypes, assessed 45 of these for synteny conservation between mouse and man, and identified 28 mouse mutations for which 29 credible candidates for human scoliosis could be identified based on mouse phenotypic and mapping data. For each of these, we have synthesized information about the mouse mutant phenotype, mapping data, information regarding molecular pathogenesis when a specific causative gene has been identified, and information regarding plausible candidates based on map position when the causative gene has not been identified. Among these were three loci for which the mutant gene had been identified and the human homologue was known. Some of the mouse mutants have phenotypes similar to human syndromes.

Animals↗

The institute for genomic research Osa1 rice genome annotation database.

We have developed a rice (Oryza sativa) genome annotation database (Osa1) that provides structural and functional annotation for this emerging model species. Using the sequence of O. sativa subsp. japonica cv Nipponbare from the International Rice Genome Sequencing Project, pseudomolecules, or virtual contigs, of the 12 rice chromosomes were constructed. Our most recent release, version 3, represents our third build of the pseudomolecules and is composed of 98% finished sequence. Genes were identified using a series of computational methods developed for Arabidopsis (Arabidopsis thaliana) that were modified for use with the rice genome. In release 3 of our annotation, we identified 57,915 genes, of which 14,196 are related to transposable elements. Of these 43,719 non-transposable element-related genes, 18,545 (42.4%) were annotated with a putative function, 5,777 (13.2%) were annotated as encoding an expressed protein with no known function, and the remaining 19,397 (44.4%) were annotated as encoding a hypothetical protein. Multiple splice forms (5,873) were detected for 2,538 genes, resulting in a total of 61,250 gene models in the rice genome. We incorporated experimental evidence into 18,252 gene models to improve the quality of the structural annotation. A series of functional data types has been annotated for the rice genome that includes alignment with genetic markers, assignment of gene ontologies, identification of flanking sequence tags, alignment with homologs from related species, and syntenic mapping with other cereal species. All structural and functional annotation data are available through interactive search and display windows as well as through download of flat files. To integrate the data with other genome projects, the annotation data are available through a Distributed Annotation System and a Genome Browser. All data can be obtained through the project Web pages at http://rice.tigr.org.

Computational Biology↗

PathwayVoyager: pathway mapping using the Kyoto Encyclopedia of Genes and Genomes (KEGG) database.

BACKGROUND: Equally important and challenging as genome annotation, is the subsequent classification of predicted genes into their respective pathways. The Kyoto Encyclopedia of Genes and Genomes (KEGG) represents a database consisting of known genes and their respective biochemical functionalities. Although accessible online, analyses of multiple genes are time consuming and are not suitable for analyzing data sets that are proprietary. RESULTS: Presented here is a new software solution that utilizes the KEGG online database for pathway mapping of partial and whole prokaryotic genomes. PathwayVoyager retrieves user-defined subsets of the KEGG database and stores the data as local, blast-formatted databases. Previously selected datasets can be re-used, reducing run-time significantly. Whole or partial genomes can be automatically analyzed using NCBI's BlastP algorithm and ORFs with similarities below the user-defined threshold will be marked on pathway maps. Multiple gene hits are sorted by similarity. Since no sequence information is transmitted over the Internet, PathwayVoyager is an ideal solution for pathway mapping and reconstruction of confidential DNA sequence data. CONCLUSION: PathwayVoyager represents an alternative approach to many already existing, more complex pathway reconstructions software solutions. This software does not require any dedicated hardware or software and is flexible and straightforward to use. It is ideally suited for environments where analyses on variable datasets are desired.

Computational Biology↗

DictyMOLD-a Dictyostelium discoideum genome browser database.

UNLABELLED: With the Dictyostelium Genome Project nearing completion, we initiated the construction of a data repository for all Dictyostelium discoideum genomic data. Up to now this database, called DictyMOLD (Dicty Map Of Linked Data), incorporates the recently completed D.discoideum chromosomes 1 and 2 sequences together with related annotations. To visualise maps, sequences and annotations and to provide access for the scientific community a perl-based browser was developed. AVAILABILITY: The DictyMOLD database is freely accessible via http://genome.imb-jena.de/dictyostelium/ CONTACT: gernot@imb-jena.de.

Animals↗

Making High-level Queries on Diverse Genome Data: A Structured Genome Document Database System Based on GXML and GQL.

Complete DNA sequences (genomes) and associated data are being made available worldwide at an astonishing rate. Through computer analysis of such data, molecular biologists hope to gain an overall understanding of the genome, such as by predicting large-scale gene networks. However, this is difficult because diverse genome data are scattered across many highly heterogeneous databases, and because existing database systems lack the facilities to expose and analyze functional relationships among the data. To address these problems, we propose a new type of genome database system. Since a genome can be thought of intuitively as a kind of 'document', our system uses a structured document language based on XML to effectively represent genomes and associated data. The information-rich structures of the genome documents help cope with data diversity and heterogeneity. A powerful query language is introduced that exposes important biological relationships among the genome data. We have obtained favorable results from several experiments, demonstrating the usefulness of our method in building a top-down view of genome functionality.

Journal Article↗

Multiple ribonuclease H-encoding genes in the Caenorhabditis elegans genome contrasts with the two typical ribonuclease H-encoding genes in the human genome.

Database searches of the Caenorhabditis elegans and human genomic DNA sequences revealed genes encoding ribonuclease H1 (RNase H1) and RNase H2 in each genome. The human genome contains a single copy of each gene, whereas C. elegans has four genes encoding RNase H1-related proteins and one gene for RNase H2. By analyzing the mRNAs produced from the C. elegans genes, examining the amino acid sequence of the predicted protein, and expressing the proteins in Esherichia coli we have identified two active RNase H1-like proteins. One is similar to other eukaryotic RNases H1, whereas the second RNase H (rnh-1.1) is unique. The rnh-1.0 gene is transcribed as a dicistronic message with three dsRNA-binding domains; the mature mRNA is transspliced with SL2 splice leader and contains only one dsRNA-binding domain. Formation of RNase H1 is further regulated by differential cis-splicing events. A single rnh-2 gene, encoding a protein similar to several other eukaryotic RNase H2L's, also has been examined. The diversity and enzymatic properties of RNase H homologues are other examples of expansion of protein families in C. elegans. The presence of two RNases H1 in C. elegans suggests that two enzymes are required in this rather simple organism to perform the functions that are accomplished by a single enzyme in more complex organisms. Phylogenetic analysis indicates that the active C. elegans RNases H1 are distantly related to one another and that the C. elegans RNase H1 is more closely related to the human RNase H1. The database searches also suggest that RNase H domains of LTR-retrotransposons in C. elegans are quite unrelated to cellular RNases H1, but numerous RNase H domains of human endogenous retroviruses are more closely related to cellular RNases H.

Amino Acid Sequence↗

Chromosomal mapping and zygosity check of transgenes based on flanking genome sequences determined by genomic walking.

Transgenes can affect transgenic mice via transgene expression or via the so-called positional effect. DNA sequences can be localized in chromosomes using recently established mouse genomic databases. In this study, we describe a chromosomal mapping method that uses the genomic walking technique to analyze genomic sequences that flank transgenes, in combination with mouse genome database searches. Genomic DNA was collected from two transgenic mouse lines harboring pCAGGS-based transgenes, and adaptor-ligated, enzyme restricted genomic libraries for each mouse line were constructed. Flanking sequences were determined by sequencing amplicons obtained by PCR amplification of genomic libraries with transgene-specific and adaptor primers. The insertion positions of the transgenes were located by BLAST searches of the Ensembl genome database using the flanking sequences of the transgenes, and the transgenes of the two transgenic mouse lines were mapped onto chromosomes 11 and 3. In addition, flanking sequence information was used to construct flanking primers for a zygosity check. The zygosity (homozygous transgenic, hemizygous transgenic and non-transgenic) of animals could be identified by differential band formation in PCR analyses with the flanking primers. These methods should prove useful for genetic quality control of transgenic animals, even though the mode of transgene integration and the specificity of flanking sequences needs to be taken into account.

5' Flanking Region↗

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Databases, Genetic↗

TcruziDB, an integrated database, and the WWW information server for the Trypanosoma cruzi genome project.

Data analysis, presentation and distribution is of utmost importance to a genome project. A public domain software, ACeDB, has been chosen as the common basis for parasite genome databases, and a first release of TcruziDB, the Trypanosoma cruzi genome database, is available by ftp from ftp://iris.dbbm.fiocruz.br/pub/genomedb/Tcr uziDB as well as versions of the software for different operating systems (ftp://iris.dbbm.fiocruz.br/pub/unixsoft/). Moreover, data originated from the project are available from the WWW server at http://www.dbbm.fiocruz.br. It contains biological and parasitological data on CL Brener, its karyotype, all available T. cruzi sequences from Genbank, data on the EST-sequencing project and on available libraries, a T. cruzi codon table and a listing of activities and participating groups in the genome project, as well as meeting reports. T. cruzi discussion lists (tcruzil@iris.dbbm.fiocruz.br and tcgenics@iris.dbbm.fiocruz.br) are being maintained for communication and to promote collaboration in the genome project.

Animals↗

A model system for studying the integration of molecular biology databases.

MOTIVATION: Integration of molecular biology databases remains limited in practice despite its practical importance and considerable research effort. The complexity of the problem is such that an experimental approach is mandatory, yet this very complexity makes it hard to design definitive experiments. This dilemma is common in science, and one tried-and-true strategy is to work with model systems. We propose a model system for this problem, namely a database of genes integrating diverse data across organisms, and describe an experiment using this model. RESULTS: We attempted to construct a database of human and mouse genes integrating data from GenBank and the human and mouse genome-databases. We discovered numerous errors in these well-respected databases: approximately 15% of genes are apparently missing from the genome-databases; links between the sequence and genome-databases are missing for another 5-10% of the cases; about a third of likely homology links are missing between the genome-databases; 10-20% of entries classified as 'genes' are apparently misclassified. By using a model system, we were able to study the problems caused by anomalous data without having to face all the hard problems of database integration. CONTACT: nat@jax.org

Animals↗

GOBASE--a database of organelle and bacterial genome information.

The organelle genome database GOBASE is now in its twelfth release, and includes 350,000 mitochondrial sequences and 118,000 chloroplast sequences, roughly a 3-fold expansion since previously documented. GOBASE also includes a fully reannotated genome sequence of Rickettsia prowazekii, one of the closest bacterial relatives of mitochondria, and will shortly expand to contain more data from bacteria from which organelles originated. All these sequences are now accessible through a single unified interface. Enhancements to the functionality of GOBASE include addition of pages for RNA structures and a page compiling data about the taxonomic distribution of organelle-encoded genes; incorporation of Gene Ontology terms; addition of features deduced from incomplete annotations to sequences in GenBank; marking of type examples in cases where single genes in single species are oversampled within GenBank; and addition of graphics illustrating gene structure and the position of neighbouring genes on a sequence. The database has been reimplemented in PostgreSQL to facilitate development and maintenance, and structural modifications have been made to speed up queries, particularly those related to taxonomy. The GOBASE database can be queried at http://gobase.bcm.umontreal.ca/ and inquiries should be directed to gobase@bch.umontreal.ca.

Chloroplasts↗

The University of Minnesota Biocatalysis/Biodegradation database: microorganisms, genomics and prediction.

The University of Minnesota Biocatalysis/Biodegradation Database (http://www.labmed.umn.edu/umbbd/ ) begins its fifth year having met its initial goals. It contains approximately 100 pathways for microbial catabolic metabolism of primarily xenobiotic organic compounds, including information on approximately 650 reactions, 600 compounds and 400 enzymes, and containing approximately 250 microorganism entries. It includes information on most known microbial catabolic reaction types and the organic functional groups they transform. Having reached its first goals, it is ready to move beyond them. It is poised to grow in many different ways, including mirror sites; fold prediction for its sequenced enzymes; closer ties to genome and microbial strain databases; and the prediction of biodegradation pathways for compounds it does not contain.

Biodegradation, Environmental↗