Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Adding semantics to genome databases: towards an ontology for molecular biology.

Molecular biology has a communication problem. There are many databases using their own labels and categories for storing data objects and some using identical labels and categories but with a different meaning. Conversely, one concept is often found under different names. Prominent examples are the concepts "gene" and "protein sequence" which are used with different semantics by major international genomic and protein databases thereby making database integration difficult and strenuous. This situation can only be improved by either defining individual semantic interfaces between each pair of databases (complexity of order n2) or by implementing one agreeable, transparent and computationally tractable semantic repository and linking each database to it (complexity of order n). Ontologies are one means to provide such semantic repository by explicitly specifying the meaning of and relation between the fundamental concepts in an application domain. Here, heuristics for building an ontology and the upper level and a database branch of a prospective Ontology for Molecular Biology are presented and compared to other ontologies with respect to suitability for molecular biology (http:/(/)igd.rz-berlin.mpg.de/www/oe/mbo.html).

Computational Biology↗

u-Genome: a database on genome design in unicellular genomes.

Unicellular eukaryotes were among the first ones to be selected for complete genome sequencing because of the small size of their genomes and their interactions with humans and a broad range of animals and plants. Currently, ten completely sequenced unicellular genome sequences have been publicly released and as the number of available unicellular genomes increases, comparative genomics analysis within this group of organisms becomes more and more instructive. However, such an analysis is difficult to carry out without a suitable platform gathering not only the original annotations but also relevant information available in public databases or obtained by applying common bioinformatics methods. With the aim of solving these difficulties, we have developed a web-accessible database named u-Genome, the unicellular genome design database. The database is unique in featuring three datasets namely (1) orthologous proteins (2) paralogous proteins and (3) statistical distributions on exons, introns, intergenic DNA and correlations between them. A tool, Uniview, designed to visualize the gene structures for individual genes in the genome is also integrated. This database is of importance in understanding unicellular genome design and architecture and evolution related studies. The database is available through a web interface at http://sege.ntu.edu.sg/wester/ugenome.

Animals↗

The PlantsP and PlantsT Functional Genomics Databases.

PlantsP and PlantsT allow users to quickly gain a global understanding of plant phosphoproteins and plant membrane transporters, respectively, from evolutionary relationships to biochemical function as well as a deep understanding of the molecular biology of individual genes and their products. As one database with two functionally different web interfaces, PlantsP and PlantsT are curated plant-specific databases that combine sequence-derived information with experimental functional-genomics data. PlantsP focuses on proteins involved in the phosphorylation process (i.e., kinases and phosphatases), whereas PlantsT focuses on membrane transport proteins. Experimentally, PlantsP provides a resource for information on a collection of T-DNA insertion mutants (knockouts) in each kinase and phosphatase, primarily in Arabidopsis thaliana, and PlantsT uniquely combines experimental data regarding mineral composition (derived from inductively coupled plasma atomic emission spectroscopy) of mutant and wild-type strains. Both databases provide extensive information on motifs and domains, detailed information contributed by individual experts in their respective fields, and descriptive information drawn directly from the literature. The databases incorporate a unique user annotation and review feature aimed at acquiring expert annotation directly from the plant biology community. PlantsP is available at http://plantsp.sdsc.edu and PlantsT is available at http://plantst.sdsc.edu.

Arabidopsis↗

Requirements and standards for organelle genome databases.

Mitochondria and plastids (collectively called organelles) descended from prokaryotes that adopted an intracellular, endosymbiotic lifestyle within early eukaryotes. Comparisons of their remnant genomes address a wide variety of biological questions, especially when including the genomes of their prokaryotic relatives and the many genes transferred to the eukaryotic nucleus during the transitions from endosymbiont to organelle. The pace of producing complete organellar genome sequences now makes it unfeasible to do broad comparisons using the primary literature and, even if it were feasible, it is now becoming uncommon for journals to accept detailed descriptions of genome-level features. Unfortunately, no database is completely useful for this task, since they have little standardization and are riddled with error. Further, the descriptors necessary to make full use of these data are generally lacking. Here, I outline what is currently wrong and what must be done to make this data useful to the scientific community.

Animals↗

Leveraging genomic databases: from an Aedes albopictus mosquito cell line to the malaria vector Anopheles gambiae via the Drosophila genome project.

An important justification for genome sequencing efforts is the anticipation that data from model organisms will provide a framework for the more rapid analysis of other, less studied genomes. In this investigation, we sequenced an internal region of 25 amino acids from a 52 kDa protein that was differentially expressed in 20-hydroxyecdysone-treated Aedes albopictus cells in culture. Within the GenBank non-mouse and non-human expressed sequence tag (EST) database, this "Aedes peptide" uncovered a putative homology to hypothetical translation products from Anopheles gambiae, Caenorhabditis elegans and Drosophila melanogaster. The hypothetical translation product from D. melanogaster, which included 462 amino acids, uncovered five expressed sequence tags (ESTs) from the malaria vector, Anopheles gambiae. When the Anopheles ESTs were aligned against the hypothetical Drosophila protein, we found that in aggregate they covered 324 amino acids, with gaps measuring 19, 30, and 87 amino acids. To approximate the complete amino acid sequence, gaps between translation products from Anopheles ESTs were replaced with corresponding amino acids from Drosophila to arrive at a calculated mass of 51 104 and a pI of 5.84 for the mosquito protein, consistent with the position of the Ae. albopictus protein on two-dimensional polyacrylamide gels. Finally, tandem mass spectrometry of a tryptic digest of the 52 kDa Ae. albopictus protein revealed 33 peptides with masses within 1 Dalton of those predicted from an in silico digestion of the reconstructed Anophleles protein. In addition to providing the first direct evidence that a hypothetical protein in Drosophila is in fact translated, this analysis provides a general approach for maximizing recovery, from existing databases, of information that can facilitate prioritization of efforts among several candidate proteins.

Aedes↗

GrainGenes, the genome database for small-grain crops.

GrainGenes, http://www.graingenes.org, is the international database for the wheat, barley, rye and oat genomes. For these species it is the primary repository for information about genetic maps, mapping probes and primers, genes, alleles and QTLs. Documentation includes such data as primer sequences, polymorphism descriptions, genotype and trait scoring data, experimental protocols used, and photographs of marker polymorphisms, disease symptoms and mutant phenotypes. These data, curated with the help of many members of the research community, are integrated with sequence and bibliographic records selected from external databases and results of BLAST searches of the ESTs. Records are linked to corresponding records in other important databases, e.g. Gramene's EST homologies to rice BAC/PACs, TIGR's Gene Indices and GenBank. In addition to this information within the GrainGenes database itself, the GrainGenes homepage at http://wheat.pw.usda.gov provides many other community resources including publications (the annual newsletters for wheat, barley and oat, monographs and articles), individual datasets (mapping and QTL studies, polymorphism surveys, variety performance evaluations), specialized databases (Triticeae repeat sequences, EST unigene sets) and pages to facilitate coordination of cooperative research efforts in specific areas such as SNP development, EST-SSRs and taxonomy. The goal is to serve as a central point for obtaining and contributing information about the genetics and biology of these cereal crops.

Alleles↗

Microbial genomic database of the Yangtze River, the third-longest river on Earth.

Microbes play an important role in mediating the nutrient cycling in the river ecosystem as a hotspot for biogeochemical processes. Due to scattered sampling efforts, however, there is a lack of a systematic study of the diversity of prokaryotic genomes in the Yangtze River, the third longest river on Earth. Here, we collected 602 metagenomic datasets of water, sediment and riparian soil samples spanning the Upper, Middle, and Lower basins of the Yangtze River over a 6,300 km continuum. We reconstructed 8,110 qualified genomes represented by 927 species-level genomes at the 95% ANI threshold, spanning 31 bacterial and five archaeal phyla. We further showed that more than half of these species (61.3% ~ 82.4%) were novel according to the genomic comparison against the curated databases, greatly expanding the known diversity of river prokaryotes. This dataset depicts an overview of microbial genomic diversity in the Yangtze River and provides a resource for in-depth investigation of metabolic potential, ecology, and evolution of riverine microbiomes.

Rivers↗

MITOMAP: a human mitochondrial genome database.

We have developed a comprehensive database (MITOMAP) for the human mitochondrial DNA (mtDNA), the first component of the human genome to be completely sequenced [Anderson et al. (1981) Nature 290, 457-465]. MITOMAP uses the mtDNA sequence as the unifying element for bringing together information on mitochondrial genome structure and function, pathogenic mutations and their clinical characteristics, population associated variation, and gene- gene interactions. As increasingly larger regions of the human genome are sequenced and characterized, the need for integrating such information will grow. Consequently, MITOMAP not only provides a valuable reference for the mitochondrial biologist, it may also provide a model for the development of information storage and retrieval systems for other components of the human genome.

Amino Acid Sequence↗

EchoBASE: an integrated post-genomic database for Escherichia coli.

EchoBASE (http://www.ecoli-york.org) is a relational database designed to contain and manipulate information from post-genomic experiments using the model bacterium Escherichia coli K-12. Its aim is to collate information from a wide range of sources to provide clues to the functions of the approximately 1500 gene products that have no confirmed cellular function. The database is built on an enhanced annotation of the updated genome sequence of strain MG1655 and the association of experimental data with the E.coli genes and their products. Experiments that can be held within EchoBASE include proteomics studies, microarray data, protein-protein interaction data, structural data and bioinformatics studies. EchoBASE also contains annotated information on 'orphan' enzyme activities from this microbe to aid characterization of the proteins that catalyse these elusive biochemical reactions.

Databases, Genetic↗

Computational method for temporal pattern discovery in biomedical genomic databases.

With the rapid growth of biomedical research databases, opportunities for scientific inquiry have expanded quickly and led to a demand for computational methods that can extract biologically relevant patterns among vast amounts of data. A significant challenge is identifying temporal relationships among genotypic and clinical (phenotypic) data. Few software tools are available for such pattern matching, and they are not interoperable with existing databases. We are developing and validating a novel software method for temporal pattern discovery in biomedical genomics. In this paper, we present an efficient and flexible query algorithm (called TEMF) to extract statistical patterns from time-oriented relational databases. We show that TEMF - as an extension to our modular temporal querying application (Chronus II) - can express a wide range of complex temporal aggregations without the need for data processing in a statistical software package. We show the expressivity of TEMF using example queries from the Stanford HIV Database.

Artificial Intelligence↗

The Ensembl genome database project.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of the human genome sequence, with confirmed gene predictions that have been integrated with external data sources, and is available as either an interactive web site or as flat files. It is also an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements from sequence analysis to data storage and visualisation. The Ensembl site is one of the leading sources of human genome sequence annotation and provided much of the analysis for publication by the international human genome project of the draft genome. The Ensembl system is being installed around the world in both companies and academic sites on machines ranging from supercomputers to laptops.

Computational Biology↗

CyanoBase, the genome database for Synechocystis sp. strain PCC6803: status for the year 2000.

CyanoBase provides an online resource for access to data on genomic information about the cyanobacterium Synechocystis sp. strain PCC6803. The database contains annotations for each protein-coding gene deduced from the entire nucleotide sequence of the genome, gene classification lists, and keyword and similarity search engines. Core portions of CyanoBase consist of annotations for each of the 3168 protein genes deduced from the entire nucleotide sequence of this genome. The contents of each gene were improved by updating with the results of similarity searches and by introducing references for analysis in bioinformatics. The database now contains repository facilities that store and provide experimental information, in addition to providing proposals for the function of each gene. This information should help to avoid unnecessary, overlapping experiments and should assist communication between scientists who wish to elucidate the function of putative genes on the cyanobacteria genome. The current URL of CyanoBase is http://www.kazusa.or.jp:8080/cyano/

Cyanobacteria↗

MBGD: microbial genome database for comparative analysis.

MBGD is a workbench system for comparative analysis of completely sequenced microbial genomes. The central function of MBGD is to create an orthologous gene classification table using precomputed all-against-all similarity relationships among genes in multiple genomes. In MBGD, an automated classification algorithm has been implemented so that users can create their own classification table by specifying a set of organisms and parameters. This feature is especially useful when the user's interest is focused on some taxonomically related organisms. The created classification table is stored into the database and can be explored combining with the data of individual genomes as well as similarity relationships among genomes. Using these data, users can carry out comparative analyses from various points of view, such as phylogenetic pattern analysis, gene order comparison and detailed gene structure comparison. MBGD is accessible at http://mbgd.genome.ad.jp/.

Algorithms↗

[Construction of rice dwarf virus genome database].

Secondary database construction is an important subject in the field of bioinformatics. As the full genomic sequences of some organisms are being completed and followed by structural and functional studies, construction of secondary database becomes essential on the agenda. The rice dwarf virus (RDV) is a pathogen infecting rice in China, Japan and the Southeastern Asia region and leading to considerable economic loss. Based on the data generated from recent genomic research and earlier biochemical studies scattered in various primary databases and scientific journals, we have constructed a compact, user-friendly and non-redundant job-oriented secondary database. This work will provide compiled useful information for plant molecular biologists as well as in achieving preliminary experiences in secondary database construction.

Databases, Factual↗

Mining proteases in the genome databases.

Protease data mining can take advantage both of the many specialist, Web-available databases that cover the genetic, protein and nucleic acid sequence information that is specific to a variety of organisms, and of a flexible, but defined, classification system. However, precomputed data, such as gene predictions, should be used with care. Unless there is definitive supporting information, ideally sequencing of a cDNA to show that the predictions are accurate, followed by expression and biochemical characterization of the predicted protein, the predicted gene and its product remains a possibility, rather than a certainty.

Animals↗

Preparing a human membrane and secreted protein-enriched cDNA library using PCR primers derived from a genomic database.

We describe here a strategy for preparing a human membrane and secreted protein (MSP)-enriched cDNA library based on human MSP- and non-MSP-encoding cDNA sequences in the databases. The signal peptide parts of the MSP-encoding cDNA sequences, which currently comprise about half of the estimated total number in humans, were analyzed for common patterns. These patterns form a 'minimal' set of polymerase chain reaction primer candidates of length varying from 9 to 21 nt. The products stemming from each primer candidate were determined and the results allowed us to obtain an 'optimal' mixed-length primer set. Ninety-six percent of the primers in this set were predicted to yield </=10% undesired products, and the desired MSP-cDNA products could be easily separated by gel electrophoresis. The present analysis establishes a methodology for preparing a cDNA library that enables the analysis of individual MSPs. This methodology may also help identify new MSPs. As many cell regulatory processes are mediated by secreted proteins and their membrane-bound receptors, the preparation of a MSP-enriched cDNA library should benefit research on MSPs.

Amino Acid Sequence↗