Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Biological database design and implementation.

We present our experience of building biological databases. Such databases have most aspects in common with other complex databases in other fields. We do not believe that biological data are that different from complex data in other fields. Our experience has led us to emphasise simplicity and conservative technology choices when building these databases. This is a short paper of advice that we hope is useful to people designing their own biological database.

Database Management Systems↗

Advances in the Exon-Intron Database (EID).

Investigation of exon-intron gene structures is a non-trivial task due to enormous expansions of the eukaryotic genomes, great variety of gene forms, and the imperfectness in sequence data. A number of available informational systems on various gene characteristics complement each other and are indispensable for many genomic studies. Among them, the Exon-Intron Database (EID) is a good choice for large-scale computational examination of exon/intron structure and splicing. It has many internal filters that control for sequence quality, consistency of gene descriptions, accordance to standards, and possible errors. New innovations in EID are described. The collection of exons and introns has been extended beyond coding regions and current versions of EID contain data on untranslated regions of gene sequences as well. Intron-less genes are included as a special part of EID. For species with entirely sequenced genomes, species-specific databases have been generated. A novel Mammalian Orthologous Intron Database (MOID) has been introduced which includes the full set of introns that come from orthologous genes that have the same positions relative to the reading frames. Examples of statistical analyses of gene sequences using EID are provided. We present the latest data on our comparison of intron positions in 11,025 orthologous genes of human, mouse and rat, and find no convincing cases of intron gain. We discuss relevant data-quality issues of genomic databases. In particular, 5% of genes in genomic databases contain internal stop codons. This fact is due to a combination of biological reasons and also to errors in sequence annotations. The EID is freely available at www.meduohio.edu/bioinfo/eid/.

Base Sequence↗

The Z curve database: a graphic representation of genome sequences.

MOTIVATION: Genome projects for many prokaryotic and eukaryotic species have been completed and more new genome projects are being underway currently. The availability of a large number of genomic sequences for researchers creates a need to find graphic tools to study genomes in a perceivable form. The Z curve is one of such tools available for visualizing genomes. The Z curve is a unique three-dimensional curve representation for a given DNA sequence in the sense that each can be uniquely reconstructed given the other. The Z curve database for more than 1000 genomes have been established here. RESULTS: The database contains the Z curves for archaea, bacteria, eukaryota, organelles, phages, plasmids, viroids and viruses, whose genomic sequences are currently available. All the 3-dimensional Z curves and their three component curves are stored in the database. The applications of the Z curve database on comparative genomics, gene prediction, computation of G+C content with a windowless technique, prediction of replication origins and terminations of bacterial and archaeal genomes and study of local deviations from the Chargaff Parity Rule 2 etc. are presented in detail. The Z curve database reported here is a treasure trove in which biologists could find useful biological knowledge.

Animals↗

The MUSC DNA Microarray Database.

SUMMARY: The Medical University of South Carolina (MUSC) DNA Microarray Database is a web-accessible archive of DNA microarray data. The database was developed using the DNA microarray project/data management system, micro ArrayDB. Annotations for each DNA microarray project and associated cRNA target information are stored in a MySQL relational database and linked to array hybridization data (raw and normalized). At the discretion of investigators, data are placed into the public domain where they can be interrogated and downloaded through a web browser. In addition to serving as an online resource of gene expression data, the MUSC DNA Microarray Database is a model for other academic DNA microarray data repositories. AVAILABILITY: Browsing and downloading of MUSC DNA Microarray Database information can be done after registration at http://proteogenomics.musc.edu/pss/home.php.

Academic Medical Centers↗

Light-weight integration of molecular biological databases.

MOTIVATION: Due to the increasing number of molecular biological databases and the exponential growth of their contents, database integration is an important topic of research in bioinformatics. Existing approaches in this area have in common that considerable efforts are needed to provide integrated access to heterogeneous data sources. RESULTS: This article describes the LIMBO architecture as a light-weight approach to molecular biological database integration. By building systems upon this architecture, the efforts needed for database integration can be significantly lowered. AVAILABILITY: As an illustration of the principle usefulness of the underlying ideas, a prototypical implementation based upon the LIMBO architecture is described. This implementation is exclusively based on freely available open source components like the PostgreSQL database management system and the BioRuby project. Additional files and modified components are available upon request from the author.

Computational Biology↗

ICBS: a database of interactions between protein chains mediated by beta-sheet formation.

MOTIVATION: Interchain beta-sheet (ICBS) interactions occur widely in protein quaternary structures, interactions between proteins and protein aggregation. These interactions play a central role in many biological processes and in diseases ranging from AIDS and cancer to anthrax and Alzheimer's. RESULTS: We have created a comprehensive database of ICBS interactions that is updated on a weekly basis and allows entries to be sorted and searched by relevance and other criteria through a simple Web interface. We derive a simple ICBS index to quantify the relative contributions of the beta-ladders in the overall interchain interaction and compute first- and second-order statistics regarding amino acid composition and pairing at different relative positions in the beta-strands. Analysis of the database reveals a 15.8% prevalence of significant ICBS interactions, the majority of which involve the formation of antiparallel beta-sheets and many of which involve the formation of dimers and oligomers. The frequencies of amino acids in ICBS interfaces are similar to those in intrachain beta-sheet interfaces. A full range of non-covalent interactions between side chains complement the hydrogen-bonding interactions between the main chains. Polar amino acids pair preferentially with polar amino acids and non-polar amino acids pair preferentially with non-polar amino acids among antiparallel (i, j) pairs. We anticipate that the statistics and insights gained from the database will guide the development of agents that control interchain beta-sheet interactions and that the database will help identify new protein interactions and targets for these agents. AVAILABILITY: The database is available at: http://www.igb.uci.edu/servers/icbs/

Binding Sites↗

POINT: a database for the prediction of protein-protein interactions based on the orthologous interactome.

One possible path towards understanding the biological function of a target protein is through the discovery of how it interfaces within protein-protein interaction networks. The goal of this study was to create a virtual protein-protein interaction model using the concepts of orthologous conservation (or interologs) to elucidate the interacting networks of a particular target protein. POINT (the prediction of interactome database) is a functional database for the prediction of the human protein-protein interactome based on available orthologous interactome datasets. POINT integrates several publicly accessible databases, with emphasis placed on the extraction of a large quantity of mouse, fruit fly, worm and yeast protein-protein interactions datasets from the Database of Interacting Proteins (DIP), followed by conversion of them into a predicted human interactome. In addition, protein-protein interactions require both temporal synchronicity and precise spatial proximity. POINT therefore also incorporates correlated mRNA expression clusters obtained from cell cycle microarray databases and subcellular localization from Gene Ontology to further pinpoint the likelihood of biological relevance of each predicted interacting sets of protein partners.

Animals↗

CSB.DB: a comprehensive systems-biology database.

SUMMARY: The open access comprehensive systems-biology database (CSB.DB) presents the results of bio-statistical analyses on gene expression data in association with additional biochemical and physiological knowledge. The main aim of this database platform is to provide tools that support insight into life's complexity pyramid with a special focus on the integration of data from transcript and metabolite profiling experiments. The central part of CSB.DB, which we describe in this applications note, is a set of co-response databases that currently focus on the three key model organisms, Escherichia coli, Saccharomyces cerevisiae and Arabidopsis thaliana. CSB.DB gives easy access to the results of large-scale co-response analyses, which are currently based exclusively on the publicly available compendia of transcript profiles. By scanning for the best co-responses among changing transcript levels, CSB.DB allows to infer hypotheses on the functional interaction of genes. These hypotheses are novel and not accessible through analysis of sequence homology. The database enables the search for pairs of genes and larger units of genes, which are under common transcriptional control. In addition, statistical tools are offered to the user, which allow validation and comparison of those co-responses that were discovered by gene queries performed on the currently available set of pre-selectable datasets. AVAILABILITY: All co-response databases can be accessed through the CSB.DB Web server (http://csbdb.mpimp-golm.mpg.de/).

Database Management Systems↗

General framework for developing and evaluating database scoring algorithms using the TANDEM search engine.

MOTIVATION: Tandem mass spectrometry (MS/MS) identifies protein sequences using database search engines, at the core of which is a score that measures the similarity between peptide MS/MS spectra and a protein sequence database. The TANDEM application was developed as a freely available database search engine for the proteomics research community. To extend TANDEM as a platform for further research on developing improved database scoring methods, we modified the software to allow users to redefine the scoring function and replace the native TANDEM scoring function while leaving the remaining core application intact. Redefinition is performed at run time so multiple scoring functions are available to be selected and applied from a single search engine binary. We introduce the implementation of the pluggable scoring algorithm and also provide implementations of two TANDEM compatible scoring functions, one previously described scoring function compatible with PeptideProphet and one very simple scoring function that quantitative researchers may use to begin their development. This extension builds on the open-source TANDEM project and will facilitate research into and dissemination of novel algorithms for matching MS/MS spectra to peptide sequences. The pluggable scoring schema is also compatible with related search applications P3 and Hunter, which are part of the X! suite of database matching algorithms. The pluggable scores and the X! suite of applications are all written in C++. AVAILABILITY: Source code for the scoring functions is available from http://proteomics.fhcrc.org

Algorithms↗

IndexToolkit: an open source toolbox to index protein databases for high-throughput proteomics.

UNLABELLED: A software package, IndexToolkit, aimed at overcoming the disadvantage of FASTA-format databases for frequent searching, is developed to utilize an indexing strategy to substantially accelerate sequence queries. IndexToolkit includes user-friendly tools and an Application Programming Interface (API) to facilitate indexing, storage and retrieval of protein sequence databases. As open source, it provides a sequence-retrieval developing framework, which is easily extensible for high-speed-request proteomic applications, such as database searching or modification discovering. We applied IndexToolkit to database searching engine pFind to demonstrate its effect. Experimental studies show that IndexToolkit is able to support significantly faster searches of protein database. AVAILABILITY: The IndexToolkit is free to use under the open source GNU GPL license. The source code and the compiled binary can be freely accessed through the website http://pfind.jdl.ac.cn/IndexToolkit. In this website, the more detailed information including screenshots and documentations for users and developers is also available.

Database Management Systems↗

A prototype object database for mitochondrial DNA variation.

Surveys of biochemical and molecular genetic variation in natural populations have generated a wealth of data, but this valuable resource has not been adequately preserved. We hope to prevent further loss by establishing a community database for population genetic surveys. We explored the feasibility of a population genetics database by developing a prototype for animal mitochondrial DNA (mtDNA) surveys. This prototype includes the specification of a format for data files that are to be submitted to the database, an open-source object database that encapsulates data with methods to display and analyze data, and a website where data can be retrieved in either its original form or extensible markup language (XML). Data from more than 50 published surveys of mtDNA variation were retrieved from the literature and entered into the database. We hope that the population genetics community will support this project by contributing both data and expertise.

Animals↗

HOWDY: an integrated database system for human genome research.

HOWDY is an integrated database system for accessing and analyzing human genomic information (http://www-alis.tokyo.jst.go.jp/HOWDY/). HOWDY stores information about relationships between genetic objects and the data extracted from a number of databases. HOWDY consists of an Internet accessible user interface that allows thorough searching of the human genomic databases using the gene symbols and their aliases. It also permits flexible editing of the sequence data. The database can be searched using simple words and the search can be restricted to a specific cytogenetic location. Linear maps displaying markers and genes on contig sequences are available, from which an object can be chosen. Any search starting point identifies all the information matching the query. HOWDY provides a convenient search environment of human genomic data for scientists unsure which database is most appropriate for their search.

Chromosome Mapping↗

EXProt: a database for proteins with an experimentally verified function.

EXProt is a non-redundant protein database containing a selection of entries from genome annotation projects and public databases, aimed at including only proteins with an experimentally verified function. In EXProt release 2.0 we have collected entries from the Pseudomonas aeruginosa community annotation project (PseudoCAP), the Escherichia coli genome and proteome database (GenProtEC) and the translated coding sequences from the Prokaryotes division of EMBL nucleotide sequence database, which are described as having an experimentally verified function. Each entry in EXProt has a unique ID number and contains information about the species, amino acid sequence, functional annotation and, in most cases, links to references in MEDLINE/PubMed and to the entry in the original database. EXProt is indexed in SRS at CMBI (http://www.cmbi.kun.nl/srs/) and can be searched with BLAST and FASTA through the EXProt web page (http://www.cmbi.kun.nl/EXProt/).

Animals↗

E-MSD: the European Bioinformatics Institute Macromolecular Structure Database.

The E-MSD macromolecular structure relational database (http://www.ebi.ac.uk/msd) is designed to be a single access point for protein and nucleic acid structures and related information. The database is derived from Protein Data Bank (PDB) entries. Relational database technologies are used in a comprehensive cleaning procedure to ensure data uniformity across the whole archive. The search database contains an extensive set of derived properties, goodness-of-fit indicators, and links to other EBI databases including InterPro, GO, and SWISS-PROT, together with links to SCOP, CATH, PFAM and PROSITE. A generic search interface is available, coupled with a fast secondary structure domain search tool.

Animals↗

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) provides a non-redundant collection of sequences representing genomic data, transcripts and proteins. Although the goal is to provide a comprehensive dataset representing the complete sequence information for any given species, the database pragmatically includes sequence data that are currently publicly available in the archival databases. The database incorporates data from over 2400 organisms and includes over one million proteins representing significant taxonomic diversity spanning prokaryotes, eukaryotes and viruses. Nucleotide and protein sequences are explicitly linked, and the sequences are linked to other resources including the NCBI Map Viewer and Gene. Sequences are annotated to include coding regions, conserved domains, variation, references, names, database cross-references, and other features using a combined approach of collaboration and other input from the scientific community, automated annotation, propagation from GenBank and curation by NCBI staff.

Animals↗

Computer databases of medical school curricula.

As the pace of curriculum reform in medical education has accelerated during the past decade, so too have demands on curriculum managers to supply increasingly detailed information about the curriculum. In response, a number of schools have joined together to begin work on designs for computer databases of the curriculum. The authors describe three of the most mature curriculum database prototypes, developed by groups at the medical schools of the University of North Carolina at Chapel Hill (UNC), The University of Maryland, and the University of Miami. All three groups have employed relational database management systems to organize information about each "instructional unit" in the preclinical curriculum, including a set of keywords defining the major concepts presented. The keywords are indexed to a controlled vocabulary, either the Medical Subject Headings (MeSH) or a MeSH derivative. The UNC database also employs a textfile management system to provide users with an overview of the entire curriculum. Future work will focus on identifying a suitable controlled vocabulary; capturing content in greater contextual detail; incorporating alternative learning formats, such as problem-based learning; creating links between content items and examination questions; and capturing information generated by student-patient interactions in clinical settings. As a result of recent collaboration with the Association of American Medical Colleges, work to define a prototype national database has begun and a consortium of interested schools is addressing further development activities.

Abstracting and Indexing↗

Protein three-dimensional structural databases: domains, structurally aligned homologues and superfamilies.

This paper reports the availability of a database of protein structural domains (DDBASE), an alignment database of homologous proteins (HOMSTRAD) and a database of structurally aligned superfamilies (CAMPASS) on the World Wide Web (WWW). DDBASE contains information on the organization of structural domains and their boundaries; it includes only one representative domain from each of the homologous families. This database has been derived by identifying the presence of structural domains in proteins on the basis of inter-secondary structural distances using the program DIAL [Sowdhamini & Blundell (1995), Protein Sci. 4, 506-520]. The alignment of proteins in superfamilies has been performed on the basis of the structural features and relationships of individual residues using the program COMPARER [Sali & Blundell (1990), J. Mol. Biol. 212, 403-428]. The alignment databases contain information on the conserved structural features in homologous proteins and those belonging to superfamilies. Available data include the sequence alignments in structure-annotated formats and the provision for viewing superposed structures of proteins using a graphical interface. Such information, which is freely accessible on the WWW, should be of value to crystallographers in the comparison of newly determined protein structures with previously identified protein domains or existing families.

Amino Acid Sequence↗

Database of repetitive elements in complete genomes and data mining using transcription factor binding sites.

Approximately 43% of the human genome is occupied by repetitive elements. Even more, around 51% of the rice genome is occupied by repetitive elements. The analysis presented here indicates that repetitive elements in complete genomes may have been very important in the evolutionary genomics. In this study, a database, called the Repeat Sequence Database, is first designed and implemented to store complete and comprehensive repetitive sequences. See http://rsdb.csie.ncu.edu.tw for more information. The database contains direct, inverted and palindromic repetitive sequences, and each repetitive sequence has a variable length ranging from seven to many hundred nucleotides. The repetitive sequences in the database are explored using a mathematical algorithm to mine rules on how combinations of individual binding sites are distributed among repetitive sequences in the database. Combinations of transcription factor binding sites in the repetitive sequences are obtained and then data mining techniques are applied to mine association rules from these combinations. The discovered associations are further pruned to remove insignificant associations and obtain a set of associations. The mined association rules facilitate efforts to identify gene classes regulated by similar mechanisms and accurately predict regulatory elements. Experiments are performed on several genomes including C. elegans, human chromosome 22, and yeast.

Algorithms↗