Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

DNA databases.

This paper presents DNA algorithms for five relational algebra database operations, selection, projection, union, set difference, and Cartesian product on so-called DNA databases. A DNA database is a database where data records are encoded as DNA strands. The five operations mentioned before are fundamental in the field of databases and perform most of the data retrieval operations on current databases.

Algorithms↗

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea↗

DNA, diseases and databases: disastrously deficient.

Recent progress in disease genetics and genome-related medicine has been substantial, with vast amounts of data being generated. However, this progress has not been matched by adequate database projects that gather and organize these data to enable their useful exploitation. This research area is complex, entailing core databases, locus-specific databases, national mutation databases, genotype-phenotype databases and patient databases--and much work is required to develop and properly integrate these various resources. To promote this, we present a timely overview of the field, emphasize its over-riding importance and discuss the disastrously deficient progress made so far. Many factors contribute to this slow progress (e.g. technological hurdles, publication requirements, the short-sighted and popularist research system). A lack of targeted funding is arguably the most fundamental problem, but one that can be solved.

Databases, Factual↗

Development, analysis, refinement, and utility of an interdisciplinary amyotrophic lateral sclerosis database.

The current status of evaluation and management provided by individual healthcare professionals (HCP) at amyotrophic lateral sclerosis (ALS) centers and clinics needs to be analyzed. This paper describes one ALS center's experiences with the development, analysis, refinement, and utility of an interdisciplinary, HCP-driven ALS database. The purpose and conceptual framework of the database, the general data that needed to be collected, and the types of reports that needed to be generated were determined, and, in collaboration with a computer programmer, data entry and database management systems were developed. Data were collected on 234 patients between September 1996 and August 1998, and were analyzed by a biostatistician. Based on review of the biostatistician's report and discussion of problems encountered with the systems, the database was then refined. Benefits of the database system included: systematization of data collection and reporting, reduction of redundant data collection by individuals, decreased variability of evaluation methods and management decisions from patient to patient, and increased availability of a variety of uniform patient information to assist team members in making care decisions. Ongoing refinement will ensure that this HCP-driven ALS database continues to be informative, practical and effective for decision-making and enhancing delivery of care.

Adult↗

Development of an animal genome database and its search system.

An animal genome database has been developed on a Unix workstation and maintained by a relational database management system. This database has focused on the comparative gene mapping between species to assist the mapping of the genes related to phenotypic traits in livestock. The linkage maps, cytogenetic maps, polymerase chain reaction primers of pig, cattle, mouse and human, and their references have been included in the database, and the correspondence among species have been stipulated in the database. In order to search the database effectively, the World Wide Web server (http://ws4.niai.affrc.go.jp/) and the electronic mail server system (e-mail: jgbase-mail@ niai.affrc.go.jp) have been developed on different Unix workstations. These servers are connected to the Internet.

Animals↗

Database-driven multi locus sequence typing (MLST) of bacterial pathogens.

MOTIVATION: Multi Locus Sequence Typing (MLST) is a newly developed typing method for bacteria based on the sequence determination of internal fragments of seven house-keeping genes. It has proved useful in characterizing and monitoring disease-causing and antibiotic resistant lineages of bacteria. The strength of this approach is that unlike data obtained using most other typing methods, sequence data are unambiguous, can be held on a central database and be queried through a web server. RESULTS: A database-driven software system (mlstdb) has been developed, which is used by public health laboratories and researchers globally to query their nucleotide sequence data against centrally held databases over the internet. The mlstdb system consists of a set of perl scripts for defining the database tables and generating the database management interface and dynamic web pages for querying the databases. AVAILABILITY: http://www.mlst.net.

Bacteria↗

BioMolQuest: integrated database-based retrieval of protein structural and functional information.

MOTIVATION: Information about a particular protein or protein family is usually distributed among multiple databases and often in more than one entry in each database. Retrieval and organization of this information can be a laborious task. This task is complicated even further by the existence of alternative terms for the same concept. RESULTS: The PDB, SWISS-PROT, ENZYME, and CATH databases have been imported into a combined relational database, BIOMOLQUEST: A powerful search engine has been built using this database as a back end. The search engine achieves significant improvements in query performance by automatically utilizing cross-references between the legacy databases. The results of the queries are presented in an organized, hierarchical way.

Abstracting and Indexing↗

The Database of Quantitative Cellular Signaling: management and analysis of chemical kinetic models of signaling networks.

MOTIVATION: Analysis of cellular signaling interactions is expected to pose an enormous informatics challenge, perhaps even larger than analyzing the genome. The complex networks arising from signaling processes are traditionally represented as block diagrams. A key step in the evolution toward a more quantitative understanding of signaling is to explicitly specify the kinetics of all chemical reaction steps in a pathway. Technical advances in proteomics and high-throughput protein interaction assays promise a flood of such quantitative data. While annotations, molecular information and pathway connectivity have been compiled in several databases, and there are several proposals for general cell model description languages, there is currently little experience with databases of chemical kinetics and reaction level models of signaling networks. RESULTS: The Database of Quantitative Cellular Signaling is a repository of models of signaling pathways. It is intended both to serve the growing field of chemical-reaction level simulation of signaling networks, and to anticipate issues in large-scale data management for signaling chemistry. AVAILABILITY: The Database of Quantitative Cellular Signaling is available at http://doqcs.ncbs.res.in. Links to the signaling model simulator, GENESIS/Kinetikit are at http://www.ncbs.res.in/~bhalla/kkit/index.html and are also provided from within the database. The database source code is available under the GNU Public License.

Abstracting and Indexing↗

MHCBN: a comprehensive database of MHC binding and non-binding peptides.

MHCBN is a comprehensive database of Major Histocompatibility Complex (MHC) binding and non-binding peptides compiled from published literature and existing databases. The latest version of the database has 19 777 entries including 17 129 MHC binders and 2648 MHC non-binders for more than 400 MHC molecules. The database has sequence and structure data of (a) source proteins of peptides and (b) MHC molecules. MHCBN has a number of web tools that include: (i) mapping of peptide on query sequence; (ii) search on any field; (iii) creation of data sets; and (iv) online data submission. The database also provides hypertext links to major databases like SWISS-PROT, PDB, IMGT/HLA-DB, GenBank and PUBMED.

Amino Acid Sequence↗

BioQuery: an object framework for building queries to biomedical databases.

SUMMARY: BioQuery is an application that helps scientists automate database searches. Users can build and store queries to public biomedical databases, and receive periodic updates on the results of those queries when new data is available. The application is implemented on a portable object framework that can provide database-searching capability to other applications. This framework is easily extensible, allowing users to develop plug-ins that provide access to new databases. BioQuery thus provides end-users with a complete database searching interface and updating service, and gives developers a toolkit to provide database-searching capability to their applications. AVAILABILITY: Free to all users: http://www.bioquery.org.

Biomedical Research↗

CBS Genome Atlas Database: a dynamic storage for bioinformatic results and sequence data.

UNLABELLED: Currently, new bacterial genomes are being published on a monthly basis. With the growing amount of genome sequence data, there is a demand for a flexible and easy-to-maintain structure for storing sequence data and results from bioinformatic analysis. More than 150 sequenced bacterial genomes are now available, and comparisons of properties for taxonomically similar organisms are not readily available to many biologists. In addition to the most basic information, such as AT content, chromosome length, tRNA count and rRNA count, a large number of more complex calculations are needed to perform detailed comparative genomics. DNA structural calculations like curvature and stacking energy, DNA compositions like base skews, oligo skews and repeats at the local and global level are just a few of the analysis that are presented on the CBS Genome Atlas Web page. Complex analysis, changing methods and frequent addition of new models are factors that require a dynamic database layout. Using basic tools like the GNU Make system, csh, Perl and MySQL, we have created a flexible database environment for storing and maintaining such results for a collection of complete microbial genomes. Currently, these results counts to more than 220 pieces of information. The backbone of this solution consists of a program package written in Perl, which enables administrators to synchronize and update the database content. The MySQL database has been connected to the CBS web-server via PHP4, to present a dynamic web content for users outside the center. This solution is tightly fitted to existing server infrastructure and the solutions proposed here can perhaps serve as a template for other research groups to solve database issues. AVAILABILITY: A web based user interface which is dynamically linked to the Genome Atlas Database can be accessed via www.cbs.dtu.dk/services/GenomeAtlas/. SUPPLEMENTARY INFORMATION: This paper has a supplemental information page which links to the examples presented: www.cbs.dtu.dk/services/GenomeAtlas/suppl/bioinfdatabase.

Algorithms↗

Babel's tower revisited: a universal resource for cross-referencing across annotation databases.

MOTIVATION: Annotation databases are widely used as public repositories of biological knowledge. However, most of these resources have been developed by independent groups which used different designs and different identifiers for the same biological entities. As we show in this article, incoherent name spaces between various databases represent a serious impediment to using the existing annotations at their full potential. Navigating between various such name spaces by mapping IDs from one database to another is a very important issue which is not properly addressed at the moment. RESULTS: We have developed a web-based resource, Onto-Translate (OT), which effectively addresses this problem. OT is able to map onto each other different types of biological entities from the following annotation databases: Swiss-Prot, TrEMBL, NREF, PIR, Gene Ontology, KEGG, Entrez Gene, GenBank, GenPept, IMAGE, RefSeq, UniGene, OMIM, PDB, Eukaryotic Promoter Database, HUGO Gene Nomenclature Committee and NetAffx. Currently, OT is able to perform 462 types of mappings between 29 different types of IDs from 17 databases concerning 53 organisms. Among these, over 300 types of translations and 15 types of IDs are not currently supported by any other tool or resource. On average, OT is able to correctly map between 96 and 99% of the biological entities provided as input. In terms of speed, sets of approximately 20 000 IDs can be translated in <30 s, in most cases. AVAILABILITY: OT is a part of Onto-Tools, which is freely available at http://vortex.cs.wayne.edu/Projects.html

Database Management Systems↗

DAtA: database of Arabidopsis thaliana annotation.

The Database of Arabidopsis thaliana Annotation (D At A) was created to enable easy access to and analysis of all the Arabidopsis genome project annotation. The database was constructed using the completed A.thaliana genomic sequence data currently in GenBank. An automated annotation process was used to predict coding sequences for GenBank records that do not include annotation. D At A also contains protein motifs and protein similarities derived from searches of the proteins in D At A with motif databases and the non-redundant protein database. The database is routinely updated to include new GenBank submissions for Arabidopsis genomic sequences and new Blast and protein motif search results. A web interface to D At A allows coding sequences to be searched by name, comment, blast similarity or motif field. In addition, browse options present lists of either all the protein names or identified motifs present in the sequenced A.thaliana genome. The database can be accessed at http://baggage. stanford.edu/group/arabprotein/

Arabidopsis↗

PASS2: a semi-automated database of protein alignments organised as structural superfamilies.

PASS2 is a nearly automated version of CAMPASS and contains sequence alignments of proteins grouped at the level of superfamilies. This database has been created to fall in correspondence with SCOP database (1.53 release) and currently consists of 110 multi-member superfamilies and 613 superfamilies corresponding to single members. In multi-member superfamilies, protein chains with no more than 25% sequence identity have been considered for the alignment and hence the database aims to address sequence alignments which represent 26 219 protein domains under the SCOP 1.53 release. Structure-based sequence alignments have been obtained by COMPARER and the initial equivalences are provided automatically from a MALIGN alignment and subsequently augmented using STAMP4.0. The final sequence alignments have been annotated for the structural features using JOY4.0. Several interesting links are provided to other related databases and genome sequence relatives. Availability of reliable sequence alignments of distantly related proteins, despite poor sequence identity and single-member superfamilies, permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. The database can be queried by keywords and also by sequence search, interfaced by PSI-BLAST methods. Structure-annotated sequence alignments and several structural accessory files can be retrieved for all the superfamilies including the user-input sequence. The database can be accessed from http://www.ncbs.res.in/%7Efaculty/mini/campass/pass.html.

Amino Acid Sequence↗

Improvements to GALA and dbERGE II: databases featuring genomic sequence alignment, annotation and experimental results.

We describe improvements to two databases that give access to information on genomic sequence similarities, functional elements in DNA and experimental results that demonstrate those functions. GALA, the database of Genome ALignments and Annotations, is now a set of interlinked relational databases for five vertebrate species, human, chimpanzee, mouse, rat and chicken. For each species, GALA records pairwise and multiple sequence alignments, scores derived from those alignments that reflect the likelihood of being under purifying selection or being a regulatory element, and extensive annotations such as genes, gene expression patterns and transcription factor binding sites. The user interface supports simple and complex queries, including operations such as subtraction and intersections as well as clustering and finding elements in proximity to features. dbERGE II, the database of Experimental Results on Gene Expression, contains experimental data from a variety of functional assays. Both databases are now run on the DB2 database management system. Improved hardware and tuning has reduced response times and increased querying capacity, while simplified query interfaces will help direct new users through the querying process. Links are available at http://www.bx.psu.edu/.

Animals↗

CAGE Basic/Analysis Databases: the CAGE resource for comprehensive promoter analysis.

Cap-analysis gene expression (CAGE) Basic and Analysis Databases store an original resource produced by CAGE, which measures expression levels of transcription starting sites by sequencing large amounts of transcript 5' ends, termed CAGE tags. Millions of human and mouse high-quality CAGE tags derived from different conditions in >20 tissues consisting of >250 RNA samples are essential for identification of novel promoters and promoter characterization in the aspect of expression profile. CAGE Basic Database is a primary database of the CAGE resource, RNA samples, CAGE libraries, CAGE clone and tag sequences and so on. CAGE Analysis Database stores promoter related information, such as counts of related transcripts, CpG islands and conserved genome region. It also provides expression profiles at base pair and promoter levels. Both databases are based on the same framework, CAGE tag starting sites, tag clusters for defining promoters and transcriptional units (TUs). Their associations and TU attributes are available to find promoters of interest. These databases were provided for Functional Annotation Of Mouse 3 (FANTOM3), an international collaboration research project focusing on expanding the transcriptome and subsequent analyses. Now access is free for all users through the World Wide Web at http://fantom3.gsc.riken.jp/.

Animals↗

The TIGR Plant Transcript Assemblies database.

The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.

DNA, Complementary↗

QxDB: a generic database to support mathematical modelling in biology.

QxDB (quantitative x-modelling database) is a web-based generic database package designed especially to house quantitative and structural information. Its development was motivated by the need for centralized access to such results for development of mathematical models, but its usefulness extends to the general research community of both modellers and experimentalists. Written in PHP (Hyper Preprocessor) and MYSQL, the database is easily adapted to new fields of research and ported to Apache-based web servers. Unlike most existing databases, experimental and observational results curated in QxDB are supplemented by comments from the experts who contribute input to the database, giving their evaluations of experimental techniques, breadth of validity of results, experimental conditions, and the like, thus providing the visitor with a basis for gauging the quality (or appropriateness) of each item for his/her needs. QxDB can be easily customized by adapting the contents of the database table containing the descriptors that characterize each data record according to an informal ontology of the research domain. We will illustrate this adaptability of QxDB by presenting two examples, the first dealing with modelling in oncology and the second with mechanical properties of cells and tissues.

Biology↗