Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Using XML technology for the ontology-based semantic integration of life science databases.

Several hundred internet accessible life science databases with constantly growing contents and varying areas of specialization are publicly available via the internet. Database integration, consequently, is a fundamental prerequisite to be able to answer complex biological questions. Due to the presence of syntactic, schematic, and semantic heterogeneities, large scale database integration at present takes considerable efforts. As there is a growing apprehension of extensible markup language (XML) as a means for data exchange in the life sciences, this article focuses on the impact of XML technology on database integration in this area. In detail, a general architecture for ontology-driven data integration based on XML technology is introduced, which overcomes some of the traditional problems in this area. As a proof of concept, a prototypical implementation of this architecture based on a native XML database and an expert system shell is described for the realization of a real world integration scenario.

Algorithms↗

A dynamic clinical dental relational database.

The traditional approach to relational database design is based on the logical organization of data into a number of related normalized tables. One assumption is that the nature and structure of the data is known at the design stage. In the case of designing a relational database to store historical dental epidemiological data from individual clinical surveys, the structure of the data is not known until the data is presented for inclusion into the database. This paper addresses the issues concerned with the theoretical design of a clinical dynamic database capable of adapting the internal table structure to accommodate clinical survey data, and presents a prototype database application capable of processing, displaying, and querying the dental data.

Database Management Systems↗

Coding and consent: moral challenges of the database project in Iceland.

A major moral problem in relation to the deCODE genetics database project in Iceland is that the heavy emphasis placed on technical security of healthcare information has precluded discussion about the issue of consent for participation in the database. On the other hand, critics who have emphasised the issue of consent have most often demanded that informed consent for participation in research be obtained. While I think that individual consent is of major significance, I argue that this demand for informed consent is neither suitable nor desirable in this case. I distinguish between three aspects of the database and show that different types of consent are appropriate for each. In particular, I describe the idea of a written authorisation based on general information about the database as an alternative to informed consent and presumed consent in database research.

Databases, Factual↗

Genomic BLAST: custom-defined virtual databases for complete and unfinished genomes.

BLAST (Basic Local Alignment Search Tool) searches against DNA and protein sequence databases have become an indispensable tool for biomedical research. The proliferation of the genome sequencing projects is steadily increasing the fraction of genome-derived sequences in the public databases and their importance as a public resource. We report here the availability of Genomic BLAST, a novel graphical tool for simplifying BLAST searches against complete and unfinished genome sequences. This tool allows the user to compare the query sequence against a virtual database of DNA and/or protein sequences from a selected group of organisms with finished or unfinished genomes. The organisms for such a database can be selected using either a graphic taxonomy-based tree or an alphabetical list of organism-specific sequences. The first option is designed to help explore the evolutionary relationships among organisms within a certain taxonomy group when performing BLAST searches. The use of an alphabetical list allows the user to perform a more elaborate set of selections, assembling any given number of organism-specific databases from unfinished or complete genomes. This tool, available at the NCBI web site http://www.ncbi.nlm.nih.gov/cgi-bin/Entrez/genom_table_cgi, currently provides access to over 170 bacterial and archaeal genomes and over 40 eukaryotic genomes.

Amino Acid Sequence↗

Applying database technology to clinical and basic research bioinformatics projects.

This paper describes the application of database technology to medical information with the goal of providing medical and clinical researchers with the tools necessary to plan bioinformatics projects. Commercial database management systems were utilized, standard database design practices were applied, a user interface was created, data entered, and the development of analysis tools, including data mining technologies is underway. Databases were constructed based on animal and cell culture models of diabetes and clinical data. Bioinformatics is a useful tool in both basic research and clinical settings. The advantages of relational databases and an approach to managing bioinformatics projects are discussed.

Animals↗

Databases in MS research: pitfalls and promises.

A database is an organized repository of data. Prospective collection of patient information in a database ('databasing') has been attempted by a few consortia of MS investigators over the past 10 years. This approach promises to facilitate epidemiologic research in MS and investigation of the natural history of the disease and how it might be altered by long-term treatments such as interferon beta. Databasing has some advantages over clinical trials in assessing new therapies, primarily because the focus is on long-term effectiveness in an entire population rather than short-term statistical significance in a highly selected population. The limitations of databasing and strategies to overcome these limitations are addressed.

Clinical Trials as Topic↗

Tools for loading MEDLINE into a local relational database.

BACKGROUND: Researchers who use MEDLINE for text mining, information extraction, or natural language processing may benefit from having a copy of MEDLINE that they can manage locally. The National Library of Medicine (NLM) distributes MEDLINE in eXtensible Markup Language (XML)-formatted text files, but it is difficult to query MEDLINE in that format. We have developed software tools to parse the MEDLINE data files and load their contents into a relational database. Although the task is conceptually straightforward, the size and scope of MEDLINE make the task nontrivial. Given the increasing importance of text analysis in biology and medicine, we believe a local installation of MEDLINE will provide helpful computing infrastructure for researchers. RESULTS: We developed three software packages that parse and load MEDLINE, and ran each package to install separate instances of the MEDLINE database. For each installation, we collected data on loading time and disk-space utilization to provide examples of the process in different settings. Settings differed in terms of commercial database-management system (IBM DB2 or Oracle 9i), processor (Intel or Sun), programming language of installation software (Java or Perl), and methods employed in different versions of the software. The loading times for the three installations were 76 hours, 196 hours, and 132 hours, and disk-space utilization was 46.3 GB, 37.7 GB, and 31.6 GB, respectively. Loading times varied due to a variety of differences among the systems. Loading time also depended on whether data were written to intermediate files or not, and on whether input files were processed in sequence or in parallel. Disk-space utilization depended on the number of MEDLINE files processed, amount of indexing, and whether abstracts were stored as character large objects or truncated. CONCLUSIONS: Relational database (RDBMS) technology supports indexing and querying of very large datasets, and can accommodate a locally stored version of MEDLINE. RDBMS systems support a wide range of queries and facilitate certain tasks that are not directly supported by the application programming interface to PubMed. Because there is variation in hardware, software, and network infrastructures across sites, we cannot predict the exact time required for a user to load MEDLINE, but our results suggest that performance of the software is reasonable. Our database schemas and conversion software are publicly available at http://biotext.berkeley.edu.

Database Management Systems↗

PASS2: an automated database of protein alignments organised as structural superfamilies.

BACKGROUND: The functional selection and three-dimensional structural constraints of proteins in nature often relates to the retention of significant sequence similarity between proteins of similar fold and function despite poor sequence identity. Organization of structure-based sequence alignments for distantly related proteins, provides a map of the conserved and critical regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination. The Protein Alignment organised as Structural Superfamily (PASS2) database represents continuously updated, structural alignments for evolutionary related, sequentially distant proteins. DESCRIPTION: An automated and updated version of PASS2 is, in direct correspondence with SCOP 1.63, consisting of sequences having identity below 40% among themselves. Protein domains have been grouped into 628 multi-member superfamilies and 566 single member superfamilies. Structure-based sequence alignments for the superfamilies have been obtained using COMPARER, while initial equivalencies have been derived from a preliminary superposition using LSQMAN or STAMP 4.0. The final sequence alignments have been annotated for structural features using JOY4.0. The database is supplemented with sequence relatives belonging to different genomes, conserved spatially interacting and structural motifs, probabilistic hidden markov models of superfamilies based on the alignments and useful links to other databases. Probabilistic models and sensitive position specific profiles obtained from reliable superfamily alignments aid annotation of remote homologues and are useful tools in structural and functional genomics. PASS2 presents the phylogeny of its members both based on sequence and structural dissimilarities. Clustering of members allows us to understand diversification of the family members. The search engine has been improved for simpler browsing of the database. CONCLUSIONS: The database resolves alignments among the structural domains consisting of evolutionarily diverged set of sequences. Availability of reliable sequence alignments of distantly related proteins despite poor sequence identity and single-member superfamilies permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. PASS2 is accessible at http://www.ncbs.res.in/~faculty/mini/campass/pass2.html

Amino Acid Sequence↗

MICA: desktop software for comprehensive searching of DNA databases.

BACKGROUND: Molecular biologists work with DNA databases that often include entire genomes. A common requirement is to search a DNA database to find exact matches for a nondegenerate or partially degenerate query. The software programs available for such purposes are normally designed to run on remote servers, but an appealing alternative is to work with DNA databases stored on local computers. We describe a desktop software program termed MICA (K-Mer Indexing with Compact Arrays) that allows large DNA databases to be searched efficiently using very little memory. RESULTS: MICA rapidly indexes a DNA database. On a Macintosh G5 computer, the complete human genome could be indexed in about 5 minutes. The indexing algorithm recognizes all 15 characters of the DNA alphabet and fully captures the information in any DNA sequence, yet for a typical sequence of length L, the index occupies only about 2L bytes. The index can be searched to return a complete list of exact matches for a nondegenerate or partially degenerate query of any length. A typical search of a long DNA sequence involves reading only a small fraction of the index into memory. As a result, searches are fast even when the available RAM is limited. CONCLUSION: MICA is suitable as a search engine for desktop DNA analysis software.

Algorithms↗

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface↗

MPID-T: database for sequence-structure-function information on T-cell receptor/peptide/MHC interactions.

UNLABELLED: Normal adaptive immune responses operate under major histocompatibility complex (MHC) restriction by binding to specific, short antigenic peptides and presenting them to appropriate T-cell receptors (TcRs). Sequence-structure-function information is critical in understanding the principles governing peptide/MHC (pMHC) and TcR/pMHC recognition and binding. A new database for sequence-structure-function information on TcR/pMHC interactions, MHC-Peptide Interaction Database version T (MPID-T), is now available with the latest available Protein Data Bank (PDB) data and interaction parameters on TcR/pMHC complexes. MPID-T is a manually curated MySQL database containing experimentally determined structures of 187 pMHC complexes and 16 TcR/pMHC complexes available in the PDB. Each structure is manually verified, classified, and analysed for intermolecular interactions (i) between the MHC and its corresponding bound peptide and (ii) between TcR and its bound pMHC complex where TcR structural information is available. The MPID-T database retrieval system has precomputed interaction parameters that include solvent accessibility, hydrogen bonds, gap volume and gap index. Structural visualisation of the TcR/pMHC complex, pMHC complex, MHC or the bound peptide can be performed using freely available graphics applications such as MDL Chime or RasMol, while structural alignment (based on MHC class and peptide length) can be viewed using the Jmol molecular viewer or an MDL Chime-compatible web browser client. MPID-T contains structural descriptors for in-depth characterisation of TcR/pMHC and pMHC interactions. The ultimate purpose of MPID-T is to enhance the understanding of the binding mechanism underlying TcR/pMHC and pMHC interactions by mapping the TcR footprint on the MHC and its bound peptide, as this eventually determines T-cell recognition and binding. AVAILABILITY: The MPID-T database retrieval system is available at http://surya.bic.nus.edu.sg/mpidt CONTACT: Joo Chuan Tong (jctong@i2r.a-star.edu.sg).

Animals↗

[Application of the JJ1017 code to master-table of RIS database].

We are developing an open-type Radiology Information System (RIS) under the project name KPECK. Part of the RIS has already been employed in Kouri Hospital of Kansai Medical University. The RIS is based on a database of the history of clinical study and exposure. We tried using the JJ1017 (ver1.0) code for the master-table of the database. The JJ1017 code, which is standardized by JIRA and JAHIS, is used in communicating information between the RIS and medical modalities. Through construction of the database, we found a technique by which the JJ1017 code could be applied to its master-table. In coordinating the JJ1017 code with the master-table of the database, we extended various study codes and systematically coordinated them with the architecture of the database.

Database Management Systems↗

Integrating a modern knowledge-based system architecture with a legacy VA database: the ATHENA and EON projects at Stanford.

We present a methodology and database mediator tool for integrating modern knowledge-based systems, such as the Stanford EON architecture for automated guideline-based decision-support, with legacy databases, such as the Veterans Health Information Systems & Technology Architecture (VISTA) systems, which are used nation-wide. Specifically, we discuss designs for database integration in ATHENA, a system for hypertension care based on EON, at the VA Palo Alto Health Care System. We describe a new database mediator that affords the EON system both physical and logical data independence from the legacy VA database. We found that to achieve our design goals, the mediator requires two separate mapping levels and must itself involve a knowledge-based component.

Artificial Intelligence↗

A database system for the analysis of biochemical pathways.

To provide support for the analysis of biochemical pathways a database system based on a model that represents the characteristics of the domain is needed. This domain has proven to be difficult to model by using conventional data modelling techniques. We are building an ontology for biochemical pathways, which acts as the basis for the generation of a database on the same domain, allowing the definition of complex queries and complex data representation. The ontology is used as a modelling and analysis tool which allows the expression of complex semantics based on a first-order logic representation language. The induction capabilities of the system can help the scientist in formulating and testing research hypotheses that are difficult to express with the standard relational database mechanisms. An ontology representing the shared formalisation of the knowledge in a scientific domain can also be used as data integration tool clarifying the mapping of concepts to the developers of different databases. In this paper we describe the general structure of our system, concentrating on the ontology-based database as the key component of the system.

Biochemical Phenomena↗

Programmed database system at the Chang Gung Craniofacial Center: part I.

BACKGROUND: A database is a system for the management of information. Databases of different forms are widely used in everyday life from telephone books to online library catalogs. The Craniofacial Center at Chang Gung Memorial Hospital has seen over 20,000 patients during the past 20 years. All of the patient records need to digitally input into a computer database. METHODS: A database was custom designed using Paradox 8. The ACDSee Photo browser and DOS linked them to the original program. The Paradox 8 was programmed to a standard mode for the diagnosis and treatment data input to prevent typographical errors. RESULTS: We collected the records of 25,200 patients from 1987 to 2002, of which 24,331 underwent operations. The data for 14,828 patients were registered as complete and/or incomplete cleft and the proportions of unilateral to bilateral and female to male are presented in Table 1. CONCLUSION: This new database system was designed to ensure the accuracy of data input using a standard model that is capable of correct data programming using the custom designed coding system for the Craniofacial Center. The system also provides easy and reliable data retrieval when using the powerful search tools.

Cleft Lip↗

A knowledge-based time-oriented active database approach for intelligent abstraction, querying and continuous monitoring of clinical data.

Query and interpretation of time-oriented medical data involves two subtasks: Temporal-reasoning--intelligent analysis of time-oriented data, and temporal-maintenance--effective storage, query, and retrieval of these data. Integration of these tasks into one system, known as temporal-mediator, has been proven to be beneficial to biomedical applications such as monitoring, therapy, quality assessment, visualization and exploration of time-oriented data. One potential problem in existing temporal-mediation approaches is lack of sufficient responsiveness when querying or continuously monitoring the database for complex abstract concepts that are derived from the raw data, especially regarding a large patient group. We propose a new approach: the knowledge-based time-oriented active database, a temporal extension of the active-database concept, and a merger of temporal reasoning and temporal maintenance within a persistent database framework. The approach preserves the efficiency of databases in handling data storage and retrieval, while enabling specification and performance of complex temporal reasoning using an incremental-computation approach. We implemented our approach within the Momentum system. Initial experiments are encouraging; an evaluation is underway

Algorithms↗

Internet-capable publication database system.

Scientific databases are generally accessible to the public via the Internet. Reports of most peer-reviewed (quotable) research is thus available to researchers and others. However, other reports and information of interest to researchers and teachers such as poster presentations at congresses, articles describing techniques and teaching material, and details of vocational and continuing education courses (nonquotable literature) generally do not appear in such databases. This nonquotable literature is often of great use to teachers. A project was therefore initiated at the Münster Dental Clinic which aimed to address the problem by developing a database of all publications and other printed material produced by the staff (faculty). After a systematic search, all such publications (quotable and nonquotable) were entered in the database which is partially accessible via the Internet and fully accessible via the Münster Dental Clinic's Intranet. The complete list can be found in the protected Intranet areas, which can be accessed by all the Dental Clinic's staff members. The database also permits Münster Clinic staff to access the Internet and locate those publications that are on the Internet by year of publication and topic.

Database Management Systems↗

Partitioning medical image databases for content-based queries on a Grid.

OBJECTIVES: In this paper we study the impact of executing a medical image database query application on the grid. For lowering the total computation time, the image database is partitioned into subsets to be processed on different grid nodes. METHODS: A theoretical model of the application complexity and estimates of the grid execution overhead are used to efficiently partition the database. RESULTS: We show results demonstrating that smart partitioning of the database can lead to significant improvements in terms of total computation time. CONCLUSIONS: Grids are promising for content-based image retrieval in medical databases.

Database Management Systems↗