Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Knowledge discovery in clinical databases based on variable precision rough set model.

Since a large amount of clinical data are being stored electronically, discovery of knowledge from such clinical databases is one of the important growing research area in medical informatics. For this purpose, we develop KDD-R (a system for Knowledge Discovery in Databases using Rough sets), an experimental system for knowledge discovery and machine learning research using variable precision rough sets (VPRS) model, which is an extension of original rough set model. This system works in the following steps. First, it preprocesses databases and translates continuous data into discretized ones. Second, KDD-R checks dependencies between attributes and reduces spurious data. Third, the system computes rules from reduced databases. Finally, fourth, it evaluates decision making. For evaluation, this system is applied to a clinical database of meningoencephalitis, whose computational results show that several new findings are obtained.

Artificial Intelligence↗

A CORBA server for the Radiation Hybrid DataBase.

Modern biology depends on a wide range of software interacting with a large number of data sources, varying both in size, complexity and structure. The range of important databases in molecular biology and genetics makes it crucial to overcome the problems which this multiplicity presents. At EMBL-EBI we have started to use CORBA technology to support interoperability between a variety of databases, as well as to facilitate the integration of tools that access these databases. Within the Radiation Hybrid DataBase project we are confronted daily with the interoperation and linking issues. In this paper we present a CORBA infrastructure implemented to access the Radiation Hybrid DataBase.

Animals↗

Physician's working diagnosis compared to the Euricterus Real Life Data Diagnostic Tool Trial in three jaundice databases: Euricterus Dutch, independent prospective and independent retrospective.

BACKGROUND/AIMS: In the European Union Euricterus Project on (sub)Icterus proforma, the history and physical examination items were to be used for the physician's working diagnosis (PWD) and 'among others, for the development of the real life data electronic diagnostic tool, Trial. Trial delivers diagnosis probabilities based on Bayes' Theorem (B), completed by Trial Algorithm (TA). We wanted to compare the diagnostic accuracies (PWD and Trial probabilities as a percentage of the final diagnosis (FD) in a patient population) in 3 Dutch databases. METHODOLOGY: The inclusion criteria for both Euricterus and Trial were age > or = 16 and bilirubin > or = 20 mmol/l. Euricterus data gathering took place at the bedside on a proforma with (among other questions) 79 questions on history and physical examination as well as the diagnosis levels for the PWD (1 alternative possible) and FD (17 disease categories, dc). Trial was developed on the data of 7,104 Euricterus patients and its data-entry Demo has the same questions. It calculates the probability of each diagnosis of the 17 dc as a percentage, as each significant finding is encountered (BO, Bayesian Overall). It can simultaneously calculate the resemblance of the patient's signs and symptoms to each disease concomitantly (BV, Bayesian Vertical), and to any subset of a disease. Any probability is further tested for compatibility using TA, a subset of BV, delivering TA-PWD, TA-BO and TA-BV. The Trial test patients came from 3 databases: a Euricterus Dutch Patients Random Sample EDRS (n = 184, internal database) and 2 independent databases: prospective P (n = 80) and retrospective R (n = 152), totalling 416 patients. RESULTS: The accuracies of PWD and Trial showed no differences between the databases, and the results are therefore pooled (n = 416). With testing on the highest probability found, the PWD accuracy was 78%, TA-PWD 81%, TA-BO 74% and TA-BV 72%. The true FD's were mentioned (at any probability) in the PWD in 86%, TA-PWD in 92%, TA-BO in 94% and TA-BV in 91% of the patients. Testing only patients whose FD was "certain" or whose data were without omissions did not improve accuracy. Testing on probability > 95% improved BO and BV accuracy, but not TA-BO or TA-BV. CONCLUSIONS: The Physician's Working Diagnosis accuracy was approximately 80% and did not greatly improve after TA. The Trial TA-BO and TA-BV accuracies were only slightly less than the PWD. For well-trained physicians, Trial strengthens the physician's judgment, and for those less trained (or those to be trained), it delivers a (sub)icterus diagnostic disease probability at nearly consultant level.

Algorithms↗

Advanced query mechanisms for biological databases.

Existing query interfaces for biological databases are either based on fixed forms or textual query languages. Users of a fixed form-based query interface are limited to performing some pre-defined queries providing a fixed view of the underlying database, while users of a free text query language-based interface have to understand the underlying data models, specific query languages and application schemas in order to formulate queries. Further, operations on application-specific complex data (e.g., DNA sequences, proteins), which are usually provided by a variety of software packages with their own format requirements and peculiarities, are not available as part of, nor integrated with biological query interfaces. In this paper, we describe generic tools that provide powerful and flexible support for interactively exploring biological databases in a uniform and consistent way, that is via common data models, formats, and notations, in the framework of the Object-Protocol Model (OPM). These tools include (i) a Java graphical query construction tool with support for automatic generation of Web query forms that can be either used for further specifying conditions, or can be saved and customized; (ii) query processors for interpreting and executing queries that may involve complex application-specific objects, and that could span multiple heterogeneous databases and file systems; and (iii) utilities for automatic generation of HTML pages containing query results, that can be browsed using a Web browser. These tools avoid the restrictions imposed by traditional fixed-form query interfaces, while providing users with simple and intuitive facilities for formulating ad-hoc queries across heterogeneous databases, without the need to understand the underlying data models and query languages.

Animals↗

Consistency of variables in PCS and JASTRO great area database.

PURPOSE: To examine whether the Patterns of Care Study (PCS) reflects the data for the major areas in Japan, the consistency of variables in the PCS and in the major area database of the Japanese Society for Therapeutic Radiology and Oncology (JASTRO) were compared. METHODS AND PATIENTS: Patients with esophageal or uterine cervical cancer were sampled from the PCS and JASTRO databases. From the JASTRO database, 147 patients with esophageal cancer and 95 patients with uterine cervical cancer were selected according to the eligibility criteria for the PCS. From the PCS, 455 esophageal and 432 uterine cervical cancer patients were surveyed. Six items for esophageal cancer and five items for uterine cervical cancer were selected for a comparative analysis of PCS and JASTRO databases. RESULTS: Esophageal cancer: Age (p=.0777), combination of radiation and surgery (p=.2136), and energy of the external beam (p=.6400) were consistent for PCS and JASTRO. However, the dose of the external beam for the non-surgery group showed inconsistency (p=.0467). Uterine cervical cancer: Age (p=.6301) and clinical stage (p=.8555) were consistent for the two sets of data. However, the energy of the external beam (p<.0001), dose rate of brachytherapy (p<.0001), and brachytherapy utilization by clinical stage (p<.0001) showed inconsistencies. CONCLUSION: It appears possible that the JASTRO major area database could not account for all patients' backgrounds and factors and that both surveys might have an imbalance in the stratification of institutions including differences in equipment and staffing patterns.

Adult↗

ONTOFUSION: ontology-based integration of genomic and clinical databases.

ONTOFUSION is an ontology-based system designed for biomedical database integration. It is based on two processes: mapping and unification. Mapping is a semi-automated process that uses ontologies to link a database schema with a conceptual framework-named virtual schema. There are three methodologies for creating virtual schemas, according to the origin of the domain ontology used: (1) top-down--e.g. using an existing ontology, such as the UMLS or Gene Ontology--, (2) bottom-up--building a new domain ontology-- and (3) a hybrid combination. Unification is an automated process for integrating ontologies and hence the database to which they are linked. Using these methods, we employed ONTOFUSION to integrate a large number of public genomic and clinical databases, as well as biomedical ontologies.

Data Collection↗

A local alignment metric for accelerating biosequence database search.

We introduce a metric for local sequence alignments that has utility for accelerating optimal alignment searches without loss of sensitivity. The metric's triangle inequality property permits identification of redundant database entries guaranteed to have optimal alignments to the query sequence that fall below a specified score threshold, thereby permitting comparisons to these entries to be skipped. We prove the existence of the metric for a variety of scoring systems, including the most commonly used ones, and show that a triangle inequality can be established as well for nucleotide-to-protein sequence comparisons. We discuss a database clustering and search strategy that takes advantage of the triangle inequality. The strategy permits moderate but significant acceleration of searches against the widely used "nr" protein database. It also provides a theoretically based method for database clustering in general and provides a standard against which to compare heuristic clustering strategies.

Algorithms↗

eccDNABase: A Comprehensive and High-Quality Database for Extrachromosomal Circular DNA.

Extrachromosomal circular DNA (eccDNA) refers to small, circular DNA molecules that originate from chromosomal sequences and are prevalent across nearly all eukaryotic organisms. In humans, eccDNAs are widely distributed in normal tissues, cancerous tissues, and body fluids, where they play important roles in tumorigenesis and are often associated with poor clinical outcomes. Given their biological and clinical significance, a well-integrated and high-quality database is essential for advancing eccDNA-related research. To address this need, we developed eccDNABase, a comprehensive and curated resource for browsing, searching, and analyzing eccDNAs across multiple species. The database systematically catalogs eccDNA-disease associations from diverse tissues and organisms. Currently, eccDNABase contains 1,875,452 eccDNA-disease associations, encompassing 8,398 ecDNA entries across nine species, 63 diseases, and healthy individuals. Each entry provides detailed information, including eccDNA ID, type, chromosomal localization, species, tissue or cell line source, disease name and Disease Ontology ID, overlap length and percentage with genes, oncogene overlap, detection method, and links to literature and source databases. Given its extensive and curated datasets, eccDNABase serves as a valuable resource for both basic and translational research, offering deeper insights into the role of eccDNA in health and disease. The database is publicly accessible at http://cgga.org.cn/eccDNABase/.

Humans↗

wFleaBase: the Daphnia genome database.

BACKGROUND: wFleaBase is a database with the necessary infrastructure to curate, archive and share genetic, molecular and functional genomic data and protocols for an emerging model organism, the microcrustacean Daphnia. Commonly known as the water-flea, Daphnia's ecological merit is unequaled among metazoans, largely because of its sentinel role within freshwater ecosystems and over 200 years of biological investigations. By consequence, the Daphnia Genomics Consortium (DGC) has launched an interdisciplinary research program to create the resources needed to study genes that affect ecological and evolutionary success in natural environments. DISCUSSION: These tools include the genome database wFleaBase, which currently contains functions to search and extract information from expressed sequenced tags, genome survey sequences and full genome sequencing projects. This new database is built primarily from core components of the Generic Model Organism Database project, and related bioinformatics tools. SUMMARY: Over the coming year, preliminary genetic maps and the nearly complete genomic sequence of Daphnia pulex will be integrated into wFleaBase, including gene predictions and ortholog assignments based on sequence similarities with eukaryote genes of known function. wFleaBase aims to serve a large ecological and evolutionary research community. Our challenge is to rapidly expand its content and to ultimately integrate genetic and functional genomic information with population-level responses to environmental challenges. URL: http://wfleabase.org/.

Animals↗

PathwayVoyager: pathway mapping using the Kyoto Encyclopedia of Genes and Genomes (KEGG) database.

BACKGROUND: Equally important and challenging as genome annotation, is the subsequent classification of predicted genes into their respective pathways. The Kyoto Encyclopedia of Genes and Genomes (KEGG) represents a database consisting of known genes and their respective biochemical functionalities. Although accessible online, analyses of multiple genes are time consuming and are not suitable for analyzing data sets that are proprietary. RESULTS: Presented here is a new software solution that utilizes the KEGG online database for pathway mapping of partial and whole prokaryotic genomes. PathwayVoyager retrieves user-defined subsets of the KEGG database and stores the data as local, blast-formatted databases. Previously selected datasets can be re-used, reducing run-time significantly. Whole or partial genomes can be automatically analyzed using NCBI's BlastP algorithm and ORFs with similarities below the user-defined threshold will be marked on pathway maps. Multiple gene hits are sorted by similarity. Since no sequence information is transmitted over the Internet, PathwayVoyager is an ideal solution for pathway mapping and reconstruction of confidential DNA sequence data. CONCLUSION: PathwayVoyager represents an alternative approach to many already existing, more complex pathway reconstructions software solutions. This software does not require any dedicated hardware or software and is flexible and straightforward to use. It is ideally suited for environments where analyses on variable datasets are desired.

Computational Biology↗

An on-line database for zebrafish development and genetics research

We have built a relational database of zebrafish developmental and genetic research information accessible via the World Wide Web. Our team of biologists and computer scientists employed a user-centered design process, using input from the research community to tailor the contents and usability of the database. The database supports the broad range of data types generated by zebrafish research including text, images and graphical information about mutations, gene expression patterns and the genetic map. Data are entered both by the database staff and directly by authorized users. The database also maintains links among data, scientists and laboratories, thus facilitating information exchange within the research community.Copyright 1997 Academic Press Limited Copyright 1997Academic Press Limited

Journal Article↗

Tracking the violent criminal offender through DNA typing profiles--a national database system concept.

Implementation of standard methods for the conduct of restriction fragment length polymorphism analysis into the protocols of United States crime laboratories offers an unprecedented opportunity for the establishment of a national computer database system to enable interchange of DNA typing information. The FBI Laboratory, in concert with crime laboratory representatives, has taken the initiative in planning and implementing such a database system. The Combined DNA Index System (CODIS) will be composed of three sub-indices: a statistical database, which will contain frequencies of DNA fragment alleles in various population groups; an investigative database which will enable linkage of violent crimes through a common subject; and a convicted felon database that will serve to maintain DNA typing profiles for comparison to profiles developed from violent crimes where the suspect may be unknown.

Computer Communication Networks↗

Database technology in health care.

In this paper we will introduce the concepts of database technology in a way that will make it easy to relate the issues of the technology to problems in health care. After the objectives of the database approach have been defined, the major components of databases and their function will be discussed. The remainder of this paper presents the scientific and operational issues associated with database technology in health care. In the scientific exposition we will begin with the logical design issues, those that assure that the data will reflect the medical environment correctly, and then we will discuss the choices that are available for the physical implementation of a database on a computer system. The operational aspects will range from data entry to output presentation. The importance and growth of these systems has been well documented. Rather than providing a survey of the field, this exposition is intended to link general concepts to the practices observed by us and others.

Computers↗

Frequencies of Ty1- copia and Ty3- gypsy retroelements within the Triticeae EST databases.

The frequency of Ty1- copia-type and Ty3- gypsy-type retrotransposons in the International Triticeae EST Consortium (ITEC) database (61,942 sequences: 82% wheat, 10% barley, 8% rye) and the DuPont EST database (86,628 wheat sequences) was estimated using BLASTN searches. These ESTs were obtained from 94 cDNA libraries from different tissues (leaves, roots, spikes, flowers and seeds) and different growing conditions, excluding subtracted and normalized cDNA libraries. Triticeae EST databases were screened using four different Ty1- -copia-type, 12 reverse transcriptase sequences, and three Ty3- gypsy-type Triticeae retrotransposon sequences. Using a selection threshold of BLASTN scores higher than 100 or E values smaller than e(-20), 0.145% of the ESTs were found to be significantly similar to at least one of the retrotransposons used in the search (0.064% Ty1- copia, 0.081% Ty3- gypsy). This percentage increased to 0.176% when the BLASTN threshold was changed to E<e(-10). The percentage of ESTs similar to retrotransposons was significantly higher ( P < 0.05) in cDNA libraries from leaf tissues than in cDNA libraries from roots, anthers, or spikes. In addition, the percentage of ESTs similar to retrotransposons in cDNA libraries from plants under stress conditions (0.25% at E<e(-20), and 0.30% at E<e(-10)) was three to four folds higher ( P < 0.0001) than in cDNA libraries from plants grown under normal conditions (0.07% at E<e(-20), and 0.09% at E<e(-10)). Identification of retrotransposons within the Triticeae EST databases provides an indirect estimation of the patterns of transcriptional activity of these repetitive elements and is important to improve the annotation of genomic sequences used to search these EST databases.

Journal Article↗

Construction and testing of a microsatellite database containing more than 500 tomato varieties.

The aim of this study was to evaluate the suitability of sequence tagged microsatellite site (STMS) markers for varietal identification and discrimination in tomato. For this purpose, a set of 20 STMS primer pairs was used to construct a database containing the molecular description of the most common varieties (>500) of tomato grown in Europe. The database was built and tested by a consortium of five European laboratories each using a different STMS detection system. In this way, it could be demonstrated that the STMS markers and database were suitable for use in network activities where a common database is being established on a continuing basis with data from different laboratories.Microsatellite polymorphism in tomato was found to be relatively low. The number of alleles per locus ranged from 2 to 8 with an average of 4.7 alleles per locus. Nevertheless, more than 90% of the varieties had different microsatellite profiles. A "blind testing" exercise showed that in general, identification of unknown samples (or detecting the most similar variety) with the 20 markers and the database was relatively easy for homogeneous varieties but less certain with heterogeneous varieties when using pools of 6 individuals.

Journal Article↗

Database for sensorineural hearing loss.

We are creating a bank of EBV immortalized lymphoblast cells and extracted DNA taken from the blood of deaf children and their relatives, in order to study the molecular basis of hereditary deafness. We have established a corresponding database for sensorineural hearing loss that records clinical data for each entered specimen. The purpose of this paper is to present the content and design of the computerized relational database. The data model is designed first to identify known etiologies of deafness, either acquired or syndromic, and then to characterize the clinical features of the deaf individual, and both their affected and non-affected family members. The application operates in a graphical environment of visual prompts and message panels. The database is organized by sections which record demographic data, presenting complaints, otologic history, birth and perinatal history, developmental history, symptoms of chronic airway obstruction, family history, neurologic history, congenital infections, hospitalizations and surgical history, medication history, vestibular findings, audiometry, radiology, medical conditions and syndromes and physical examination. The database was developed on a commercially available software product. Our database is presented as a model for use by other clinicians and investigators.

Airway Obstruction↗

BIOLEFF: three databases on air pollution effects on vegetation.

Three databases on air pollution effects on vegetation were developed by storing bibliographic and abstract data for technical literature on the subject in a free-form database program, 'askSam'. Approximately 4 000 journal articles have been computerized in three separate database files: BIOLEFF, LICHENS and METALS. BIOLEFF includes over 2 800 articles on the effects of approximately 25 gaseous and particulate pollutants on over 2 000 species of vascular plants. LICHENS includes almost 400 papers on the effects of gaseous and heavy metal pollutants on over 735 species of lichens and mosses. METALS includes over 465 papers on the effects of heavy metals on over 830 species of vascular plants. The combined databases include articles from about 375 different journals spanning 1905 to the present. Picea abies and Phaseolus vulgaris are the most studied vascular plants in BIOLEFF, while Hypogymnia physodes is the most studied lichen species in LICHENS. Ozone and sulfur dioxide are the most studied gaseous pollutants with about two thirds of the records in BIOLEFF. The combined size of the databases is now about 5.5 megabytes.

Journal Article↗

A mathematical model that improves the validity of osteoarthritis diagnoses obtained from a computerized diagnostic database.

We developed an algorithm, using recursive partitioning, that utilized information from a computerized, diagnostic database to predict the diagnosis of osteoarthritis as determined by medical record review. The complete (inpatient and outpatient) medical records for a random sample of 400 Olmsted County, Minnesota residents with a database diagnosis consistent with osteoarthritis were reviewed, and confirmation or rejection of the diagnosis was accomplished. Of the 387 patients in our sample, only 232 (a positive predictive value of 60%) fulfilled diagnostic criteria for osteoarthritis following medical record review. A classification tree was created that used information from the diagnostic database to partition the study population according to the proportion of individuals with a "true" diagnosis of osteoarthritis (based on medical record review). The receiver-operating characteristic curve generated from these data illustrated that the algorithm substantially improved the validity of the database diagnosis, yielding a positive predictive value of 89% and a negative predictive value of 70% (sensitivity of 75% and specificity of 86%) at a selected cutoff point. This model also provides the capability of selecting the cutoff point to favor either specificity or sensitivity. These data demonstrate that a mathematical model can substantially improve the validity of computerized diagnostic databases in osteoarthritis.

Adult↗