Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

An integrated analysis and database system for full-length cDNA.

Annotation and database system of full-length cDNA sequences was developed. As the components of the system, ORF annotation system, functional annotation system based on database search results, mapping annotation system, and integrated retrieval and display system were developed. In the ORF annotation system integrated analyses using conventional tools are performed and useful retrieval interface using motif list are introduced. In the functional annotation system based on database search results, a new method that characterizes a given unknown cDNA was developed by using a profile of similarity level over words appearing in sequence database entries. In the mapping annotation system, we linked by similarity searches full-length cDNA sequences with database DNA sequences that are already mapped on chromosomes. By using these links, full-length cDNAs can be retrieved by the retrieval condition of physical mapping information. Genetic disease information mapped on the physical mapping site can also be displayed by this system. Furthermore, we constructed an integrated database system for these analyzed data, and thus enabled annotation and selection of full-length cDNAs from points of both gene function and mapping information.

Chromosome Mapping↗

Implementation of a classification hierarchy for the GeneTests/GeneClinics genetic testing databases.

The combination of a) our changing understanding of genotypic and phenotypic classification of diseases and b) the rapid growth and expansion of the number of entries in two databases targeted toward clinicians resulted in the need to develop a flexible dynamic hierarchical classification system for genetic disorders. The two databases making use of this classification schemas are the GeneClinics (GC) database - www.geneclinics.org and the GeneTests (GT) database - www.genetests.org The GC and GT databases serve respsectively as the users manual and yellow pages of genetic testing. The GeneTests/GeneClinics (GT/GC) classification hierarchy is maintained as a simple set of parent/child relationships in a relational database. The hierarchy is generated in real time in response to a user request. It is not maintained as a set of members with relationships defined by characters that are parsed to determine the structure of the tree. The GT/GC classification hierarchy entries are handled as objects by the data maintenance and search tools and may have a number of attributes and associations that create a rich tool for defining and examining genetic disorders

Databases, Genetic↗

The KEGG database.

KEGG (http://www.genome.ad.jp/kegg/) is a suite of databases and associated software for understanding and simulating higher-order functional behaviours of the cell or the organism from its genome information. First, KEGG computerizes data and knowledge on protein interaction networks (PATHWAY database) and chemical reactions (LIGAND database) that are responsible for various cellular processes. Second, KEGG attempts to reconstruct protein interaction networks for all organisms whose genomes are completely sequenced (GENES and SSDB databases). Third, KEGG can be utilized as reference knowledge for functional genomics (EXPRESSION database) and proteomics (BRITE database) experiments. I will review the current status of KEGG and report on new developments in graph representation and graph computations.

Amino Acid Sequence↗

Integr8: enhanced inter-operability of European molecular biology databases.

OBJECTIVES: The increasing production of molecular biology data in the post-genomic era, and the proliferation of databases that store it, require the development of an integrative layer in database services to facilitate the synthesis of related information. The solution of this problem is made more difficult by the absence of universal identifiers for biological entities, and the breadth and variety of available data. METHODS: Integr8 was modelled using UML (Universal Modelling Language). Integr8 is being implemented as an n-tier system using a modern object-oriented programming language (Java). An object-relational mapping tool, OJB, is being used to specify the interface between the upper layers and an underlying relational database. RESULTS: The European Bioinformatics Institute is launching the Integr8 project. Integr8 will be an automatically populated database in which we will maintain stable identifiers for biological entities, describe their relationships with each other (in accordance with the central dogma of biology), and store equivalences between identified entities in the source databases. Only core data will be stored in Integr8, with web links to the source databases providing further information. CONCLUSIONS: Integr8 will provide the integrative layer of the next generation of bioinformatics services from the EBI. Web-based interfaces will be developed to offer gene-centric views of the integrated data, presenting (where known) the links between genome, proteome and phenotype.

Computational Biology↗

Deniz: the electronic database for beta-thalassemia mutations in the Arab world.

OBJECTIVE: Data on the distribution of beta-thalassemia mutations in Arab populations are usually destined to disparate locations and much of these become increasingly difficult for an average researcher to locate. That is why we aimed at establishing an electronic database network, called Deniz, for beta-thalassemia allele frequency distributions in the Arab world at http://biobase.fatih.edu.tr. METHODS: The scheme of the database combines the benefits of the relational and hierarchical systems. Detailed statistics of the frequencies of beta-thalassemia mutations are retrieved in tabular forms. Multiple permanent connections allow flexible movement within the database. Queries are processed by the systems language and sent to the user's browser as hypertext markup language documents. RESULTS: The database catalogues the frequencies of beta-thalassemia mutations in 14 Arab countries as pooled from the analysis of 3,138 chromosomes by 36 laboratories. Of the 57 B-globin gene mutations reported in Arabs, IVS-I-110 (G-A), IVS-I-5 (G-C), IVS-I-6 (T-C), IVS-II-1 (G-A), and IVS-I-1 (G-A) are the most encountered and they account for approximately two thirds of the Arab chromosomes registered in Deniz. CONCLUSION: In addition to its importance as a hub of updated information on the distribution of beta-thalassemia mutations in Arabs, information in Deniz may be used to predict diagnostic strategies that may be offered to natives of unstudied countries. Incidence data may also give important clues on the possible origins of beta-thalassemia in the Arab world. The integration of Deniz with other databases is currently in process and researchers are invited to contribute to the growth of the database.

Alleles↗

Uniform databases in early arthritis: specific measures to complement classification criteria and indices of clinical change.

Rheumatoid arthritis (RA) is not characterized by a single pathognomonic measure such as blood pressure in hypertension or cholesterol in hyperlipidemia, which can be used in the diagnosis, prognosis, and monitoring of patient status. Measures such as swollen joints and an elevated erythrocyte sedimentation rate are certainly valuable, but many individuals with abnormal values have conditions other than RA, and many people with RA may have favorable values for one or more of these measures. Therefore, the rheumatology community has developed indices of several measures, such as classification criteria, the disease activity score (DAS), and the ACR Core Data Set with 20%, 50% and 70% improvement (ACR 20, ACR 50, ACR 70) to classify and monitor patients with RA. While these indices have greatly advanced clinical research, databases for long-term observations, including those in early RA described in this Supplement, differ in 20-50% of included data, and the software platforms for these databases differ sufficiently to render it difficult to merge the data to compare one data set to another. It has been proposed that a uniform database for early arthritis clinical research could help advance clinical research in early arthritis. One example of such a database, termed a "standard protocol to evaluate rheumatoid arthritis" (SPERA), has been in use for almost two decades in one clinical site, and has proven valuable in a number of ways, including the demonstration of early radiographic damage, development of a 28-joint count, and documentation that patient questionnaire data are correlated significantly with laboratory, joint count and radiographic data, although questionnaire data are the strongest predictors of severe outcomes including work disability and premature mortality. The use of a uniform database in no way precludes the collection of additional data at particular centers including immunogenetic, serologic, or structural magnetic resonance imaging (MRI) data. However, the availability of an infrastructure of standard data in all RA databases would enhance clinical research in early RA.

Adult↗

Automating terminological networks to link heterogeneous biomedical databases.

As cross-disciplinary research escalates, researchers are facing the challenge of linking disparate biomedical databases that have been developed without common indexes. Manually indexing these large-scale databases is laborious and often impractical. Solutions involving mediating terminologies have been proposed, but coordination of terms from the databases of interest to these mediating terminologies is also laborious, and regular synchronization between indexes is an additional problem. In this study we describe a novel method of linking heterogeneous databases using terminology networks constructed with automated mapping methods. Linkage was established between two disparate biomedical databases (SNOMED-CT and HDG), using two relevant intermediating databases (UMLS and OMIM). One gold standard of 514 distinct matches is used as proof-of-principle. In conclusion, as hypothesized, 1) Manually curated pathways provide high precision, but offer low recall, 2) the automated terminology pathways can significantly increase recall at acceptable precision. Taken together, our conclusion may suggest the combined manual and automated terminology networks could offer recall and precision in an incremental manner

Abstracting and Indexing↗

Addressing the use of phylogenetics for identification of sequences in error in the SWGDAM mitochondrial DNA database.

The SWGDAM mtDNA database is a publicly available reference source that is used for estimating the rarity of an evidence mtDNA profile. Because of the current processes for generating population data, it is unlikely that population databases are error free. The majority of the errors are due to human error and are transcriptional in nature. Phylogenetic analysis of data sets can identify some potential errors, and coupled with a review of the sequence data or alignment sheets can be a very useful tool. Seven sequences with errors have been identified by phylogenetic analysis. In addition, two samples were inadvertently modified when placed in the SWGDAM database. The corrected sequences are provided so that users can modify appropriately the current iteration of the SWGDAM database. From a practical perspective, upper bound estimates of the percentage of matching profiles obtained from a database search containing an incorrect sequence and those of a database containing the corrected sequence are not substantially different. Community wide access and review has enabled identification of errors in the SWGDAM data set and will continue to do so. The result of public accessibility is that the quality of the SWGDAM forensic dataset is always improving.

Base Sequence↗

A database to record, track and report health student rural placements.

The Spencer Gulf Rural Health School (SGRHS), South Australia, is funded by the Australian Commonwealth Government to deliver health education in the rural setting. The SGRHS required a database to record, track and report on student rural placements to satisfy Commonwealth reporting requirements, and for internal academic and administration staff use. Staff in widely separate rural locations needed to be able to access the database. A web-based relational database was created using Microsoft Access. The student rural placement database has been successfully utilised as the primary tool to record and track student placements in the SGRHS for 2 years, and has generated data for eight Commonwealth reports in this time. Future database developments include student accessible sections. With few alterations the database could be utilised by other Australian Rural Clinical Schools and University Departments of Rural Health.

Clinical Clerkship↗

Relational database design.

Relational databases are the predominant method for storing repetitive data in computers because they allow efficient and flexible storage of that data. While medical directors and underwriters are more likely to use a spreadsheet than a database program to analyze their business, the data they wish to study are often stored in corporate databases. Or the data may be complex enough to require being keyed into or downloaded into a personal computer (PC) database program for storage, even if the data are then output to a spread-sheet for numerical analysis. In many circumstances, one can benefit from an understanding of efficient database design. After a brief overview, the reader is led step-by-step through a practical explanation of database design, from a flat file to a relational model.

Database Management Systems↗

NEOBASE: databasing the neocortical microcircuit.

Mammals adapt to a rapidly changing world because of the sophisticated perceptual and cognitive function enabled by the neocortex. The neocortex, which has expanded to constitute nearly 80% of the human brain seems to have arisen from repeated duplication of a stereotypical template of neurons and synaptic circuits with subtle specializations in different brain regions and species. Determining the design and function of this microcircuitry is therefore of paramount importance to understanding normal and abnormal higher brain function. Recent advances in recording synaptically-coupled neurons has allowed rapid dissection of the neocortical microcircuitry thus yielding a massive amount of quantitative anatomical, electrical and gene expression data on the neurons and the synaptic circuits that connect the neurons. Due to the availability of the above mentioned data, it has now become imperative to database the neurons of the microcircuit and their synaptic connections. The NEOBASE project, aims to archive the neocortical microcircuit data in a manner that facilitates development of advanced data mining applications, statistical and bioinformatics analyses tools, custom microcircuit builders, and visualization and simulation applications. The database architecture is based on ROOT, a software environment that allows the construction of an object oriented database with numerous relational capabilities. The proposed architecture allows construction of a database that closely mimics the architecture of the real microcircuit, which facilitates the interface with virtually any application, allows for data format evolution, and aims for full interoperability with other databases. NEOBASE will provide an important resource and research tool for studying the microcircuit basis of normal and abnormal neocortical function. The database will be available to local as well as remote users using Grid based tools and technologies.

Animals↗

Mutated or non-mutated? Which database to choose when determining the IgVH hypermutation status in chronic lymphocytic leukemia?

It has been accepted that the hypermutation status of immunoglobulin heavy chain genes (IgVH) is one of the most important independent prognostic factors in chronic lymphocytic leukemia (CLL). According to the degree ofIgVH hypermutaion, CLL patients can be stratified into prognostic groups. Given the impact ofIgVH mutation status on clinical setting, it has become highly desirable to standardize the laboratory methodologies used for IgVH mutation status determination. To check the reliability of our laboratory results, we performed a random interlaboratory testing. From 10 CLL samples tested, in 9 cases identical results were obtained in both laboratories. In one case, the result was discordant. The discrepancy was caused by theIgVH database used. This finding prompted us to double-check our cohort of 624 CLL patients, using IgBLAST and IMGT databases. The results showed 7.5% (47/624) discrepancies between both databases. In 21 out of 47 cases, the degree of hypermutation has changed in regard to the database used, resulting in major changes in the prognostic subgroup. Other irregularities between both databases were identified, with yet to be determined significance. In the light of presented data we would like to stress the necessity to identify/compile the most comprehensiveIgVH database to be used for the determination ofIgVH mutation status in CLL.

Databases, Genetic↗

u-Genome: a database on genome design in unicellular genomes.

Unicellular eukaryotes were among the first ones to be selected for complete genome sequencing because of the small size of their genomes and their interactions with humans and a broad range of animals and plants. Currently, ten completely sequenced unicellular genome sequences have been publicly released and as the number of available unicellular genomes increases, comparative genomics analysis within this group of organisms becomes more and more instructive. However, such an analysis is difficult to carry out without a suitable platform gathering not only the original annotations but also relevant information available in public databases or obtained by applying common bioinformatics methods. With the aim of solving these difficulties, we have developed a web-accessible database named u-Genome, the unicellular genome design database. The database is unique in featuring three datasets namely (1) orthologous proteins (2) paralogous proteins and (3) statistical distributions on exons, introns, intergenic DNA and correlations between them. A tool, Uniview, designed to visualize the gene structures for individual genes in the genome is also integrated. This database is of importance in understanding unicellular genome design and architecture and evolution related studies. The database is available through a web interface at http://sege.ntu.edu.sg/wester/ugenome.

Animals↗

A nation's genes for a cure to cancer: evolving ethical, social and legal issues regarding population genetic databases.

The advent of the human genome sequence has focused research on understanding underlying genetic links to complex diseases such as cancer, asthma and heart disease. In the past few years, individual countries, such as Iceland, Estonia, Singapore and the United Kingdom, have created national databases of their citizens' DNA for comparative research. Most recently, an international consortium including Nigeria, Japan, China and the United States launched a $100 million project called the International HapMap to map the human genome according to haplotypes, blocks of DNA that contain genetic variation. Such population genetic databases present challenging ethical, social and legal issues, yet regulation of genetic information has developed sporadically, from region to region, without a consistent international standard. Without a clear understanding of the consequences of genetic research in terms of individual and community-wide discrimination and stigmatization, genetic databases raise concerns about the protection of genetic information. This Note provides a survey of the evolving landscape of population genetic databases as a legislative and public policy tool for national and international regulators. It compares different approaches to regulating the collection and use of population genetic databases in order to understand what areas of consensus are formulating a foundation for an international standard. As the first population genetics project that will span multiple countries for the collection of DNA, the International HapMap has the potential to become an influential standard for the protection of population genetic information. This Note highlights issues among the national databases and the HapMap project that raise ethical, social and legal concerns for the future and recommends further protections for both individual donors and community interests.

Access to Information↗

[Estimation of a nationwide statistics of hernia operation applying data mining technique to the National Health Insurance Database].

OBJECTIVES: The aim of this study is to develop a methodology for estimating a nationwide statistic for hernia operations with using the claim database of the Korea Health Insurance Cooperation (KHIC). METHODS: According to the insurance claim procedures, the claim database was divided into the electronic data interchange database (EDI_DB) and the sheet database (Paper_DB). Although the EDI_DB has operation and management codes showing the facts and kinds of operations, the Paper_DB doesn't. Using the hernia matched management code in the EDI_DB, the cases of hernia surgery were extracted. For drawing the potential cases from the Paper_DB, which doesn't have the code, the predictive model was developed using the data mining technique called SEMMA. The claim sheets of the cases that showed a predictive probability of an operation over the threshold, as was decided by the ROC curve, were identified in order to get the positive predictive value as an index of usefulness for the predictive model. RESULTS: Of the claim databases in 2004, 14,386 cases had hernia related management codes with using the EDI system. For fitting the models with applying the data mining technique, logistic regression was chosen rather than the neural network method or the decision tree method. From the Paper_DB, 1,019 cases were extracted as potential cases. Direct review of the sheets of the extracted cases showed that the positive predictive value was 95.3%. CONCLUSIONS: The results suggested that applying the data mining technique to the claim database in the KHIC for estimating the nationwide surgical statistics would be useful from the aspect of execution and cost-effectiveness.

Adolescent↗

A novel management database in obstetrics and gynaecology to introduce the electronic healthcare record and improve the clinical audit process.

OBJECTIVES: To design a system capable of recording complete and accurate electronic patient records with respect to obstetrics and gynaecology, with the ability to perform instant statistically summaries of data. BACKGROUND: Electronic patient records have been shown to provide numerous benefits for the clinician, with respect to patient consultation, accurate recording of data, medical audit and statistical analysis. In Northern Ireland there is no database designed to cover all the major clinical aspects of obstetrics and gynaecology. This project incorporates all aspects of obstetrics and gynaecology into a single database. METHODS: Database designed using Filemaker pro 7, Macromedia Fireworks 8, and Microsoft photodraw. Problems specific to obstetrics and gynaecology included recording multiple pregnancy data, and the lack of a unique patient number (the current system in Northern Ireland gives patients a unique hospital number, and a separate maternity number for each pregnancy). Linking all of these sources was a major component of this database. The database contains many intrinsic tabulations, relationships, programming scripts and calculations to combine files and calculate important statistical information for clinicians automatically. RESULTS: A successful audit of delivery statistics for December 05 was performed using the system. Several additional audits are currently under completion using the database. The major audit, completion date end April 06, is a 5 month summary of delivery data (Dec-April) based on mode of delivery, Robson groups, and Caesarian -Section rate among specified patient sub-groups. CONCLUSION: The system has been successful in its initial stages with obvious improvements to the medical audit process already apparent. The system should prove to be a valuable addition to the department and ultimately improve patient care. The ability to provide instant access to clinical data and statistics will simplify and improve the audit process, improving clinical governance. The management of the OB/GYN department should benefit greatly.

Computer Communication Networks↗

Creating a resource database for nursing service administration.

In response to the current information explosion in nursing service administration (NSA), the authors felt a need to collect and organize available resources for use by their faculty and graduate students. An electronic database was developed to facilitate the use of the collected print and software resources. This article describes the creation of the NSA Resource Database from the time the need for it was realized to its completion. There is discussion regarding the criteria used for writing the database, what the database screens look like and why and what the database contains. The article also discusses the use and users of the NSA Resource Database to date.

Databases, Bibliographic↗

Developing drug-use indicators with a computerized drug database and a personal computer software package.

The use of a multihospital drug- and patient-database system, a personal computer (PC), and a standard PC software package to monitor drug-use indicators is described, and a five-step method for analyzing a set of data is presented. An integrated spread-sheet, database, and graphics program (Lotus 1-2-3), which is compatible with an IBM PC, can manipulate data obtained from a multihospital database system. To demonstrate the utility of this system, a previously published procedure for analyzing drug-use indicators (e.g., length of stay, drug cost per patient, number of drugs received) for patients in two diagnosis-related groups was repeated using the database and PC software. The records of patients in each DRG were randomly selected from the database. The following steps were applied to the data: (1) data (as a whole) were characterized statistically, (2) data were examined to identify subgroups of interest, (3) groups of data were characterized statistically and compared with each other, (4) the effect of changing the characteristics of one or more subgroups was predicted, and (5) the results of the data manipulations were presented in tables and graphs. A multihospital database can serve as a source for obtaining large quantities of hospital-specific data. These data can be manipulated and presented in tabular or graphic form on a personal computer and used by hospitals to monitor various drug-use indicators.

Data Display↗