Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Database of mRNA gene expression profiles of multiple human organs.

Genome-wide expression profiling of normal tissue may facilitate our understanding of the etiology of diseased organs and augment the development of new targeted therapeutics. Here, we have developed a high-density gene expression database of 18,927 unique genes for 158 normal human samples from 19 different organs of 30 different individuals using DNA microarrays. We report four main findings. First, despite very diverse sample parameters (e.g., age, ethnicity, sex, and postmortem interval), the expression profiles belonging to the same organs cluster together, demonstrating internal stability of the database. Second, the gene expression profiles reflect major organ-specific functions on the molecular level, indicating consistency of our database with known biology. Third, we demonstrate that any small (i.e., n approximately 100), randomly selected subset of genes can approximately reproduce the hierarchical clustering of the full data set, suggesting that the observed differential expression of >90% of the probed genes is of biological origin. Fourth, we demonstrate a potential application of this database to cancer research by identifying 19 tumor-specific genes in neuroblastoma. The selected genes are relatively underexpressed in all of the organs examined and belong to therapeutically relevant pathways, making them potential novel diagnostic markers and targets for therapy. We expect this database will be of utility for developing rationally designed molecularly targeted therapeutics in diseases such as cancer, as well as for exploring the functions of genes.

Cluster Analysis↗

The SUPERFAMILY database in structural genomics.

The SUPERFAMILY hidden Markov model library representing all proteins of known structure predicts the domain architecture of protein sequences and classifies them at the SCOP superfamily level. This analysis has been carried out on all completely sequenced genomes. The ways in which the database can be useful to crystallographers is discussed, in particular with a view to high-throughput structure determination. The application of the SUPERFAMILY database to different target-selection strategies is suggested: novel folds, novel domain combinations and targeted attacks on genomes. Use of the database for more general inquiry in the context of structural studies is also explained. The database provides evolutionary relationships between target proteins and other proteins of known structure through the SCOP database, genome assignments and multiple sequence alignments.

Amino Acid Sequence↗

Likelihood ratios for evaluating DNA evidence when the suspect is found through a database search.

A crime has been committed, and a DNA profile of the perpetrator is obtained from the crime scene. A suspect with a matching profile is found. The problem of evaluating this DNA evidence in a forensic context, when the suspect is found through a database search, is analysed through a likelihood approach. The recommendations of the National Research Council of the U.S. are derived in this setting as the proper way of evaluating the evidence when finiteness of the population of possible perpetrators is not taken into account. When a finite population of possible perpetrators may be assumed, it is possible to take account of the sampling process that resulted in the actual database, so one can deal with the problem where a large proportion of the possible perpetrators belongs to the database in question. It is shown that the last approach does not in general result in a greater weight being assigned to the evidence, though it does when a sufficiently large amount of the possible perpetrators are in the database. The value of the likelihood ratio corresponding to the probable cause setting constitutes an upper bound for this weight, and the upper bound is only attained when all but one of the possible perpetrators are in the database.

Biometry↗

Generic obstetric database systems are unreliable for reporting the hypertensive disorders of pregnancy.

BACKGROUND: Obstetric outcome data can be accessed by a variety of different methods in New South Wales, Australia, including Diagnosis-Related Group (DRG) summaries and the Midwives Data Collection (MDC). The accuracy of the reporting of the obstetric complications encompassed by the hypertensive disorders of pregnancy (HDP) has been doubted owing to the inconsistency and confusing coding categories available in both coding systems. AIM: To test that there would be no disagreement in coding between DRG coding, MDC coding and a disorder-specific database. METHODS: A prospective disorder-specific database, namely the Hypertensive Disorders of Pregnancy Database (HDPDB), was maintained for 6 months and diagnoses were compared with the DRG and MDC coding systems. Medical records of all women (n = 230) who received a HDP coding during this period were examined and recoded using the Australasian Society for the Study of Hypertension in Pregnancy diagnostic groupings. The HDPDB was the gold standard for the comparison of coding. RESULTS: Significant coding errors were found in both systems available. Sixty-four percent (P < 0.001) of medical records were incorrectly coded in the DRG coding system and 56% (P < 0.001) were incorrectly coded in the MDC. CONCLUSIONS: Current database systems are unreliable for recording maternal medical conditions, such as hypertension, and accuracy can only be assured with the use of a disorder-specific database, such as the HDPDB.

Data Interpretation, Statistical↗

A relational database for storing clinical information.

In medicine, we are constantly acquiring, recording, and analyzing data. Database programs give us a powerful tool to save and report our data. While spreadsheets and even word processing software can store simple datasets, the complexity of data storage needs in medicine often requires the use of a database program. Using most database software was difficult task for the casual user until recently, new, powerful programs were introduced by several developers. Microsoft Access, which is licensed for use with the purchase of many new computers, is an extremely powerful and relatively easy to use program written for the Windows operating systems. With this program, even novices to database programming can learn to construct useful database management systems.

Database Management Systems↗

Creating and analyzing a statewide nursing quality measurement database.

PURPOSE: To explicate a replicable methodology for designing and analyzing a large ongoing reliable and valid quality database to examine nurse staffing and patient care outcomes in acute care hospitals. DESIGN: Prospective nurse staffing, process of care, and patient outcomes data based on the American Nurses Association's (ANA) nursing quality indicators collected from a voluntary convenience sample at acute care hospitals in California with rolling-site accrual. METHODS: The ongoing CalNOC database development and repository project, the largest statewide effort of its kind in the United States (US), currently includes data on hospital nurse staffing, patient days, patient falls, pressure ulcer and restraint prevalence, registered nurse (RN) education, and patients' perceptions of satisfaction with care. FINDINGS: As of May 2003, the CalNOC database contained staffing data from 842 units in 134 acute care hospitals over 20 quarters from April 1998 to March 2003. The repository also included clinical outcome information on 34,262 reported patient falls, pressure ulcer prevalence data on 41,982 patient observations, and service outcome data on patient satisfaction from 26,461 patients. Participating hospitals receive quarterly reports allowing them to benchmark their own performance against other participating hospitals. CalNOC methods have been adapted and replicated by both the Military Nursing Outcomes Database and VA Nursing Outcomes Database projects, and CalNOC nursing-sensitive measures have been endorsed by the National Quality Forum. CONCLUSIONS: This working model for collecting reliable and valid data was derived from multiple hospitals across California. The data are the basis for studies to contribute to the development of evidence-based public policy, and for ongoing study of the effects of nurse staffing on clinical and service outcomes.

Accidental Falls↗

Use of health care databases in pharmacoepidemiology.

Pharmacoepidemiology is the study of the use and effects of medications in populations. Large health care databases are often used to address research questions within pharmacoepidemiology. This paper briefly describes the kinds of research questions that can be addressed using pharmacoepidemiology databases, provides an overview of pharmacoepidemiologic databases, describes some differences between medical records and administrative databases, discusses factors that should be considered when choosing a database for a particular study, and considers what the future holds.

Databases, Factual↗

Completeness of cause of injury coding in healthcare administrative databases in the United States, 2001.

OBJECTIVES: To determine the completeness of external cause of injury coding (E-coding) within healthcare administrative databases in the United States and to identify factors that contribute to variations in E-code reporting across states. DESIGN: Cross sectional analysis of the 2001 Healthcare Cost and Utilization Project (HCUP), including 33 State Inpatient Databases (SID), a Nationwide Inpatient Sample (NIS), and nine State Emergency Department Databases (SEDD). To assess state reporting practices, structured telephone interviews were conducted with the data organizations that participate in HCUP. RESULTS: The percent of injury records with an injury E-code was 86% in HCUP's nationally representative database, the NIS. For the 33 states represented in the SID, completeness averaged 87%, with more than half of the states reporting E-codes on at least 90% of injuries. In the nine states also represented in the SEDD, completeness averaged 93%. Twenty two states had mandates for E-code reporting, but only eight had provisions for enforcing the mandates. These eight states had the highest rates of E-code completeness. CONCLUSIONS: E-code reporting in administrative databases is relatively complete, but there is significant variation in completeness across the states. States with mandates for the collection of E-codes and with a mechanism to enforce those mandates had the highest rates of E-code reporting. Nine statewide ED data systems demonstrate consistently high E-coding completeness.

Cross-Sectional Studies↗

The predicted impact of coding single nucleotide polymorphisms database.

Nonsynonymous single nucleotide polymorphisms (nsSNP) have the potential to affect the structure or function of expressed proteins and are, therefore, likely to represent modifiers of inherited susceptibility. We have classified and catalogued the predicted functionality of nsSNPs in genes relevant to the biology of cancer to facilitate sequence-based association studies. Candidate genes were identified using targeted search terms and pathways to interrogate the Gene Ontology Consortium database, Kyoto Encyclopedia of Genes and Genomes database, Iobion's Interaction Explorer PathwayAssist Program, National Center for Biotechnology Information Entrez Gene database, and CancerGene database. A total of 9,537 validated nsSNPs located within annotated genes were retrieved from National Center for Biotechnology Information dbSNP Build 123. Filtering this list and linking it to 7,080 candidate genes yielded 3,666 validated nsSNPs with minor allele frequencies > or =0.01 in Caucasian populations. The functional effect of nsSNPs in genes with a single mRNA transcript was predicted using three computational tools-Grantham matrix, Polymorphism Phenotyping, and Sorting Intolerant from Tolerant algorithms. The resultant pool of 3,009 fully annotated nsSNPs is accessible from the Predicted Impact of Coding SNPs database at http://www.icr.ac.uk/cancgen/molgen/MolPopGen_PICS_database.htm. Predicted Impact of Coding SNPs is an ongoing project that will continue to curate and release data on the putative functionality of coding SNPs.

Algorithms↗

Meta-analysis of the p53 mutation database for mutant p53 biological activity reveals a methodologic bias in mutation detection.

PURPOSE: Analyses of the pattern of p53 mutations have been essential for epidemiologic studies linking carcinogen exposure and cancer. We were concerned by the inclusion of dubious reports in the p53 databases that could lead to controversial analysis prejudicial to the scientific community. EXPERIMENTAL DESIGN: We used the universal mutation database p53 database (21,717 mutations) combined with a new p53 mutant activity database (2,300 mutants) to perform functional analysis of 1,992 publications reporting p53 alterations. This analysis was done using a statistical approach similar to that of clinical meta-analyses. RESULTS: This analysis reveals that some reports of infrequent mutations are associated with almost normal activities of p53 proteins. These particular mutations are frequently found in studies reporting multiple mutations in one tumor, silent mutations, or lacking mutation hotspots. These reports are often associated with particular methodologies, such as nested PCR, for which key controls are not satisfactory. CONCLUSIONS: We show the importance of accurate functional analysis before inferring any genetic variation. The quality of the p53 databases is essential in order to prevent erroneous analysis and/or conclusions. The availability of functional data from our new p53 web site (http://p53.free.fr and http://www.umd.be:2072/) will allow functional prescreening to identify potential artifactual data.

Databases, Factual↗

Genome SEGE: a database for 'intronless' genes in eukaryotic genomes.

BACKGROUND: A number of completely sequenced eukaryotic genome data are available in the public domain. Eukaryotic genes are either 'intron containing' or 'intronless'. Eukaryotic 'intronless' genes are interesting datasets for comparative genomics and evolutionary studies. The SEGE database containing a collection of eukaryotic single exon genes is available. However, SEGE is derived using GenBank. The redundant, incomplete and heterogeneous qualities of GenBank data are a bottleneck for biological investigation in comparative genomics and evolutionary studies. Such studies often require representative gene sets from each genome and this is possible only by deriving specific datasets from completely sequenced genome data. Thus Genome SEGE, a database for 'intronless' genes in completely sequenced eukaryotic genomes, has been constructed. AVAILABILITY: http://sege.ntu.edu.sg/wester/intronless DESCRIPTION: Eukaryotic 'intronless' genes are extracted from nine completely sequenced genomes (four of which are unicellular and five of which are multi-cellular). The complete dataset is available for download. Data subsets are also available for 'intronless' pseudo-genes. The database provides information on the distribution of 'intronless' genes in different genomes together with their length distributions in each genome. Additionally, the search tool provides pre-computed PROSITE motifs for each sequence in the database with appropriate hyperlinks to InterPro. A search facility is also available through the web server. CONCLUSIONS: The unique features that distinguish Genome SEGE from SEGE is the service providing representative 'intronless' datasets for completely sequenced genomes. 'Intronless' gene sets available in this database will be of use for subsequent bio-computational analysis in comparative genomics and evolutionary studies. Such analysis may help to revisit the original genome data for re-examination and re-annotation.

Databases, Genetic↗

AtRTPrimer: database for Arabidopsis genome-wide homogeneous and specific RT-PCR primer-pairs.

BACKGROUND: Primer design is a critical step in all types of RT-PCR methods to ensure specificity and efficiency of a target amplicon. However, most traditional primer design programs suggest primers on a single template of limited genetic complexity. To provide researchers with a sufficient number of pre-designed specific RT-PCR primer pairs for whole genes in Arabidopsis, we aimed to construct a genome-wide primer-pair database. DESCRIPTION: We considered the homogeneous physical and chemical properties of each primer (homogeneity) of a gene, non-specific binding against all other known genes (specificity), and other possible amplicons from its corresponding genomic DNA or similar cDNAs (additional information). Then, we evaluated the reliability of our database with selected primer pairs from 15 genes using conventional and real time RT-PCR. CONCLUSION: Approximately 97% of 28,952 genes investigated were finally registered in AtRTPrimer. Unlike other freely available primer databases for Arabidopsis thaliana, AtRTPrimer provides a large number of reliable primer pairs for each gene so that researchers can perform various types of RT-PCR experiments for their specific needs. Furthermore, by experimentally evaluating our database, we made sure that our database provides good starting primer pairs for Arabidopsis researchers to perform various types of RT-PCR experiments.

Arabidopsis↗

SPODOBASE: an EST database for the lepidopteran crop pest Spodoptera.

BACKGROUND: The Lepidoptera Spodoptera frugiperda is a pest which causes widespread economic damage on a variety of crop plants. It is also well known through its famous Sf9 cell line which is used for numerous heterologous protein productions. Species of the Spodoptera genus are used as model for pesticide resistance and to study virus host interactions. A genomic approach is now a critical step for further new developments in biology and pathology of these insects, and the results of ESTs sequencing efforts need to be structured into databases providing an integrated set of tools and informations. DESCRIPTION: The ESTs from five independent cDNA libraries, prepared from three different S. frugiperda tissues (hemocytes, midgut and fat body) and from the Sf9 cell line, are deposited in the database. These tissues were chosen because of their importance in biological processes such as immune response, development and plant/insect interaction. So far, the SPODOBASE contains 29,325 ESTs, which are cleaned and clustered into non-redundant sets (2294 clusters and 6103 singletons). The SPODOBASE is constructed in such a way that other ESTs from S. frugiperda or other species may be added. User can retrieve information using text searches, pre-formatted queries, query assistant or blast searches. Annotation is provided against NCBI, UNIPROT or Bombyx mori ESTs databases, and with GO-Slim vocabulary. CONCLUSION: The SPODOBASE database provides integrated access to expressed sequence tags (EST) from the lepidopteran insect Spodoptera frugiperda. It is a publicly available structured database with insect pest sequences which will allow identification of a number of genes and comprehensive cloning of gene families of interest for scientific community. SPODOBASE is available from URL: http://bioweb.ensam.inra.fr/spodobase.

Animals↗

The SDH mutation database: an online resource for succinate dehydrogenase sequence variants involved in pheochromocytoma, paraganglioma and mitochondrial complex II deficiency.

BACKGROUND: The SDHA, SDHB, SDHC and SDHD genes encode the subunits of succinate dehydrogenase (succinate: ubiquinone oxidoreductase), a component of both the Krebs cycle and the mitochondrial respiratory chain. SDHA, a flavoprotein and SDHB, an iron-sulfur protein together constitute the catalytic domain, while SDHC and SDHD encode membrane anchors that allow the complex to participate in the respiratory chain as complex II. Germline mutations of SDHD and SDHB are a major cause of the hereditary forms of the tumors paraganglioma and pheochromocytoma. The largest subunit, SDHA, is mutated in patients with Leigh syndrome and late-onset optic atrophy, but has not as yet been identified as a factor in hereditary cancer. DESCRIPTION: The SDH mutation database is based on the recently described Leiden Open (source) Variation Database (LOVD) system. The variants currently described in the database were extracted from the published literature and in some cases annotated to conform to current mutation nomenclature. Researchers can also directly submit new sequence variants online. Since the identification of SDHD, SDHC, and SDHB as classic tumor suppressor genes in 2000 and 2001, studies from research groups around the world have identified a total of 120 variants. Here we introduce all reported paraganglioma and pheochromocytoma related sequence variations in these genes, in addition to all reported mutations of SDHA. The database is now accessible online. CONCLUSION: The SDH mutation database offers a valuable tool and resource for clinicians involved in the treatment of patients with paraganglioma-pheochromocytoma, clinical geneticists needing an overview of current knowledge, and geneticists and other researchers needing a solid foundation for further exploration of both these tumor syndromes and SDHA-related phenotypes.

Codon, Nonsense↗

Coupling computer-interpretable guidelines with a drug-database through a web-based system--The PRESGUID project.

BACKGROUND: Clinical Practice Guidelines (CPGs) available today are not extensively used due to lack of proper integration into clinical settings, knowledge-related information resources, and lack of decision support at the point of care in a particular clinical context. OBJECTIVE: The PRESGUID project (PREScription and GUIDelines) aims to improve the assistance provided by guidelines. The project proposes an online service enabling physicians to consult computerized CPGs linked to drug databases for easier integration into the healthcare process. METHODS: Computable CPGs are structured as decision trees and coded in XML format. Recommendations related to drug classes are tagged with ATC codes. We use a mapping module to enhance computerized guidelines coupling with a drug database, which contains detailed information about each usable specific medication. In this way, therapeutic recommendations are backed up with current and up-to-date information from the database. RESULTS: Two authoritative CPGs, originally diffused as static textual documents, have been implemented to validate the computerization process and to illustrate the usefulness of the resulting automated CPGs and their coupling with a drug database. We discuss the advantages of this approach for practitioners and the implications for both guideline developers and drug database providers. Other CPGs will be implemented and evaluated in real conditions by clinicians working in different health institutions.

Computer Graphics↗

A database of antimalarial drug resistance.

A large investment is required to develop, license and deploy a new antimalarial drug. Too often, that investment has been rapidly devalued by the selection of parasite populations resistant to the drug action. To understand the mechanisms of selection, detailed information on the patterns of drug use in a variety of environments, and the geographic and temporal patterns of resistance is needed. Currently, there is no publically-accessible central database that contains information on the levels of resistance to antimalaria drugs. This paper outlines the resources that are available and the steps that might be taken to create a dynamic, open access database that would include current and historical data on clinical efficacy, in vitro responses and molecular markers related to drug resistance in Plasmodium falciparum and Plasmodium vivax. The goal is to include historical and current data on resistance to commonly used drugs, like chloroquine and sulfadoxine-pyrimethamine, and on the many combinations that are now being tested in different settings. The database will be accessible to all on the Web. The information in such a database will inform optimal utilization of current drugs and sustain the longest possible therapeutic life of newly introduced drugs and combinations. The database will protect the valuable investment represented by the development and deployment of novel therapies for malaria.

Animals↗

Case mix, outcome and length of stay for admissions to adult, general critical care units in England, Wales and Northern Ireland: the Intensive Care National Audit & Research Centre Case Mix Programme Database.

INTRODUCTION: The present paper describes the methods of data collection and validation employed in the Intensive Care National Audit & Research Centre Case Mix Programme (CMP), a national comparative audit of outcome for adult, critical care admissions. The paper also describes the case mix, outcome and activity of the admissions in the Case Mix Programme Database (CMPD). METHODS: The CMP collects data on consecutive admissions to adult, general critical care units in England, Wales and Northern Ireland. Explicit steps are taken to ensure the accuracy of the data, including use of a dataset specification, of initial and refresher training courses, and of local and central validation of submitted data for incomplete, illogical and inconsistent values. Criteria for evaluating clinical databases developed by the Directory of Clinical Databases were applied to the CMPD. The case mix, outcome and activity for all admissions were briefly summarised. RESULTS: The mean quality level achieved by the CMPD for the 10 Directory of Clinical Databases criteria was 3.4 (on a scale of 1 = worst to 4 = best). The CMPD contained validated data on 129,647 admissions to 128 units. The median age was 63 years, and 59% were male. The mean Acute Physiology and Chronic Health Evaluation II score was 16.5. Mortality was 20.3% in the CMP unit and was 30.8% at ultimate discharge from hospital. Nonsurvivors stayed longer in intensive care than did survivors (median 2.0 days versus 1.7 days in the CMP unit) but had a shorter total hospital length of stay (9 days versus 16 days). Results for the CMPD were comparable with results from other published reports of UK critical care admissions. CONCLUSIONS: The CMP uses rigorous methods to ensure data are complete, valid and reliable. The CMP scores well against published criteria for high-quality clinical databases.

Adult↗

Relational databases: a transparent framework for encouraging biology students to think informatically.

We discuss how relational databases constitute an ideal framework for representing and analyzing large-scale genomic data sets in biology. As a case study, we describe a Drosophila splice-site database that we recently developed at Wesleyan University for use in research and teaching. The database stores data about splice sites computed by a custom algorithm using Drosophila cDNA transcripts and genomic DNA and supports a set of procedures for analyzing splice-site sequence space. A generic Web interface permits the execution of the procedures with a variety of parameter settings and also supports custom structured query language queries. Moreover, new analytical procedures can be added by updating special metatables in the database without altering the Web interface. The database provides a powerful setting for students to develop informatic thinking skills.

Algorithms↗