Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Future need for development of occupational exposure databases in Hungary.

The data on the occupational exposure measured by the hygienic network laboratories in Hungary were collected during the last 20 years. The data refer to the air pollutant chemicals and to the biological monitoring. The structure of the database of occupational exposure that has been used in the last decade is intended to be changed according to the guideline of the European Working Group on Exposure Registers in Europe. The current databases on the occupational exposure contain data only related to the substances responsible for air pollution. It would be desirable to complement the database with data from the biological monitoring, which also characterize the occupational exposure. It would then be possible to harmonize two distinct databases, and the risk assessment of the employees would become more thorough. This practice would require close cooperation of several organizations. It is desirable to set minimum requirements against the measurement techniques, thus, giving results that are accepted in the exposure database. It is necessary to encourage the use of direct reading instruments in collecting the exposure data and to complement the requirements of the strategies of data collection and the system of evaluation.

Databases, Factual↗

Cochrane systematic reviews in acupuncture: methodological diversity in database searching.

BACKGROUND: Since the early 1970s, the efficacy of acupuncture for treating clinical conditions has been evaluated in several hundred randomized trials. Results from these trials have been synthesized in systematic reviews. A well-designed systematic review provides the highest level of evidence for establishing the efficacy of a clinical intervention. OBJECTIVES: The present study assesses the source of original literature contributing to Cochrane reviews on acupuncture. Databases searched to retrieve original studies are evaluated. The distribution of controlled trials in acupuncture across different topic areas and journals, the ability of the reviews to provide conclusive results, and the proportion of original studies indexed with MEDLINE are evaluated. METHODS: Systematic reviews on acupuncture were extracted from the Cochrane Database of Systematic Reviews. The key search term used was "acupuncture." When more than one systematic review was retrieved on the same topic, the most recent review was included. Indexing of individual clinical trials with MEDLINE was searched using the Single Citation Matcher in PubMed. RESULTS: A total of 94 papers were retrieved from the Cochrane database, of which 10 were included in the analysis. The most common subject areas were related to chronic pain. Considerable heterogeneity was observed in the number of databases searched (median 5, range 3-12). A total of 69% (74/108) papers were indexed with PubMed. Only 13% (14/108) of the papers were published in the primary acupuncture journals. Conclusive statements about the efficacy of acupuncture were made in only 2 of the 10 systematic reviews. CONCLUSIONS: Considerable methodological diversity exists in the comprehensiveness of database searches for Cochrane systematic reviews on acupuncture. This diversity makes the reviews prone to bias and adds another layer of complexity in interpreting the acupuncture literature.

Abstracting and Indexing↗

A conceptual database model for genomic research.

We describe a conceptual model for genome databases that facilitates the process of building, maintaining, and disseminating physically anchored genetic linkage maps. The model has been implemented as a relational database at the Roman L. Hruska U.S. Meat Animal Research Center (MARC). Development of consensus maps using disparate data from different reference pedigrees or laboratories is supported. The model is of use to quantitative and population geneticists interested in loci that affect phenotypes and marker-assisted selection, and it is sufficiently flexible for centralized, species genome databases facilitating comparative mapping. The MARC genome database is used to assemble, maintain, and disseminate physically anchored genetic linkage maps for cattle, swine, and sheep currently based on more than 100,000 genotypes from 1,000 markers. Integrated with linkage analysis software, this database permits frequent updates of physically anchored genetic linkage maps.

Algorithms↗

Tracking the epidemiology of human genes in the literature: the HuGE Published Literature database.

Completion of the human genome sequence has inspired a new wave of epidemiologic studies on the prevalence of gene variants and their associations with diseases in human populations. In 2001, the Human Genome Epidemiology (HuGE) Network launched the HuGE Published Literature database (HuGE Pub Lit), a searchable, online knowledge base of published, population-based epidemiologic studies of human genes. The database contains links to PubMed articles and can be searched by gene, disease, interacting factor, type of study design or analysis, or any combination of terms in these categories. The search output contains a link to each identified article, along with a table summarizing key features of the reported study. As of September 6, 2005, some 17,665 articles were indexed in the database. Most described gene-disease associations (86%); fewer evaluated gene-gene or gene-environment interactions (17%), the prevalence of gene variants (10%), or genetic tests (3%). Although not comprehensive, this database is a unique tool for epidemiologic researchers and others concerned with the role of genetic variation in population health. Here, the authors provide an overview of the database and its characteristics and uses.

Centers for Disease Control and Prevention, U.S.↗

Relational database for drug-use review of Tennessee Medicaid claims.

The development of a relational database from Tennessee Medicaid files for the purpose of retrospective drug-use review (DUR) and application of the database for DUR of angiotensin-converting-enzyme inhibitors (ACEIs) are described. Computer queries were designed to create profiles of physicians' or pharmacies' experiences from claims data and other Medicaid data. Outlying patients (patients for whom at least one DUR criterion was unmet) were grouped according to their physicians or pharmacies. Thresholds for defining outlying physicians and pharmacies (i.e., those with more than a specified number of outlying patients) were based on the provider population instead of the patient population as a whole; aggregating patient outliers by provider allowed trends of inappropriate practices to be detected. As the threshold for outlying providers rose, the number of such providers fell, as did the number of outlying patients with whom they were associated. Stratification of the outliers by provider for specific drug-drug interactions and drug-disease complications afforded the option to set individual thresholds for outlying providers based on individual subsets; for example, for ACEIs, a threshold of greater than five patient outliers could be set for the criterion of no concurrent potassium supplements and a threshold of greater than three, for the criterion of no unmonitored concurrent lithium therapy. Tennessee patients formerly covered by Medicaid are now enrolled in managed care plans, and the flexibility of the database has allowed it to be modified accordingly. The relational database allows flexibility in the analysis of certain patterns of drug use. Such a database may be useful to other Medicaid programs that are converting to managed care models.

Angiotensin-Converting Enzyme Inhibitors↗

Evaluating a benchmarking database and identifying cost reduction opportunities by diagnosis-related group.

Pharmacy cost data from the University HealthSystem Consortium (UHC) Clinical Database for specific diagnosis-related groups (DRGs) were reviewed to assess their applicability to a university medical center and to identify opportunities to reduce costs. UHC headquarters was contacted by telephone to determine UHC's data collection methods. Pharmacy costs for DRG 302 (kidney transplant) at the University of Kansas Medical Center (KUMC) were compared with the costs shown in the UHC Clinical Database. Appropriate drug use for DRGs 302 and 480 (liver transplant) was assessed by contacting transplant pharmacists and pharmacy administrators at the five top-performing hospitals (in terms of cost per DRG) as listed in the UHC database to find opportunities for reducing pharmacy costs. KUMC's actual pharmacy costs for DRG 302 ($4635) were 46% lower than those listed in the UHC Clinical Database ($8546). There was a disparity between the amount of both intravenous immune globulin (IVIG) and lymphocyte immune globulin used by KUMC and the top-performing hospitals. Guidelines for use of IVIG, acyclovir, and azathioprine in liver transplant patients at KUMC were revised. A potential cost saving of $53,000 was identified in relation to the use of lymphocyte immune globulin in kidney transplant patients. Data in the UHC Clinical Database were not representative of pharmacy costs at a university medical center for DRG 302 (kidney transplant), overstating pharmacy costs by 46%; benchmarking was found to be a useful tool for identifying opportunities for reducing costs.

Benchmarking↗

Evaluation of electronic databases used to identify solid oral dosage forms.

The ability of electronic drug identification databases to identify solid oral dosage forms by their imprint codes was studied. The following seven commercially available electronic drug identification databases were selected to identify 500 solid oral dosage forms by their imprint codes: Clinical Pharmacology (Gold Standard Media, Tampa, FL), eFacts (Facts and Comparison, St. Louis, MO), Ident-A-Drug (Therapeutic Research, Stockton, CA), Identidex (Micromedex, Greenwood Village, CO), Clinical Reference Library (Lexi-Comp, Hudson, OH), Physicians' Desk Reference (PDR) Electronic Library (Medical Economics, Montvale, NJ), and RxList (RxList LLC, San Francisco, CA), Chi-square test was used to compare the percentages of medications identified by each of the seven electronic references. The ability of the databases to identify medication by specific characteristics, such as brand name versus generic, prescription versus nonprescription, commercially available for more than one year versus less than one year, colored versus white drug products, and controlled versus noncontrolled substances was evaluated. A logistic regression model was used to determine the probability of a drug product being identified by one of the electronic references based on these characteristics. All seven electronic databases combined identified 95.6% of the unknown medications by imprint code, color, shape, and scoring. Ident-A-Drug and Identidex identified the most drugs. The PDR Electronic Library and Facts and Comparisons Identified the least number of drugs. Solid oral dosage forms more likely to be identified were those that were on the market for more than a year, brand-name products, and prescription medications. Generic products on the market for less than a year and nonprescription products were particularly difficult to identify. A combination of electronic drug identification databases provides the best method of drug identification in an institutional setting.

Capsules↗

Database searching with DNA and protein sequences: an introduction.

This review of sequence database searching aims to set out current practice in the area, in order to give practical guidelines to the experimental biologist. It describes the basic principles behind the programs and enumerates the range of databases available in the public domain. Of these, the most important are the equivalent DNA databases European Molecular Biology Laboratory (EMBL), GenBank and DNA Databank of Japan (DDBJ), and the protein databases Swiss-Prot and TrEMBL. The commonly used BLAST and FASTA algorithms are described in detail and alternative approaches mentioned briefly. Scoring matrices used to compare amino acid types during protein database searches are compared, with an emphasis on the PAM and BLOSUM series of observed substitution matrices.

Algorithms↗

The role of pattern databases in sequence analysis.

In the wake of the numerous now-fruitful genome projects, we are entering an era rich in biological data. The field of bioinformatics is poised to exploit this information in increasingly powerful ways, but the abundance and growing complexity both of the data and of the tools and resources required to analyse them are threatening to overwhelm us. Databases and their search tools are now an essential part of the research environment. However, the rate of sequence generation and the haphazard proliferation of databases have made it difficult to keep pace with developments. In an age of information overload, researchers want rapid, easy-to-use, reliable tools for functional characterisation of newly determined sequences. But what are those tools? How do we access them? Which should we use? This review focuses on a particular type of database that is increasingly used in the task of routine sequence analysis--the so-called pattern database. The paper aims to provide an overview of the current status of pattern databases in common use, outlining the methods behind them and giving pointers on their diagnostic strengths and weaknesses.

Amino Acid Motifs↗

Searching the expressed sequence tag (EST) databases: panning for genes.

The genomes of living organisms contain many elements, including genes coding for proteins. The portions of the genes expressed as mature mRNA, collectively known as the transcriptome, represent only a small part of the genome. The expressed sequence tag (EST) databases contain an increasingly large part of the transcriptome of many species. For this reason, these databases are probably the most abundant source of new coding sequences available today. However, the raw data deposited in the EST databases are to a large extent unorganised, unannotated, redundant and of relatively low quality. This paper reviews some of the characteristics of the EST data, and the methods that can be used to find novel protein sequences within them. It also documents a collection of databases, software and web sites that can be useful to biologists interested in mining the EST databases over the Internet, or in establishing a local environment for such analyses.

Algorithms↗

Plant genome databases: from references to inference tools.

Plant genome databases play an important role in the archiving and dissemination of data arising from the international genome projects. Recent developments in bioinformatics, such as new software tools, programming languages and standards, have produced better access across the Internet to the data held within them. An increasing emphasis is placed on data analysis and indeed many resources now provide tools allied to the databases, to aid in the analysis and interpretation of the data. However, a considerable wealth of information lies untapped by considering the databases as single entities and will only be exploited by linking them with a wide range of data sources. Data from research programs such as comparative mapping and germplasm studies may be used as tools, to gain additional knowledge but without additional experimentation. To date, the current plant genome databases are not yet linked comprehensively with each other or with these additional resources, although they are clearly moving toward this. Here, the current wealth of public plant genome databases is reviewed, together with an overview of initiatives underway to bind them to form a single plant genome infrastructure.

Computational Biology↗

CLEANUP: a fast computer program for removing redundancies from nucleotide sequence databases.

A key concept in comparing sequence collections is the issue of redundancy. The production of sequence collections free from redundancy is undoubtedly very useful, both in performing statistical analyses and accelerating extensive database searching on nucleotide sequences. Indeed, publicly available databases contain multiple entries of identical or almost identical sequences. Performing statistical analysis on such biased data makes the risk of assigning high significance to non-significant patterns very high. In order to carry out unbiased statistical analysis as well as more efficient database searching it is thus necessary to analyse sequence data that have been purged of redundancy. Given that a unambiguous definition of redundancy is impracticable for biological sequence data, in the present program a quantitative description of redundancy will be used, based on the measure of sequence similarity. A sequence is considered redundant if it shows a degree of similarity and overlapping with a longer sequence in the database greater than a threshold fixed by the user. In this paper we present a new algorithm based on an "approximate string matching' procedure, which is able to determine the overall degree of similarity between each pair of sequences contained in a nucleotide sequence database and to generate automatically nucleotide sequence collections free from redundancies.

Algorithms↗

Automated protein sequence database classification. I. Integration of compositional similarity search, local similarity search, and multiple sequence alignment.

MOTIVATION: Genome sequencing projects require the periodic application of analysis tools that can classify and multiply align related protein sequence domains. Full automation of this task requires an efficient integration of similarity and alignment techniques. RESULTS: We have developed a fully automated process that classifies entire protein sequence databases, resulting in alignment of the homologous sequences. The successive steps of the procedure are based on compositional and local sequence similarity searches followed by multiple sequence alignments. Global similarities are detected from the pairwise comparison of amino acid and dipeptide compositions of each protein. After the elimination of all but one sequence from each detected cluster of closely related proteins, the remaining sequences are compiled in a suffix tree which is self-compared to detect local sequence similarities. Sets of proteins which share similar sequence segments are then weighted according to their closeness and multiply aligned using a fast hierarchical dynamic programming algorithm. Computational strategies were devised to minimize computer processing time and memory space requirements. The accuracy of the sequence classifications has been evaluated for 12 462 primary structures distributed over 341 known families. The percentage of sequences with missed or incorrect family assignments was 6.8% on the test set. This low error level is only twice that of the manually constructed PROSITE database ( 3.4% ) and is substantially better than that found for the automatically built PRODOM database ( 34.9% ). AVAILABILITY: The resulting database, called DOMO, is available through database search routine SRS at Infobiogen (http://www.infobiogen.fr/srs5/), EBI (http://srs.ebi.ac.uk:5000/) and EMBL (http://www.embl-heidelberg.de/srs5/) World Wide Web sites. CONTACT: gracy@infobiogen.fr

Algorithms↗

Searching DNA databases for similarities to DNA sequences: when is a match significant?

MOTIVATION: Searching DNA sequences against a DNA database is an essential element of sequence analysis. However, few systematic studies have been carried out to determine when a match between two DNA sequences has biological significance and this is limiting the use that can be made of DNA searching algorithms. RESULTS: A test set of DNA sequences has been constructed consisting of artificially evolved and real sequences. This set has been used to test various database searching algorithms (BLAST, BLAST2, FASTA and Smith-Waterman) on a subset of the EMBL database. The results of this analysis have been used to determine the sensitivity and coverage of all of the algorithms. Guidelines have been produced which can be used to assess the significance of DNA database search results. The Smith-Waterman algorithm was shown to have the best coverage, but the worst sensitivity, whereas the default BLASTN algorithm (word length set to 11) was shown to have good sensitivity, but poor coverage. A sensible compromise between speed, sensitivity and coverage can be obtained using either the FASTA or BLAST (word length set to 6) algorithms. However, analysis of the results also showed that no algorithm works well when the length of the probe sequence is <200 bases. In general, matches can accurately be identified between coding regions of DNA sequences when there is >35% sequence identity between the corresponding proteins. Searching a DNA sequence against a DNA sequence database can, therefore, be a useful tool in sequence analysis. AVAILABILITY: The test sets used are available via anonymous ftp from mbisg2.sbc.man.ac.uk in the directory /pub/cabios/testdata/ CONTACT: I.Anderson@stud.man.ac.uk; abrass@man.ac.uk

Algorithms↗

The TRANSPATH signal transduction database: a knowledge base on signal transduction networks.

UNLABELLED: TRANSPATH is an information system on gene-regulatory pathways, and an extension module to the TRANSFAC database system (Wingender et al., Nucleic Acids Res., 28, 316-319, 2000). It focuses on pathways involved in the regulation of transcription factors in different species, mainly human, mouse and rat. Elements of the relevant signal transduction pathways like complexes, signaling molecules, and their states are stored together with information about their interaction in an object-oriented database. The database interface provides clickable maps and automatically generated pathway cascades as additional ways to explore the data. All information is validated with references to the original publications. Also, references to other databases are provided (TRANSFAC, SWISS-PROT, EMBL, PubMed and others). AVAILABILITY: The database is available over (http://transpath.gbf.de) for interactive perusal. As an exchange format for the data, eXtensible Markup Language (XML) flatfiles and a Document Type Definition (DTD) are provided.

Algorithms↗

Semi-automated update and cleanup of structural RNA alignment databases.

UNLABELLED: We have developed a series of programs which assist in maintenance of structural RNA databases. A main program BLASTs the RNA database against GenBank and automatically extends and realigns the sequences to include the entire range of the RNA query sequences. After manual update of the database, other programs can examine base pair consistency and phylogenetic support. The output can be applied iteratively to refine the structural alignment of the RNA database. Using these tools, the number of potential misannotations per sequence was reduced from 20 to 3 in the Signal Recognition Particle RNA database. AVAILABILITY: A quick-server and programs are available at http://www.bioinf.au.dk/rnadbtool/

Base Sequence↗

Statistical analysis in dBASE-compatible databases.

Database management in clinical and experimental research often requires statistical analysis of the data in addition to the usual functions for storing, organizing, manipulating and reporting. With most database systems, transfer of data to a dedicated statistics package is a relatively simple task. However, many statistics programs lack the powerful features found in database management software. dBASE IV and compatible programs are currently among the most widely used database management programs. d4STAT is a utility program for dBASE, containing a collection of statistical functions and tests for data stored in the dBASE file format. By using d4STAT, statistical calculations may be performed directly on the data stored in the database without having to exit dBASE IV or export data. Record selection and variable transformations are performed in memory, thus obviating the need for creating new variables or data files. The current version of the program contains routines for descriptive statistics, paired and unpaired t-tests, correlation, linear regression, frequency tables, Mann-Whitney U-test, Wilcoxon signed rank test, a time-saving procedure for counting observations according to user specified selection criteria, survival analysis (product limit estimate analysis, log-rank test, and graphics), and normal t and chi-squared distribution functions.

Database Management Systems↗

Redesigning, implementing and integrating Escherichia coli genome software tools with an object-oriented database system.

This paper reports our exploratory work to redesign, implement and integrate a collection of genome software tools with an object-oriented database system. Our software tools deal with genome data from Escherichia coli K-12, a bacterium that has been studied intensively and provides richer data sets than any other living organism. The object-oriented DBMS used for the integration is ONTOS, a commercial object-oriented system from Ontologic Inc. This redesign and implementation task was performed in two steps. First, C programs were converted into C++, and then the C++ version programs were modified and integrated with an object-oriented modeling of the data to form an ONTOS database application. The first step helps us develop a conceptual view for a DBMS-independent object-oriented construct. The second step elucidates what additional DBMS-dependent modification steps are needed to provide persistency to the objects. Examples are included to illustrate steps of the redesign and implementation. Overall, the outcome of this project demonstrates that programs and data can be successfully integrated with an object-oriented database, while providing the objects with persistency and shareability. This paper includes discussions using concrete examples on what advantage the object-oriented database approach provides over the relational database approach.

Base Sequence↗