Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Adaptation of international nutrition databases and data-entry system tools to a specific population.

OBJECTIVE: To develop a nutritional dietary intake database based on available reliable international nutritional databases adapted to the local needs of a specific population. DESIGN: The Negev Nutritional Study (NNS) is a survey of a random sample of the Negev population regarding their dietary intake using 24-hour dietary recalls. A nutritional database for the Israeli population was developed based on adaptation and modification of the US Department of Agriculture's database. A data-entry system was developed based on the logic of the US Food Information Analysis System. The system was designed as bilingual (English and Hebrew). Local foods and recipes were collected during the NNS, which included 1465 24-hour diet interviews. RESULTS: During the course of the NNS, 383 basic Israeli recipes were constructed. In total 1362 Israeli products were added to the database, and each was given a code, specific gravity and portion size. Most of the added products were cereals and grains and dairy products. The added recipes were collected from the interviewees in the NNS and from the most popular cookbooks. CONCLUSIONS: This paper describes the process undertaken to develop an Israeli food composition database as well as the data-entry system. This knowledge may aid other research groups in developing a computerised, nation-specific nutritional database and data-entry system adapted to their own specific local needs.

Databases, Factual↗

Overview of the national spinal cord injury statistical center database.

OBJECTIVE: An evaluation of the history, design, and status of the database of the National Spinal Cord Injury Statistical Center (NSCISC) was undertaken to identify its continued relevance. RESEARCH DESIGN: A systematic review was conducted of goals, content, and quality control procedures, as well as its suitability and public availability for conducting future epidemiologic and health services research. RESULTS: The NSCISC database contains information on approximately 29,000 persons injured since 1973 and treated at any regional model spinal cord injury system within 1 year of injury. The NSCISC database is structured longitudinally with data collected at discharge, 1 year after injury, 5 years after injury, and every 5 years thereafter. The database includes information on demographics, injury severity, medical complications, surgical procedures, types and amounts of therapy, length of stay, charges, and both short-term and long-term treatment outcomes. Strengths include large sample size, use of valid and reliable measures, geographic and patient diversity, comprehensiveness, availability of long-term prospective follow-up information, good case identification, and rigorous quality control procedures. Limitations include lack of population basis, inclusion of only model system patients, losses to follow-up, and other missing data. Recent content additions include detailed information on each treatment phase, depression, substance abuse, environmental barriers to community integration, and patient identifying information. A process exists for researchers to gain access to the data. CONCLUSIONS: The database remains a valuable resource. Future plans include linkage to other databases to enhance research capability, a published research compendium, and development of a user's guide to facilitate database usage.

Databases as Topic↗

Exposure databases and exposure surveillance: promise and practice.

Based on recent developments in occupational health and a review of industry practices, it is argued that integrated exposure database and surveillance systems hold considerable promise for improving workplace health and safety. A foundation from which to build practical and effective exposure surveillance systems is proposed based on the integration of recent developments in electronic exposure databases, the codification of exposure assessment practice, and the theory and practice of public health surveillance. The merging of parallel, but until now largely separate, efforts in these areas into exposure surveillance systems combines unique strengths from each subdiscipline. The promise of exposure database and surveillance systems, however, is yet to be realized. Exposure surveillance practices in general industry are reviewed based on the published literature as well as an Internet survey of three prominent industrial hygiene e-mail lists. Although the benefits of exposure surveillance are many, relatively few organizations use electronic exposure databases, and even fewer have active exposure surveillance systems. Implementation of exposure databases and surveillance systems can likely be improved by the development of systems that are more responsive to workplace or organizational-level needs. An overview of exposure database software packages provides guidance to readers considering the implementation of commercially available systems. Strategies for improving the implementation of exposure database and surveillance systems are outlined. A companion report in this issue on the development and pilot testing of a workplace-level exposure surveillance system concretely illustrates the application of the conceptual framework proposed.

Databases, Factual↗

Ambulatory care databases for managed care organizations.

The uses, advantages, and limitations of ambulatory care databases are discussed, and processes for extracting and using the data are described. Claims databases allow health systems, including managed care organizations, to generate descriptive statistics on patients, providers, and diseases; to conduct comprehensive cost and resource-use analyses; and to build economic models of diseases. The use of health care databases has several advantages over clinical trials, including lower data collection costs, shorter times for analysis, larger numbers of patients, and less inconvenience to patients and providers. These databases allow the effectiveness of a treatment, instead of its efficacy, to be assessed and drug-switching patterns within disease categories to be detected. This information can be used for determining the cost and outcome implications of new treatments and formulary changes, as well as for monitoring disease management programs. Limitations of health care databases include the omission of services not covered by health plans, the risk of coding errors, the absence of indicators of disease severity, and the lack of data that would assist with clinical outcome analysis. Medical and pharmacy claims data do not contain all the information contained in patients' medical records. Despite their limitations, ambulatory care databases are useful for describing patient, provider, and disease characteristics. The databases are also useful for predicting and estimating the implications of a change in the formulary, measuring the effects of treatment guidelines, and monitoring disease management programs.

Ambulatory Care Information Systems↗

Pharmacy database for tracking drug costs and utilization.

A pharmacy database for tracking drug costs and physician prescribing trends is described. Accuracy problems plagued data systems used to make drug-use-policy decisions at a tertiary care teaching hospital because of structural deficiencies within the systems and their nonclinical orientation. To resolve these problems, a programmer analyst, a clinical supervisor, and a clinical pharmacist developed a hierarchical database of drug costs. The database was designed to be valid for tracking drug costs according to patterns of clinical use. Internal controls were created that could identify and correct cost-tabulation errors arising within the ordering, order-entry, and billing processes. The database was able to tabulate drug costs according to the clinical service on which the patient was being treated at the time so that reports could compare aggregate prescribing trends from one time period to another for the same service. Similarly, the database could track and report drug use by disease or financial classification. Flagging elements were introduced to the database for cancer chemotherapy and antimicrobial drug products to enable reporting by these categories and by therapeutic subcategories within the antimicrobial category. Routine monthly reports were distributed to end users. Development of a database for tracking drug costs and utilization allowed a teaching hospital to derive the cost of medications from billing-charge information and to report data to health care professionals on the basis of important factors like clinical services.

Academic Medical Centers↗

Evaluation of drug information databases for personal digital assistants.

PURPOSE: Core and supplemental drug information databases available for use with personal digital assistants (PDAs) were evaluated. METHODS: Ten core (or standalone) databases, six drug interaction analyzers, and three dietary supplement databases used with the Palm and Pocket PC operating systems were selected for study. The databases were rated for scope (the absence or presence of an answer to a drug information question), completeness (the comprehensiveness of an answer), and ease of use (the number of hypertext links needed to reach the desired answer). A total of 14 weighted categories, consisting of 146 and 30 drug questions for the core and supplemental databases, respectively, were used to determine the overall scores. RESULTS: The best overall performers were, in order of total scores, Lexi-Drugs Platinum, Tarascon Pocket Pharmacopoeia, ePocrates Rx Pro, and Clinical Pharmacology OnHand. The databases with the lowest composite scores were Triple i Prescribing Guide and A2Z Drugs. CONCLUSION: Drug information databases for PDAs varied in scope, completeness, and ease of use. The results may help clinicians find the most appropriate product for their practice setting.

Computers, Handheld↗

Saccharomyces genome database: underlying principles and organisation.

A scientific database can be a powerful tool for biologists in an era where large-scale genomic analysis, combined with smaller-scale scientific results, provides new insights into the roles of genes and their products in the cell. However, the collection and assimilation of data is, in itself, not enough to make a database useful. The data must be incorporated into the database and presented to the user in an intuitive and biologically significant manner. Most importantly, this presentation must be driven by the user's point of view; that is, from a biological perspective. The success of a scientific database can therefore be measured by the response of its users - statistically, by usage numbers and, in a less quantifiable way, by its relationship with the community it serves and its ability to serve as a model for similar projects. Since its inception ten years ago, the Saccharomyces Genome Database (SGD) has seen a dramatic increase in its usage, has developed and maintained a positive working relationship with the yeast research community, and has served as a template for at least one other database. The success of SGD, as measured by these criteria, is due in large part to philosophies that have guided its mission and organisation since it was established in 1993. This paper aims to detail these philosophies and how they shape the organisation and presentation of the database.

Databases, Nucleic Acid↗

An extensible network query unification system for biological databases.

Database federation enables biological researchers to utilize resources more effectively, creating an environment in which the researcher can query multiple data sources without spending time learning new query mechanisms or issuing redundant queries which need to be integrated. Several mechanisms exist to federate databases. The ENQUire system is a network database federation system which uses a World-Wide-Web (WWW) interface to connect the users to various databases. Generic queries entered via a query generator form are sent in parallel to multiple databases, and the results are presented to the user in a unified format. All forms building, query generation, and results translation is done on the fly, and individual database translation modules can be added dynamically. ENQUire is a flexible answer to the problems of database federation on the WWW.

Computer Communication Networks↗

LIGAND: chemical database for enzyme reactions.

MOTIVATION: The existing molecular biology databases focus on the sequence and structural aspects of biological macromolecules, i.e. DNAs, RNAs and proteins. However, in order to understand the functional aspects, it is essential to computerize the interaction of these molecules. Furthermore, living cells contain additional molecules, such as metabolic compounds and metal ions, that may also be considered as parts of the basic building blocks of life, but are not well organized in public databases. LIGAND chemical database is our attempt to solve these problems, at least for enzymatic reactions. RESULTS: LIGAND consists of two sections: ENZYME and COMPOUND. The ENZYME section is an extension of previous studies (Suyama et al. , Comput. Applic. Biosci., 9, 9-15, 1993), and it is a flat-file representation of 3303 enzymes and 2976 enzymatic reactions in the chemical equation format that can be parsed by machine. The COMPOUND section has been newly constructed for information on the nomenclature and chemical structures of compounds. It contains 5383 chemical compounds. Both ENZYME and COMPOUND entries contain rich cross-reference information, most of which is automatically generated by the DBGET/LinkDB system, thus providing the linkage between chemical and biological databases. LIGAND is updated daily, tightly coupled with the KEGG metabolic pathway database, and forms the basis for reconstruction and computation of pathways. AVAILABILITY: LIGAND can be accessed through the DBGET/LinkDB and KEGG systems in the Japanese GenomeNet database service via http://www.genome.ad.jp/. The flat-file format of the LIGAND database can be downloaded by anonymous FTP via ftp://kegg. genome.adjp/molecules/ligand/. CONTACT: goto@kuicr.kyoto-u.ac.jp; nishioka@scl.kyoto-u.ac.jp; kanehisa@kuicr.kyoto-u.ac.jp

Computational Biology↗

Blocks+: a non-redundant database of protein alignment blocks derived from multiple compilations.

MOTIVATION: As databanks grow, sequence classification and prediction of function by searching protein family databases becomes increasingly valuable. The original Blocks Database, which contains ungapped multiple alignments for families documented in Prosite, can be searched to classify new sequences. However, Prosite is incomplete, and families from other databases are now available to expand coverage of the Blocks Database. RESULTS: To take advantage of protein family information present in several existing compilations, we have used five databases to construct Blocks+, a unified database that is built on the PROTOMAT/BLOSUM scoring model and that can be searched using a single algorithm for consistent sequence classification. The LAMA blocks-versus-blocks searching program identifies overlapping protein families, making possible a non-redundant hierarchical compilation. Blocks+ consists of all blocks derived from PROSITE, blocks from Prints not present in PROSITE, blocks from Pfam-A not present in PROSITE or Prints, and so on for ProDom and Domo, for a total of 1995 protein families represented by 8909 blocks, doubling the coverage of the original Blocks Database. A challenge for any procedure aimed at non-redundancy is to retain related but distinct families while discarding those that are duplicates. We illustrate how using multiple compilations can minimize this potential problem by examining the SNF2 family of ATPases, which is detectably similar to distinct families of helicases and ATPases. AVAILABILITY: http://blocks.fhcrc.org/

Adenosine Triphosphatases↗

MDB: a database system utilizing automatic construction of modules and STAR-derived universal language.

MOTIVATION: The value of information greatly increases if stored in databases. The objective was to construct a multi-purpose database system primarily designed to store and provide access to three-dimensional structures of biological molecules including theoretical models. RESULTS: A dictionary defining data format and structure for three-dimensional models of biological molecules (MDB dictionary) was developed. The dictionary was written using universal, standardized data description language. This language can be applied to describe data with no restrictions on their origin or type, including metadata. Thus both the data definitions (format) and database descriptions are created using the uniform language and processed with universal software. A database and data design technique that allowed use of dictionaries to automatically construct relational databases was developed. This technique was employed to construct the MDB database system. Data design developed and applied in the MDB project makes it possible to carry out data curation utilizing the database engine to identify errors. It also allows storage and query of data at different levels of consistency with the standard format specifications, i.e. both the correctly formatted data, and data that requires further curation. AVAILABILITY: The MDB dictionary is available at http://www.gwer.ch/proteinstructure/mdb and as part of the PDB resources at http://pdb.rutgers.edu/mmcif/.

Computational Biology↗

An improved FORTRAN 77 recombinant DNA database management system with graphic extensions in GKS.

We have improved an existing clone database management system written in FORTRAN 77 and adapted it to our software environment. Improvements are that the database can be interrogated for any type of information, not just keywords. Also, recombinant DNA constructions can be represented in a simplified 'shorthand', whereafter a program assembles the full nucleotide sequence from the contributing fragments, which may be obtained from nucleotide sequence databases. Another improvement is the replacement of the database manager by programs, running in batch to maintain the databank and verify its consistency automatically. Finally, graphic extensions are written in Graphical Kernel System, to draw linear and circular restriction maps of recombinants. Besides restriction sites, recombinant features can be presented from the feature lines of recombinant database entries, or from the feature tables of nucleotide databases. The clone database management system is fully integrated into the sequence analysis software package from the Pasteur Institute, Paris, and is made accessible through the same menu. As a result, recombinant DNA sequences can directly be analysed by the sequence analysis programs.

Algorithms↗

Automated assembly of protein blocks for database searching.

A system is described for finding and assembling the most highly conserved regions of related proteins for database searching. First, an automated version of Smith's algorithm for finding motifs is used for sensitive detection of multiple local alignments. Next, the local alignments are converted to blocks and the best set of non-overlapping blocks is determined. When the automated system was applied successively to all 437 groups of related proteins in the PROSITE catalog, 1764 blocks resulted; these could be used for very sensitive searches of sequence databases. Each block was calibrated by searching the SWISS-PROT database to obtain a measure of the chance distribution of matches, and the calibrated blocks were concatenated into a database that could itself be searched. Examples are provided in which distant relationships are detected either using a set of blocks to search a sequence database or using sequences to search the database of blocks. The practical use of the blocks database is demonstrated by detecting previously unknown relationships between oxidoreductases and by evaluating a proposed relationship between HIV Vif protein and thiol proteases.

Algorithms↗

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural three dimensional (3-D) and sequence one dimensional(1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in Swissprot using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 27% of all Swissprot-stored sequences.

Amino Acid Sequence↗

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural (3-D) and sequence (1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in SwissProt using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 29% of all SwissProt-stored sequences.

Amino Acid Sequence↗

MIPS: a database for protein sequences, homology data and yeast genome information.

The MIPS group (Martinsried Institute for Protein Sequences) at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, collects, processes and distributes protein sequence data within the framework of the tripartite association of the PIR-International Protein Sequence Database (,). MIPS contributes nearly 50% of the data input to the PIR-International Protein Sequence Database. The database is distributed on CD-ROM together with PATCHX, an exhaustive supplement of unique, unverified protein sequences from external sources compiled by MIPS. Through its WWW server (http://www.mips.biochem.mpg.de/ ) MIPS permits internet access to sequence databases, homology data and to yeast genome information. (i) Sequence similarity results from the FASTA program () are stored in the FASTA database for all proteins from PIR-International and PATCHX. The database is dynamically maintained and permits instant access to FASTA results. (ii) Starting with FASTA database queries, proteins have been classified into families and superfamilies (PROT-FAM). (iii) The HPT (hashed position tree) data structure () developed at MIPS is a new approach for rapid sequence and pattern searching. (iv) MIPS provides access to the sequence and annotation of the complete yeast genome (), the functional classification of yeast genes (FunCat) and its graphical display, the 'Genome Browser' (). A CD-ROM based on the JAVA programming language providing dynamic interactive access to the yeast genome and the related protein sequences has been compiled and is available on request.

Academies and Institutes↗

IARC Database of p53 gene mutations in human tumors and cell lines: updated compilation, revised formats and new visualisation tools.

Since 1989, about 570 different p53 mutations have been identified in more than 8000 human cancers. A database of these mutations was initiated by M. Hollstein and C. C. Harris in 1990. This database originally consisted of a list of somatic point mutations in the p 53 gene of human tumors and cell lines, compiled from the published literature and made available in a standard electronic form. The database is maintained at the International Agency for Research on Cancer (IARC) and updated versions are released twice a year (January and July). The current version (July 1997) contains records on 6800 published mutations and will surpass the 8000 mark in the January 1998 release. The database now contains information on somatic and germline mutations in a new format to facilitate data retrieval. In addition, new tools are constructed to improve data analysis, such as a Mutation Viewer Java applet developed at the European Bioinformatics Institute (EBI) to visualise the location and impact of mutations on p53 protein structure. The database is available in different electronic formats at IARC (http://www.iarc. fr/p53/homepage.htm ) or from the EBI server (http://www.ebi.ac.uk ). The IARC p53 website also provides reports on database analysis and links with other p53 sites as well as with related databases. In this report, we describe the criteria for inclusion of data, the revised format and the new visualisation tools. We also briefly discuss the relevance of p 53 mutations to clinical and biological questions.

Computer Communication Networks↗

FlyNets and GIF-DB, two internet databases for molecular interactions in Drosophila melanogaster.

GIF-DB and FlyNets are two WWW databases describing molecular (protein-DNA, protein-RNA and protein-protein) interactions occuring in the fly Drosophila melanogaster (http://gifts.univ-mrs.fr/GIFTS_home_page.html ). GIF-DB is a specialised database which focuses on molecular interactions involved in the process of embryonic pattern formation, whereas FlyNets is a new and more general database, the long-term goal of which is to report on any published molecular interaction occuring in the fly. The information content of both databases is distributed in specific lines arranged into an EMBL- (or GenBank-) like format. These databases achieve a high level of integration with other databases such as FlyBase, EMBL, GenBank and SWISS-PROT through numerous hyperlinks. In addition, we also describe SOS-DGDB, a new collection of annotated Drosophila gene sequences, in which binding sites for regulatory proteins are directly visible on the DNA primary sequence and hyperlinked both to GIF-DB and TRANSFAC database entries.

Animals↗