Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Pivot/Remote: a distributed database for remote data entry in multi-center clinical trials.

1. INTRODUCTION. Data collection is a critical component of multi-center clinical trials. Clinical trials conducted in intensive care units (ICU) are even more difficult because the acute nature of illnesses in ICU settings requires that masses of data be collected in a short time. More than a thousand data points are routinely collected for each study patient. The majority of clinical trials are still "paper-based," even if a remote data entry (RDE) system is utilized. The typical RDE system consists of a computer housed in the CC office and connected by modem to a centralized data coordinating center (DCC). Study data must first be recorded on a paper case report form (CRF), transcribed into the RDE system, and transmitted to the DCC. This approach requires additional monitoring since both the paper CRF and study database must be verified. The paper-based RDE system cannot take full advantage of automatic data checking routines. Much of the effort (and expense) of a clinical trial is ensuring that study data matches the original patient data. 2. METHODS. We have developed an RDE system, Pivot/Remote, that eliminates the need for paper-based CRFs. It creates an innovative, distributed database. The database resides partially at the study clinical centers (CC) and at the DCC. Pivot/Remote is descended from technology introduced with Pivot [1]. Study data is collected at the bedside with laptop computers. A graphical user interface (GUI) allows the display of electronic CRFs that closely mimic the normal paper-based forms. Data entry time is the same as for paper CRFs. Pull-down menus, displaying the possible responses, simplify the process of entering data. Edit checks are performed on most data items. For example, entered dates must conform to some temporal logic imposed by the study. Data must conform to some acceptable range of values. Calculations, such as computing the subject's age or the APACHE II score, are automatically made as the data is entered. Data that is collected serially (BP, HR, etc.) can be displayed graphically in a trend form along with other related variables. An audit trail is created that automatically tracks all changes to the original data, making it possible to reconstruct the CRF to any point in time. On-line help provides information on the study protocol as well as assistance with the use of the system. Electronic security makes it possible to lock certain parts of the CRF once it has been monitored. Completed CRFs are transmitted to the DCC via electronic mail where it is reviewed and merged into the study database. Questions about subject data are transmitted back to the CC via electronic mail. This approach to maintaining the study database is unique in that the study data files are distributed among the CC and DCC. Until a subject's CRF is monitored (verified against the original patient data residing in the hospital record), it logically resides at the CC where it was collected. Copies are transmitted to the DCC and are only read there. Any pre-monitoring changes must be made to the data at the CC. Once the subject's CRF is monitored, it logically moves to the DCC, and any subsequent changes are made at the DCC with copies of the CRF flowing back to the CC. 3. DISCUSSION. Pivot/Remote eliminates the need for paper forms by utilizing portable computers that can be used at the patient bedside. A GUI makes it possible to quickly enter data. Because the user gets instant feedback on possible error conditions, time is saved because the original data is close at hand. The ability to display trended data or variables in the context of other data allows detection of erroneous conditions beyond simple range checks. The logical construction of the database minimizes the problem of managing dual databases (at the CC and DCC) and keeps CC personnel in the loop until all changes are made.

Computer Communication Networks↗

A prototype Internet autopsy database. 1625 consecutive fetal and neonatal autopsy facesheets spanning 20 years.

OBJECTIVE: To demonstrate that cause-of-death statements can be generated by a computer algorithm from an autopsy database composed of diagnostic terms. DATA SOURCES: Over 49 000 autopsy facesheets contributed by over a dozen institutions were collected from a publicly accessible Internet autopsy database. This database is available at the following web site: http:@www.med.jhu.edu/pathology/iad.html STUDY SELECTION: To test the feasibility of creating and using a publicly available autopsy database, and to identify the technical and medicolegal problems that may arise with such a novel resource, a prototype study was designed by selecting autopsy facesheets from fetal and neonatal deaths. An algorithm was developed to determine the cause of death from the listing of anatomic diagnoses. DATA EXTRACTION: One thousand six hundred twenty-five fetal and neonatal autopsy facesheets were selected encompassing fetal and neonatal deaths occurring up to 28 days after birth. DATA SYNTHESIS: The algorithm determined causes of death from autopsy facesheet data in all cases. On review by an experienced pediatric pathologist, these automatically generated cause-of-death statements required no modification or only slight modification in over 90% of cases. CONCLUSIONS: A large multi-institutional autopsy database composed of demographic and diagnostic information has been deposited on the Internet. This information can be freely downloaded and used by any researcher without violating patient confidentiality. As a demonstration of one possible application of the database, fetal and neonatal autopsies generated cause-of-death statements using a computer algorithm. One can anticipate that the wealth of information contained in autopsy facesheets can be assembled into a database that will serve the public interest.

Algorithms↗

A regional perinatal database in southern Sweden--a basis for quality assurance in obstetrics and neonatology.

BACKGROUND: In order to ensure as few avoidable adverse outcomes of pregnancy as possible, it is necessary to continuously evaluate the quality of both obstetric and neonatal care. The eleven southernmost hospitals in Sweden have joined together in a project of developing a regional database, with special emphasis on rapid output of information in order to identify changing trends. METHODS: A regional computerized database has been developed, collecting variables and quality indicators agreed upon by all participants. Specific protocols have been designed for obstetric care, neonatal care and autopsy findings. All participating units transfer information on paper forms or via local computerized information systems. The regional database thus receives data on about 20,000 deliveries annually. RESULTS: Data collection started on September 1, 1994. The first results are due in March 1996, and thereafter on a regular basis every 3 months. Special methods for rapid analysis of raw data have been developed with the help of commercially available data analysis tools. CONCLUSIONS: It is possible to construct an information system with different computer platforms and different database tools at each participating facility, as long as the database systems are locally controllable. That a software is commercially available is no guarantee that data transfer to a central database is possible. Experience from participating sites also indicates that a specialized database is needed for registering obstetric data, as general computerized record-keeping systems are unable to cope with an event concerning more than one subject at a time.

Female↗

Prototype implementation of the integrated genomic database.

We aim to develop an open software system to handle human genome data. The system, called Integrated Genomic Database (IGD), will integrate information from many genomic databases and experimental resources into a comprehensive target-end database (IGD TED). Users will access front-end client systems (IGD FRED) to download data of interest to their computers and merge them with their own local data. FREDs will provide persistent storage of, and instant access to, retrieved data; a friendly graphical interface; tools for querying, browsing, analyzing, and editing local data; interface to external analysis; and tools for communicating with the outside world. The TED will be accessible over the network (online and offline) as a read-only resource for multiple clients. It collects data from major databases for nucleotide and protein sequences and structures, genome maps, experimental reagents, phenotypes, and bibliographic data, and sets of raw data produced at genome centers and laboratories. Beside character-based access via Gopher, WAIS, FTP, and several query language interfaces to the TED, we will develop a specialized front-end client, IGD FRED, with its own database manager, based on the ACEDB program. The FRED will support graphical display methods for sequence feature maps, chromosomal genetic and physical maps, and experimental objects like clone grids, etc. FRED will also provide an interface to important analysis software packages and tools for submitting data to external databases in their own format.

Computer Communication Networks↗

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea↗

Development, analysis, refinement, and utility of an interdisciplinary amyotrophic lateral sclerosis database.

The current status of evaluation and management provided by individual healthcare professionals (HCP) at amyotrophic lateral sclerosis (ALS) centers and clinics needs to be analyzed. This paper describes one ALS center's experiences with the development, analysis, refinement, and utility of an interdisciplinary, HCP-driven ALS database. The purpose and conceptual framework of the database, the general data that needed to be collected, and the types of reports that needed to be generated were determined, and, in collaboration with a computer programmer, data entry and database management systems were developed. Data were collected on 234 patients between September 1996 and August 1998, and were analyzed by a biostatistician. Based on review of the biostatistician's report and discussion of problems encountered with the systems, the database was then refined. Benefits of the database system included: systematization of data collection and reporting, reduction of redundant data collection by individuals, decreased variability of evaluation methods and management decisions from patient to patient, and increased availability of a variety of uniform patient information to assist team members in making care decisions. Ongoing refinement will ensure that this HCP-driven ALS database continues to be informative, practical and effective for decision-making and enhancing delivery of care.

Adult↗

Development of an animal genome database and its search system.

An animal genome database has been developed on a Unix workstation and maintained by a relational database management system. This database has focused on the comparative gene mapping between species to assist the mapping of the genes related to phenotypic traits in livestock. The linkage maps, cytogenetic maps, polymerase chain reaction primers of pig, cattle, mouse and human, and their references have been included in the database, and the correspondence among species have been stipulated in the database. In order to search the database effectively, the World Wide Web server (http://ws4.niai.affrc.go.jp/) and the electronic mail server system (e-mail: jgbase-mail@ niai.affrc.go.jp) have been developed on different Unix workstations. These servers are connected to the Internet.

Animals↗

Database-driven multi locus sequence typing (MLST) of bacterial pathogens.

MOTIVATION: Multi Locus Sequence Typing (MLST) is a newly developed typing method for bacteria based on the sequence determination of internal fragments of seven house-keeping genes. It has proved useful in characterizing and monitoring disease-causing and antibiotic resistant lineages of bacteria. The strength of this approach is that unlike data obtained using most other typing methods, sequence data are unambiguous, can be held on a central database and be queried through a web server. RESULTS: A database-driven software system (mlstdb) has been developed, which is used by public health laboratories and researchers globally to query their nucleotide sequence data against centrally held databases over the internet. The mlstdb system consists of a set of perl scripts for defining the database tables and generating the database management interface and dynamic web pages for querying the databases. AVAILABILITY: http://www.mlst.net.

Bacteria↗

BioMolQuest: integrated database-based retrieval of protein structural and functional information.

MOTIVATION: Information about a particular protein or protein family is usually distributed among multiple databases and often in more than one entry in each database. Retrieval and organization of this information can be a laborious task. This task is complicated even further by the existence of alternative terms for the same concept. RESULTS: The PDB, SWISS-PROT, ENZYME, and CATH databases have been imported into a combined relational database, BIOMOLQUEST: A powerful search engine has been built using this database as a back end. The search engine achieves significant improvements in query performance by automatically utilizing cross-references between the legacy databases. The results of the queries are presented in an organized, hierarchical way.

Abstracting and Indexing↗

DAtA: database of Arabidopsis thaliana annotation.

The Database of Arabidopsis thaliana Annotation (D At A) was created to enable easy access to and analysis of all the Arabidopsis genome project annotation. The database was constructed using the completed A.thaliana genomic sequence data currently in GenBank. An automated annotation process was used to predict coding sequences for GenBank records that do not include annotation. D At A also contains protein motifs and protein similarities derived from searches of the proteins in D At A with motif databases and the non-redundant protein database. The database is routinely updated to include new GenBank submissions for Arabidopsis genomic sequences and new Blast and protein motif search results. A web interface to D At A allows coding sequences to be searched by name, comment, blast similarity or motif field. In addition, browse options present lists of either all the protein names or identified motifs present in the sequenced A.thaliana genome. The database can be accessed at http://baggage. stanford.edu/group/arabprotein/

Arabidopsis↗

PASS2: a semi-automated database of protein alignments organised as structural superfamilies.

PASS2 is a nearly automated version of CAMPASS and contains sequence alignments of proteins grouped at the level of superfamilies. This database has been created to fall in correspondence with SCOP database (1.53 release) and currently consists of 110 multi-member superfamilies and 613 superfamilies corresponding to single members. In multi-member superfamilies, protein chains with no more than 25% sequence identity have been considered for the alignment and hence the database aims to address sequence alignments which represent 26 219 protein domains under the SCOP 1.53 release. Structure-based sequence alignments have been obtained by COMPARER and the initial equivalences are provided automatically from a MALIGN alignment and subsequently augmented using STAMP4.0. The final sequence alignments have been annotated for the structural features using JOY4.0. Several interesting links are provided to other related databases and genome sequence relatives. Availability of reliable sequence alignments of distantly related proteins, despite poor sequence identity and single-member superfamilies, permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. The database can be queried by keywords and also by sequence search, interfaced by PSI-BLAST methods. Structure-annotated sequence alignments and several structural accessory files can be retrieved for all the superfamilies including the user-input sequence. The database can be accessed from http://www.ncbs.res.in/%7Efaculty/mini/campass/pass.html.

Amino Acid Sequence↗

Characteristics of the U.S. EPA's Office of Pesticide Programs' toxicity information databases.

The United States Environmental Protection Agency's Office of Pesticide Programs (OPP) requires that data from toxicity testing be submitted to the OPP to support the registration of pesticide chemicals. Once the toxicity data are submitted, they are entered into various toxicity databases. The studies are listed in an archival database to catalog and allow retrieval of the study for review. Reviews of toxicity studies are then placed into a separate database that can be retrieved to support a regulatory position. Toxicity information for health effects other than cancer and gene mutations from chronic exposure is reviewed through a reference dose (RfD) approach, and these decisions and supporting data are entered into an RfD database. Carcinogenicity data are reviewed by a peer review process, and these decisions are entered into a newly developed database to show the regulatory decision with supporting data. The mutagenicity data are reviewed and acceptable data are entered into the Genetic Activity Profile system to catalog and display the submitted information. These databases contain the information used for hazard evaluations as part of the OPP review of pesticide chemicals.

Animals↗

EXProt--a database for EXPerimentally verified Protein functions.

EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).

Amino Acid Sequence↗

Non-sequence databases for biological activity and physicochemical properties.

A biological activity database and a physicochemical property database are described. They are intended to complement the protein sequence database of PIR-International. The Biological Activity Database and the Physicochemical Property Database contain information regarding the biological activity and the physicochemical properties of proteins, respectively. In addition they also provide information about wild-type molecules with which information concerning variant molecules may be compared. Data on artificial variant molecules are stored in the Artificial Variant Database which is described separately.

Amino Acid Sequence↗

A database model for studies of cocaine-dependent pregnant women and their families.

The database management functions for the Mothers Project are arranged into administrative and analytic task groups, and separate systems are devised for each. The task groups can be distinguished not only by differences in data structure but also by interface requirements. The administrative database system uses a relational database technology, whereas the analytic database system employs more traditional flat-file methods. Although the database management systems are complex, they are based on standard database practices, used in widely available software packages, and run on inexpensive desktop computing equipment.

Cocaine↗

Creation and maintenance of Helix, a Web based database of medical genetics laboratories, to serve the needs of the genetics community.

Helix (healthlinks.washington.edu/helix) is a web accessible database that serves as the main U.S. directory of laboratories offering genetic testing. The database was designed to address the previously unmet need for a centralized, continuously updated source of information about clinical and research genetic testing to keep pace with the rapid rate of gene discovery resulting from the Human Genome Project. The Helix project began in 1992 at the University of Washington and Children's Hospital and Regional Medical Center. It has evolved from a single user stand alone relational database to a fully Web enabled database queried and maintained via the web and linked to other web accessible genomic databases. As of February, 1998 it lists more than 500 diseases and 290 laboratories, with over 5,200 registered users making approximately 250 queries/day (90% via the Internet). We describe the iterative design, implementation, population and assessment of the database over a six year period.

Database Management Systems↗

Human gene mutation database-a biomedical information and research resource.

Although 20 years have elapsed since the first single basepair substitution underlying an inherited disease in humans was characterised at the DNA level, the initiative has only recently been taken to establish central database resources for pathological genetic variants. Disease-associated gene lesions are currently collected and publicised by the Human Gene Mutation Database (HGMD) in Cardiff, locus-specific mutation databases, and to some extent also by the Genome Database (GDB) and Online Mendelian Inheritance in Man (OMIM). To date, HGMD represents the only comprehensive and publicly available database of gene lesions underlying human inherited disease. By July 1999, HGMD contained over 18,000 different mutations from some 900 human genes, the majority being single basepair substitutions. In addition to its potential as an information resource for clinicians and genetic counsellors, HGMD has allowed molecular geneticists to address a variety of biological questions through meta-analysis of the collated data. HGMD also promises to assist research workers in optimising mutation search strategies for a given gene. A questionnaire sent out to, and answered by, the editors of 20 key journals revealed that human genetics journals are increasingly reluctant to publish mutation reports. Electronic data submission and publication facilities are therefore urgently required. The World Wide Web (WWW) provides an excellent medium within which to combine the centralised management of basic mutation data, including rigorous quality control, with the possibility of publishing additional mutation-related information. In response to these needs, HGMD has both instituted a collaboration with Springer-Verlag GmbH, Heidelberg, to potentiate free online submission and electronic publication of human gene mutation data and developed links with the curators of locus-specific mutation databases.

Databases, Factual↗

A dynamic two-dimensional polyacrylamide gel electrophoresis database: the mycobacterial proteome via Internet.

Proteome analysis by two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) and mass spectrometry, in combination with protein chemical methods, is a powerful approach for the analysis of the protein composition of complex biological samples. Data organization is imperative for efficient handling of the vast amount of information generated. Thus we have constructed a 2-D PAGE database to store and compare protein patterns of cell-associated and culture-supernatant proteins of different mycobacterial strains. In accordance with the guidelines for federated 2-DE databases, we developed a program that generates a dynamic 2-D PAGE database for the World-Wide-Web to organise and publish, via the internet, our results from proteome analysis of different Mycobacterium tuberculosis as well as Mycobacterium bovis BCG strains. The uniform resource locator for the database is http://www.mpiib-berlin.mpg.de/2D-PAGE and can be read with a Java compatible browser. The interactive hypertext markup language documents displayed are generated dynamically in each individual session from a rational data file, a 2-D gel image file and a map file describing the protein spots as polygons. The program consists of common gateway interface scripts written in PERL, minimizing the administrative workload of the database. Furthermore, the database facilitates not only interactive use, but also worldwide active participation of other scientific groups with their own data, requiring only minimal computer hardware and knowledge of information technology.

Bacterial Proteins↗