Search PubMedSearch

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

A scientific relational database combined with a report generator for endoscopy in networks: EndoNet.

BACKGROUND AND STUDY AIMS: The flexibility required in academic endoscopy units is not provided by the available database systems. In a project involving substantial cooperation between endoscopists and computer scientists, we have developed an adaptable database, combined with a report generator embedded in the hospital's intranet. PATIENTS AND METHODS: Six workstations in different areas of the hospital were clustered with a UNIX operating system to implement multi-user capability and access control. A relational database was used to design an application appropriate to the specific needs of the endoscopy unit in a teaching hospital engaged in scientific research. Both the terminology used in standardized endoscopy nomenclature and a free text block facility were included. A graphical user interface was developed to assemble pertinent data, generate the reports, and supervise the database. RESULTS: A total of 4936 examinations including 2988 patients were entered consecutively during continuous routine operation of the system. Complete report generation required five minutes (median; range 1-9 minutes). Both structured items and free text were used in all the reports. Querying of the database was possible, concerning matters such as the need for repeated endoscopic therapy in acute gastrointestinal bleeding (4%), the search for Helicobacter pylori in appropriate patients (64%), the rate of accidental pancreatic duct visualization in endoscopic retrograde cholangiography (24%), and links between examinations and active trials (2%). Indicating improved report quality, the number and the diameter of esophageal varices in patients with varices were more frequently reported with the new report system than with previous typed reports (P<0.001). An anonymous questionnaire revealed that the readability of the computer-generated reports was better than that of the previous typewritten reports (P=0.01). CONCLUSIONS: This report describes the creation of a database application and a report generator meeting the needs of scientific and routine use, and the successful application of this system in an academic endoscopy unit.

Computer Communication Networks

A database for cell signaling networks.

We developed a data and knowledge base for cellular signal transduction in human cells, to make this rapidly growing information available. The database includes all the biological properties of cellular signal transduction, including biological reactions that transfer cellular signals and molecular attributes characterized by sequences, structures, and functions. Since the database is based on the object-oriented technique, highly flexible methods of data definition and modification are necessary to handle this diverse and complex biological information. The database includes attractive graphical representations of signaling cascades and the three-dimensional structure of molecules. The database is a novel application of ACEDB, which was the database originally developed to store the C. elegans genome. The database can be accessed through the Internet at http://geo.nihs.go.jp/csndb.html.

Cells

Post-processing of BLAST results using databases of clustered sequences.

MOTIVATION: When evaluating the results of a sequence similarity search, there are many situations where it can be useful to determine whether sequences appearing in the results share some distinguishing characteristic. Such dependencies between database entries are often not readily identifiable, but can yield important new insights into the biological function of a gene or protein. RESULTS: We have developed a program called CBLAST that sorts the results of a BLAST sequence similarity search according to sequence membership in user-defined 'clusters' of sequences. To demonstrate the utility of this application, we have constructed two cluster databases. The first describes clusters of nucleotide sequences representing the same gene, as documented in the UNIGENE database, and the second describes clusters of protein sequences which are members of the protein families documented in the PROSITE database. Cluster databases and the CBLAST post-processor provide an efficient mechanism for identifying and exploring relationships and dependencies between new sequences and database entries.

Algorithms

A set-theoretic approach to database searching and clustering.

MOTIVATION: In this paper, we introduce an iterative method of database searching and apply it to design a database clustering algorithm applicable to an entire protein database. The clustering procedure relies on the quality of the database searching routine and further improves its results based on a set-theoretic analysis of a highly redundant yet efficient to generate cluster system. RESULTS: Overall, we achieve unambiguous assignment of 80% of SWISS-PROT sequences to non-overlapping sequence clusters in an entirely automatic fashion. Our results are compared to an expert-generated clustering for validation. The database searching method is fast and the clustering technique does not require time-consuming all-against-all comparison. This allows for fast clustering of large amounts of sequences. AVAILABILITY: The resulting clustering for the PIR1 (Release 51) and SWISS-PROT (Release 34) databases is available over the Internet from http://www.dkfz-heidelberg.de/tbi/services/modest/b rowsesysters.pl. CONTACT: a.krause@dkfz-heidelberg.de; m.vingron@dkfz-heidelberg.de

Algorithms

Genome-related datasets within the E. coli Genetic Stock Center database.

The contents of the E. coli Genetic Stock Center database and the availability in electronic form of the subset of information most relevant to sequence databases are described. The database uses the long-standing Stock Center records (developed and curated by Dr B.J.Bachmann) in describing genotypes of mutant derivatives of E.coli K-12 in terms of alleles, structural mutations, mating type, and plasmids as well as the derivation, names and originators of the strain, and references. The database includes descriptions of mutations, mutation properties, genes, gene properties, and gene products, with EC number identifiers for enzymes. Sequence information is not included, but entries refer to sequence database accession numbers for sequenced regions. A gene is described as a subtype of a more general category of chromosome interval called Site. Since sites are used to describe any chromosomal interval, mapping information is associated with sites. Alleles are described as mutations of those sites and they are not primary map objects, but inherit map position information from the corresponding site description. The database design is intended to preserve richness of detail where it is known and uncertainty of measurements or information as it occurs in order to represent the stock center records as accurately as possible.

Bacterial Proteins

Histone and histone fold sequences and structures: a database.

A database of aligned histone protein sequences has been constructed based on the results of homology searches of the major public sequence databases. In addition, sequences of proteins identified as containing the histone fold motif and structures of all known histone and histone fold proteins have been included in the current release. Database resources include information on conflicts between similar sequence entries in different source databases, multiple sequence alignments, and links to the Entrez integrated information retrieval system at the National Center for Biotechnology Information (NCBI). The database currently contains over 1000 protein sequences. All sequences and alignments in this database are available through the World Wide Web at: http: //www.ncbi.nlm.nih.gov/Baxevani/HISTONES/ .

Amino Acid Sequence

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database is a comprehensive database of DNA and RNA sequences directly submitted from researchers and genome sequencing groups and collected from the scientific literature and patent applications. In collaboration with DDBJ and GenBank the database is produced, maintained and distributed at the European Bioinformatics Institute (EBI) and constitutes Europe's primary nucleotide sequence resource. Database releases are produced quarterly and are distributed on CD-ROM. EBI's network services allow access to the most up-to-date data collection via Internet and World Wide Web interface, providing database searching and sequence similarity facilities plus access to a large number of additional databases.

Academies and Institutes

Histone Sequence Database: new histone fold family members.

Searches of the major public protein databases with core and linker chicken and human histone sequences have resulted in the compilation of an annotated set of histone protein sequences. In addition, new database searches with two distinct motif search algorithms have identified several members of the histone fold family, including human DRAP1 and yeast CSE4. Database resources include information on conflicts between similar sequence entries in different source databases, multiple sequence alignments, links to the Entrez integrated information retrieval system, structures for histone and histone fold proteins, and the ability to visualize structural data through Cn3D. The database currently contains >1000 protein sequences, which are searchable by protein type, accession number, organism name, or any other free text appearing in the definition line of the entry. All sequences and alignments in this database are available through the World Wide Web at http://www.nhgri.nih. gov/DIR/GTB/HISTONES or http://www.ncbi.nlm.nih. gov/Baxevani/HISTONES

Amino Acid Sequence

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl.html) constitutes Europe's primary nucleotide sequence resource. Main sources for DNA and RNA sequences are direct submissions from individual researchers, genome sequencing projects and patent applications. While automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO), the preferred submission tool for individual submitters is Webin (WWW). Through all stages, dataflow is monitored by EBI biologists communicating with the sequencing groups. In collaboration with DDBJ and GenBank the database is produced, maintained and distributed at the European Bioinformatics Institute (EBI). Database releases are produced quarterly and are distributed on CD-ROM. Network services allow access to the most up-to-date data collection via Internet and World Wide Web interface. EBI's Sequence Retrieval System (SRS) is a Network Browser for Databanks in Molecular Biology, integrating and linking the main nucleotide and protein databases, plus many specialised databases. For sequence similarity searching a variety of tools (e.g. Blitz, Fasta, Blast etc) are available for external users to compare their own sequences against the most currently available data in the EMBL Nucleotide Sequence Database and SWISS-PROT.

Amino Acid Sequence

The RESID Database of protein structure modifications.

Because the number of post-translational modifications requiring standardized annotation in the PIR-International Protein Sequence Database was large and steadily increasing, a database of protein structure modifications was constructed in 1993 to assist in producing appropriate feature annotations for covalent binding sites, modified sites and cross-links. In 1995 RESID was publicly released as a PIR-International text database distributed on CD-ROM and accessible through the ATLAS program. In 1998 it was made available on the PIR Web site at http://www-nbrf.georgetown.edu/pir/searchdb++ +.html . The RESID Database includes such information as: systematic and frequently observed alternate names; Chemical s Service registry numbers; atomic formulas and weights; enzyme activities; indicators forN-terminal, C-terminal or peptide chain cross-link modifications; keywords; and literature citations with database cross-references. The RESID Database can be used to predict atomic masses for peptides, and is being enhanced to provide molecular structures for graphical presentation on the PIR Web site using widely available molecular viewing programs.

Binding Sites

ProClass Protein Family Database.

ProClass is a protein family database that organizes non-redundant sequence entries into families defined collectively by PROSITE patterns and PIR superfamilies. By combining global similarities and functional motifs into a single classification scheme, ProClass helps to reveal domain and family relationships and classify multi-domain proteins. The database currently consists of more than 120 000 sequence entries, approximately 60% of which is classified into about 3500 families. To maximize family information retrieval, the database provides links to various protein family/domain and structural class databases and contains multiple motif alignments of all PROSITE patterns as well as global alignments of PIR superfamilies. The motif sequences are retrieved from both PIR-International and SWISS-PROT databases, including a large number of new members detected by our GeneFIND family identification system. ProClass can be used to support full-scale genomic annotation, because of its high classification rate. The ProClass database is available for on-line search and record retrieval from our WWW server at http://diana.uthct.edu/proclass.html

Amino Acid Sequence

Grasping at molecular interactions and genetic networks in Drosophila melanogaster using FlyNets, an Internet database.

FlyNets (http://gifts.univ-mrs.fr/FlyNets/FlyNets_home_page.++ +html) is a WWW database describing molecular interactions (protein-DNA, protein-RNA and protein-protein) in the fly Drosophila melanogaster. It is composed of two parts, as follows. (i) FlyNets-base is a specialized database which focuses on molecular interactions involved in Drosophila development. The information content of FlyNets-base is distributed among several specific lines arranged according to a GenBank-like format and grouped into five thematic zones to improve human readability. The FlyNets database achieves a high level of integration with other databases such as FlyBase, EMBL, GenBank and SWISS-PROT through numerous hyperlinks. (ii) FlyNets-list is a very simple and more general databank, the long-term goal of which is to report on any published molecular interaction occuring in the fly, giving direct web access to corresponding s in Medline and in FlyBase. In the context of genome projects, databases describing molecular interactions and genetic networks will provide a link at the functional level between the genome, the proteome and the transcriptome worlds of different organisms. Interaction databases therefore aim at describing the contents, structure, function and behaviour of what we herein define as the interactome world.

Animals

An acute care physical therapy clinical practice database for outcomes research.

Clinical practice databases are frequently used to assess outcomes in various medical specialties. Formulating a computerized physical therapy medical record requires standardization of clinical assessments among the users. The purpose of this article is to describe an acute care physical therapy database system that emphasizes high-quality measures of function. The logic underlying the development of a physical therapy computerized medical record is described. Selected uses of the database are demonstrated by projects that assess data quality, generate clinical hypotheses, manage clinical data, develop clinical measures, and generate pilot data on patient variability. Patients seen in physical therapy for total joint replacement, pain, and decreased ambulation were studied to demonstrate some of the present capabilities of the database. Clinical practice databases contribute to the overall research mission, provided the data are of high quality. The use of databases in conjunction with randomized clinical trials may serve an important role in determining effective physical therapy interventions to reduce disability.

Database Management Systems

Benefits and limitations of database analysis for outcome prediction in cardiac surgery.

Large databases are being used for outcome prediction analysis with increasing frequency. This review examines four separate databases used to provide risk analysis in the cardiac surgery population. Populations in the databases range in size from 3500 to over 116,000 patients. All of the databases were applied on the clinical, and in one instance, institutional level. Outcome prediction from databases is not without its limitations. Data collection, model bias, and methodologic variation all contribute to weaknesses in the application of databases for outcome prediction.

Cardiac Surgical Procedures

Adopting a corporate perspective on databases. Improving support for research and decision making.

The Veterans Health Administration (VHA) is at the forefront of designing and managing health care information systems that accommodate the needs of clinicians, researchers, and administrators at all levels. Rather than using one single-site, centralized corporate database VHA has constructed several large databases with different configurations to meet the needs of users with different perspectives. The largest VHA database is the Decentralized Hospital Computer Program (DHCP), a multisite, distributed data system that uses decoupled hospital databases. The centralization of DHCP policy has promoted data coherence, whereas the decentralization of DHCP management has permitted system development to be done with maximum relevance to the users'local practices. A more recently developed VHA data system, the Event Driven Reporting system (EDR), uses multiple, highly coupled databases to provide workload data at facility, regional, and national levels. The EDR automatically posts a subset of DHCP data to local and national VHA management. The development of the EDR illustrates how adoption of a corporate perspective can offer significant database improvements at reasonable cost and with modest impact on the legacy system.

Database Management Systems

A database of cardiac arrhythmias.

OBJECTIVE: To describe a database of cardiac arrhythmia recordings, useful for the development and testing of ECG rhythm processing or monitoring algorithms and devices. METHODS: The raw data were acquired within the Wisconsin-Dane County emergency medical technician-defibrillation program and contained emergency rhythm recordings of an average length of 30 minutes. The raw data were integrated into a software platform designed for the annotation and visualization of the recordings. RESULTS: Currently the database contains the following arrhythmia episodes: ventricular fibrillation (56), asystole (65), electromechanical dissociation (31), and other arrhythmias (42). The software, resident on personal computers, also can transmit any of the database recordings, through a digital-to-analog converter board, to a device under test. CONCLUSIONS: The database technique described will provide a useful means of objectively assessing electronic devices for their ability to detect arrhythmias. The database is unique in that it contains lengthy episodes of arrhythmias. The database will be extended to include additional cases.

Arrhythmias, Cardiac

Linking large administrative databases: a method for conducting emergency medical services cohort studies using existing data.

OBJECTIVE: To evaluate probabilistic matching for linking a cohort of cardiac arrest (CA) patients identified in the Metro Toronto Ambulance (MTA) database in Toronto, Ontario, Canada, to their appropriate record in either the Vital Statistics Information System (VSIS) or the Canadian Institute of Health Information (CIHI) databases and thus establish their clinical outcomes. METHODS: A linkage of a large administrative database was performed. A cohort of patients who suffered an out-of-hospital CA during the calendar years 1988-1993 was identified. To determine the patients' outcomes, the cohort was probabilistically linked to patient records in the VSIS and CIHI databases. Identifying variables used during the process of linking records included: names (first and last); New York State Identification and Intelligence System (NYSIIS) code; date of event; date of death; city; admitting hospital number; mode of admission to hospital; age; and sex. RESULTS: A cohort of 7,079 CA patients was identified from the MTA database; 6,448 (91%) patients were accurately linked to records in 1 of the 2 outcome databases (CIHI, VSIS). Missing data for > or = 1 of the linking variables were responsible for unlinked records. Using these longitudinal data, it was possible to determine the number of patients surviving their out-of-hospital CAs to be admitted to hospital (n = 833) (16%). No differences in survival rates (p = 0.06) or median lengths of hospital stay among the survivors (p = 0.15) were observed between admitting hospitals. CONCLUSIONS: Probabilistic matching is an effective method by which researchers can use existing administrative data to determine outcomes of population cohorts. This is especially valuable in situations where controlled intervention studies are not feasible or may be inappropriate. In this analysis, in-hospital management of admitted CA patients, as determined by hospital-specific survival rates and length of stay, suggests no measurable differences in the care provided to these patients by hospitals in Toronto.

Algorithms

Multicenter evaluation of the updated and extended API (RAPID) Coryne database 2.0.

In a multicenter study, 407 strains of coryneform bacteria were tested with the updated and extended API (RAPID) Coryne system with database 2.0 (bioMérieux, La-Balme-les-Grottes, France) in order to evaluate the system's capability of identifying these bacteria. The design of the system was exactly the same as for the previous API (RAPID) Coryne strip with database 1.0, i.e., the 20 biochemical reactions covered were identical, but database 2.0 included both more taxa and additional differential tests. Three hundred ninety strains tested belonged to the 49 taxa covered by database 2.0, and 17 strains belonged to taxa not covered. Overall, the system correctly identified 90.5% of the strains belonging to taxa included, with additional tests needed for correct identification for 55.1% of all strains tested. Only 5.6% of all strains were not identified, and 3.8% were misidentified. Identification problems were observed in particular for Corynebacterium coyleae, Propionibacterium acnes, and Aureobacterium spp. The numerical profiles and corresponding identification results for the taxa not covered by the new database 2.0 were also given. In comparison to the results from published previous evaluations of the API (RAPID) Coryne database 1.0, more additional tests had to be performed with version 2.0 in order to completely identify the strains. This was the result of current changes in taxonomy and to provide for organisms described since the appearance of version 1.0. We conclude that the new API (RAPID) Coryne system 2.0 is a useful tool for identifying the diverse group of coryneform bacteria encountered in the routine clinical laboratory.

Actinomycetales