Search PubMedSearch

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Microcomputer-assisted filing system of cardiac catheterization records using a relational database management system.

To efficiently store and retrieve cardiac catheterization records, we have developed a computer-assisted database, which comprises a 16-bit microcomputer with dual floppy disk drives, a 20 MB random-access memory, hard disk drive, and a line printer. All programmings were accomplished using a relational database management system (R:base 5000, Microrim, Inc.). Data inquiry procedures could be performed with direct operational commands of the system as well as with preprogrammed command files, and final results of searches were printed out with a line printer. The major advantages of the present system described in this report include: (1) the relatively easy and rapid creation of the database, (2) ease of modification of the database structures even after the system design is finished, (3) operational commands in combination with conditional operator(s) are flexible and powerful enough to allow the end user to retrieve data based on various kinds of criteria, (4) a high-level programming language provided by the R:base automates a series of database procedures with relative ease, (5) relational capabilities of the database management system can enhance the possibility of reconstruction of a new data file from a single or several preexisting data files, and (6) the system can be realized at reasonable cost.

Cardiac Catheterization

The gene-protein database of Escherichia coli: edition 4.

The gene-protein database of Escherichia coli has as its core an index that links each of the protein spots from a two-dimensional polyacrylamide gel to the gene that encodes the protein. Additional information about each protein and its gene is generated from two-dimensional gel analysis or collated from the literature to form the database. Earlier editions of the database have provided periodic updates of information. The current edition does this, but also introduces a new reference gel image produced by an electrophoresis system recently adopted in this laboratory. The new gel system was chosen because it offers an improved opportunity for other investigations to produce close replicas of the reference gel pattern, thereby allowing easier access to the information of the database and encouraging independent contribution to the database. The new gel format also is larger and hence more compatible with computer assisted image analysis, which has become essential for a project of this magnitude. This edition continues the use of the former reference gel images, but adds a reference image of an equilibrium gel of E. coli strain W3110 produced by the new standardized gel system. At this time, 55% of the protein spots annotated on the previous equilibrium reference gel for this organism have been located on the new reference image, and these identifications are included in the tables of the database.

Bacterial Proteins

Mouse liver protein database: a catalog of proteins detected by two-dimensional gel electrophoresis.

Alterations in the abundance or structure of mouse liver proteins are being studied using two-dimensional gel electrophoresis (2-DE) to build a database of protein changes correlating with exposure to ionizing radiation or toxic chemicals. Thus far, studies have included the analysis of proteins from the offspring of exposed parents or from the exposed individuals themselves. In order to characterize and identify proteins found altered by such exposures, sex- and strain-related differences in protein patterns have been analyzed, and the subcellular locations of a large portion of the mapped proteins have been determined. As part of these studies, data are collected and stored using a variety of computer hardware and software tools that allow the accumulation of information on the origin of samples, gel identification, experiment description, and protein similarities and differences. This accumulation of information constitutes the mouse liver protein database. Relational database software is used to tie the different facets of the database together so that the results of a variety of experiments can be compared and interrelated. The database optimizes the information obtained from 2-DE gel sets by allowing use of the data for many purposes, including monitoring of gel resolution to ensure the collection of high quality data and correlation of protein effects induced by different agents. This first edition of the Argonne National Laboratory mouse liver protein database lays the foundation for future work and communication that should elucidate the significance of observed protein effects as possible markers of exposure to toxic agents.

Animals

HSC-2DPAGE and the two-dimensional gel electrophoresis database of dog heart proteins.

A two-dimensional gel electrophoresis database of dog (Canis familiaris) proteins is presented. The database contains 1212 protein spots which have been characterised in terms of their pI and Mr. This database has been integrated into the HSC-2DPAGE database which is accessible on the Internet via the World Wide Web with the uniform resource location (URL): (http://www.harefield.nthames.nhs.uk/nhli/ protein/index.html). Identifications for 80 of the protein spots have been obtained by visual cross-matching with the human heart protein database in HSC-2DPAGE (42 spots), N-terminal microsequence analysis (25 spots) and peptide mass fingerprinting (20 spots). This database is being used in studies of alterations in protein expression in models of heart failure and heart disease.

Amino Acid Sequence

Mining the human proteome: experience with the human lymphoid protein database.

We have undertaken an effort in the past five years aimed at developing a database of lymphoid proteins detectable by two-dimensional (2-D) polyacrylamide gel electrophoresis. The database contains 2-D patterns and derived information pertaining to: (i) polypeptide constituents of unstimulated and stimulated mature T cells and immature thymocytes; (ii) cultured T cells and cell lines that have been manipulated by transfection with a variety of constructs or by treatment with specific agents; (iii) single cell-derived T and B cell clones; (iv) cells obtained from patients with lymphoproliferative disorders and leukemia; and (v) a variety of other relevant cell populations. The database has experienced a substantial expansion in 2-D patterns it contains, numbering currently 9167 individual 2-D patterns. This number represents a fraction of the 30,682 2-D patterns maintained in our databases. The capacity to design and undertake experiments, produce high-quality 2-D patterns, and to undertake simple or rudimentary analyses of 2-D patterns to meet the basic needs of the experiments for which the 2-D gels were produced has exceeded the capacity to fully and uniformly integrate information generated from any gel image or experiment, across all images and experiments. While only a fraction of the information in the 2-D patterns contained in the lymphoid database has been mined, novel findings derived from querying the database point to the merits of this protein based approach. Additional resources have recently been put into place to mine more effectively data pertaining to protein expression in lymphoid cells.

Cell Cycle

The NIDDK liver transplantation database.

UNLABELLED: The NIDDK Liver Transplantation Database was established to prospectively investigate questions related to the experience of patients evaluated for and undergoing liver transplantation. This article presents the study design, methods, and quality of data collection, along with some of the overall results. METHODS: An initial 4-year planning phase was used to develop data collection instruments and quality control procedures regarding assessment for transplantation, liver donors, and the recipients' pre-, peri- and postoperative course. During the 1990-1995 implementation phase, three clinical centers refined the data collection instruments and enrolled and followed consecutive liver transplant candidates who consented to be included in the protocol. RESULTS: The Database contains more than 49,000 data forms from 1563 candidates, 1002 donors, and 916 transplant recipients followed up to 5 years after transplantation. Overall, 95% of protocol forms were completed. The Database includes uniformly defined histology results of liver biopsies performed per protocol and for complications throughout follow-up. In addition, the Database maintains an inventory of available sera for the Serum Bank. All test results of studies performed on the sera are added to the Database. Of 1563 evaluated patients, 59% were deemed eligible for liver transplantation. Of the others who were too well or had contraindications, 15% became eligible later. Characteristics of patients in this study were generally comparable to those of patients nationally. CONCLUSIONS: The NIDDK Liver Transplantation Database has yielded comprehensive and high quality data and is a rich resource for extensive analysis about many important clinical aspects of liver transplantation.

Adolescent

Practice databases and their uses in clinical research.

A few large clinical information databases have been established within larger medical information systems. Although they are smaller than claims databases, these clinical databases offer several advantages: accurate and timely data, rich clinical detail, and continuous parameters (for example, vital signs and laboratory results). However, the nature of the data vary considerably, which affects the kinds of secondary analyses that can be performed. These databases have been used to investigate clinical epidemiology, risk assessment, post-marketing surveillance of drugs, practice variation, resource use, quality assurance, and decision analysis. In addition, practice databases can be used to identify subjects for prospective studies. Further methodologic developments are necessary to deal with the prevalent problems of missing data and various forms of bias if such databases are to grow and contribute valuable clinical information.

Clinical Medicine

Database challenges and solutions in neuroscientific applications.

In the scientific community, the quality and progress of various endeavors depend in part on the ability of researchers to share and exchange large quantities of heterogeneous data with one another efficiently. This requires controlled sharing and exchange of information among autonomous, distributed, and heterogeneous databases. In this paper, we focus on a neuroscience application, Neuroanatomical Rat Brain Viewer (NeuART Viewer) to demonstrate alternative database concepts that allow neuroscientists to manage and exchange data. Requirements for the NeuART application, in combination with an underlying network-aware database, are described at a conceptual level. Emphasis is placed on functionality from the user's perspective and on requirements that the database must fulfill. The most important functionality required by neuroscientists is the ability to construct brain models using information from different repositories. To accomplish such a task, users need to browse remote and local sources and summaries of data and capture relevant information to be used in building and extending the brain models. Other functionalities are also required, including posing queries related to brain models, augmenting and customizing brain models, and sharing brain models in a collaborative environment. An extensible object-oriented data model is presented to capture the many data types expected in this application. After presenting conceptual level design issues, we describe several known database solutions that support these requirements and discuss requirements that demand further research. Data integration for heterogeneous databases is discussed in terms of reducing or eliminating semantic heterogeneity when translations are made from one system to another. Performance enhancement mechanisms such as materialized views and spatial indexing for three-dimensional objects are explained and evaluated in the context of browsing, incorporating, and sharing. Policies for providing the system with fault tolerance and avoiding possible intellectual property abuses are presented. Finally, two existing systems are evaluated and compared using the identified requirements.

Animals

A new way of building a database of EEG findings.

Whereas computer-based electroencephalography (EEG) is widely applied, the EEG interpretations are usually not stored in a way that favours exploitation of modern computer technology. This paper reports an EEG description system facilitating categorization of EEG data in a computerized database. The system interactively communicates with the digital EEG system and also with the general patient administrative system. The main new quality of this system is the methods for data input and automatic data retrieval from several systems, rather than the establishment of a database of EEG data itself. The EEGs are visually analysed and categorized. Manually marked EEG events are automatically transferred to the database and such events as well as defined electrode positions within these epochs are directly linked to their corresponding descriptions. The database is updated without demand for filling in the events in the database in a second operation. Thereby, the EEG interpreter builds the database while analysing the EEG. This system provides an improved accessibility of EEG data for clinical, normative, educational and scientific use.

Brain

Method to correlate tandem mass spectra of modified peptides to amino acid sequences in the protein database.

A method to correlate uninterpreted tandem mass spectra of modified peptides, produced under low-energy (10-50 eV) collision conditions, with amino acid sequences in a protein database has been developed. The fragmentation patterns observed in the tandem mass spectra of peptides containing covalent modifications is used to directly search and fit linear amino acid sequences in the database. Specific information relevant to sites of modification is not contained in the character-based sequence information of the databases. The search method considers each putative modification site as both modified and unmodified in one pass through the database and simultaneously considers up to three different sites of modification. The search method will identify the correct sequence if the tandem mass spectrum did not represent a modified peptide. This approach is demonstrated with peptides containing modifications such as S-carboxymethylated cysteine, oxidized methionine, phosphoserine, phosphothreonine, or phosphotyrosine. In addition, a scanning approach is used in which neutral loss scans are used to initiate the acquisition of product ion MS/MS spectra of doubly charged phosphorylated peptides during a single chromatographic run for data analysis with the database-searching algorithm. The approach described in this paper provides a convenient method to match the nascent tandem mass spectra of modified peptides to sequences in a protein database and thereby identify previously unknown sites of modification.

Algorithms

Database diversity assessment: new ideas, concepts, and tools.

We present some new ideas for characterizing and comparing large chemical databases. The comparison of the contents of large databases is not trivial since it implies pairwise comparison of hundreds of thousands of compounds. We have developed methods for categorizing compounds into groups or series based on their ring-system content, using precalculated structure-based hashcodes. Two large databases can then be compared by simply comparing their hashcode tables. Furthermore, the number of distinct ring-system combinations can be used as an indicator of database diversity. We also present an independent technique for diversity assessment called the saturation diversity approach. This method is based on picking as many mutually dissimilar compounds as possible from a database or a subset thereof. We show that both methods yield similar results. Since the two methods measure very different properties, this probably says more about the properties of the databases studied than about the methods.

Benzene Derivatives

Use of the UK General Practice Research Database for pharmacoepidemiology.

The last decade has seen a surge in the use of computerized health care data for pharmacoepidemiology. Of all European databases, the General Practice Research Database (GPRD) in the UK, has been the most widely used for pharmacoepidemiological research. Since 1994, this database has belonged to the UK Department of Health, and is maintained by the Office of National Statistics (ONS). Currently, around 1500 general practitioners with a population coverage in excess of 3 million, systematically provide their computerized medical data anonymously to ONS. Validation studies of the GPRD have documented the recording of medical data into general practitioners' computers to be near to complete. The GPRD collects truly population-based data, has a size that makes it possible to follow-up large cohorts of users of specific drugs, and includes both outpatient and inpatient clinical information. The access to original medical records is excellent. Desirable improvements to the GPRD would be additional computerized information on certain variables and linkage to other health care databases. Most published studies to date have been in the area of drug safety. The General Practice Research Database has proved that valuable data can be collected in a general practice setting. The full potential of this rich computerized database has yet to come. This experience should serve to encourage others to develop similar population-based data in other countries.

Databases, Factual

A scientific relational database combined with a report generator for endoscopy in networks: EndoNet.

BACKGROUND AND STUDY AIMS: The flexibility required in academic endoscopy units is not provided by the available database systems. In a project involving substantial cooperation between endoscopists and computer scientists, we have developed an adaptable database, combined with a report generator embedded in the hospital's intranet. PATIENTS AND METHODS: Six workstations in different areas of the hospital were clustered with a UNIX operating system to implement multi-user capability and access control. A relational database was used to design an application appropriate to the specific needs of the endoscopy unit in a teaching hospital engaged in scientific research. Both the terminology used in standardized endoscopy nomenclature and a free text block facility were included. A graphical user interface was developed to assemble pertinent data, generate the reports, and supervise the database. RESULTS: A total of 4936 examinations including 2988 patients were entered consecutively during continuous routine operation of the system. Complete report generation required five minutes (median; range 1-9 minutes). Both structured items and free text were used in all the reports. Querying of the database was possible, concerning matters such as the need for repeated endoscopic therapy in acute gastrointestinal bleeding (4%), the search for Helicobacter pylori in appropriate patients (64%), the rate of accidental pancreatic duct visualization in endoscopic retrograde cholangiography (24%), and links between examinations and active trials (2%). Indicating improved report quality, the number and the diameter of esophageal varices in patients with varices were more frequently reported with the new report system than with previous typed reports (P<0.001). An anonymous questionnaire revealed that the readability of the computer-generated reports was better than that of the previous typewritten reports (P=0.01). CONCLUSIONS: This report describes the creation of a database application and a report generator meeting the needs of scientific and routine use, and the successful application of this system in an academic endoscopy unit.

Computer Communication Networks

A database for cell signaling networks.

We developed a data and knowledge base for cellular signal transduction in human cells, to make this rapidly growing information available. The database includes all the biological properties of cellular signal transduction, including biological reactions that transfer cellular signals and molecular attributes characterized by sequences, structures, and functions. Since the database is based on the object-oriented technique, highly flexible methods of data definition and modification are necessary to handle this diverse and complex biological information. The database includes attractive graphical representations of signaling cascades and the three-dimensional structure of molecules. The database is a novel application of ACEDB, which was the database originally developed to store the C. elegans genome. The database can be accessed through the Internet at http://geo.nihs.go.jp/csndb.html.

Cells

Post-processing of BLAST results using databases of clustered sequences.

MOTIVATION: When evaluating the results of a sequence similarity search, there are many situations where it can be useful to determine whether sequences appearing in the results share some distinguishing characteristic. Such dependencies between database entries are often not readily identifiable, but can yield important new insights into the biological function of a gene or protein. RESULTS: We have developed a program called CBLAST that sorts the results of a BLAST sequence similarity search according to sequence membership in user-defined 'clusters' of sequences. To demonstrate the utility of this application, we have constructed two cluster databases. The first describes clusters of nucleotide sequences representing the same gene, as documented in the UNIGENE database, and the second describes clusters of protein sequences which are members of the protein families documented in the PROSITE database. Cluster databases and the CBLAST post-processor provide an efficient mechanism for identifying and exploring relationships and dependencies between new sequences and database entries.

Algorithms

A set-theoretic approach to database searching and clustering.

MOTIVATION: In this paper, we introduce an iterative method of database searching and apply it to design a database clustering algorithm applicable to an entire protein database. The clustering procedure relies on the quality of the database searching routine and further improves its results based on a set-theoretic analysis of a highly redundant yet efficient to generate cluster system. RESULTS: Overall, we achieve unambiguous assignment of 80% of SWISS-PROT sequences to non-overlapping sequence clusters in an entirely automatic fashion. Our results are compared to an expert-generated clustering for validation. The database searching method is fast and the clustering technique does not require time-consuming all-against-all comparison. This allows for fast clustering of large amounts of sequences. AVAILABILITY: The resulting clustering for the PIR1 (Release 51) and SWISS-PROT (Release 34) databases is available over the Internet from http://www.dkfz-heidelberg.de/tbi/services/modest/b rowsesysters.pl. CONTACT: a.krause@dkfz-heidelberg.de; m.vingron@dkfz-heidelberg.de

Algorithms

Genome-related datasets within the E. coli Genetic Stock Center database.

The contents of the E. coli Genetic Stock Center database and the availability in electronic form of the subset of information most relevant to sequence databases are described. The database uses the long-standing Stock Center records (developed and curated by Dr B.J.Bachmann) in describing genotypes of mutant derivatives of E.coli K-12 in terms of alleles, structural mutations, mating type, and plasmids as well as the derivation, names and originators of the strain, and references. The database includes descriptions of mutations, mutation properties, genes, gene properties, and gene products, with EC number identifiers for enzymes. Sequence information is not included, but entries refer to sequence database accession numbers for sequenced regions. A gene is described as a subtype of a more general category of chromosome interval called Site. Since sites are used to describe any chromosomal interval, mapping information is associated with sites. Alleles are described as mutations of those sites and they are not primary map objects, but inherit map position information from the corresponding site description. The database design is intended to preserve richness of detail where it is known and uncertainty of measurements or information as it occurs in order to represent the stock center records as accurately as possible.

Bacterial Proteins

Histone and histone fold sequences and structures: a database.

A database of aligned histone protein sequences has been constructed based on the results of homology searches of the major public sequence databases. In addition, sequences of proteins identified as containing the histone fold motif and structures of all known histone and histone fold proteins have been included in the current release. Database resources include information on conflicts between similar sequence entries in different source databases, multiple sequence alignments, and links to the Entrez integrated information retrieval system at the National Center for Biotechnology Information (NCBI). The database currently contains over 1000 protein sequences. All sequences and alignments in this database are available through the World Wide Web at: http: //www.ncbi.nlm.nih.gov/Baxevani/HISTONES/ .

Amino Acid Sequence