Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Accuracy of endoscopic databases for assessing patient symptoms: comparison with self-reported questionnaires in patients infected with the human immunodeficiency virus.

BACKGROUND: Endoscopic databases are increasingly used for clinical research, but their validity as research instruments has not been assessed. We compared the accuracy of endoscopic indications recorded in an endoscopic database with patient symptom questionnaires. METHODS: All patients infected with the human immunodeficiency virus referred to the outpatient gastroenterology practice were prospectively evaluated using recognized symptom questionnaires. For patients undergoing esophagogastroduodenoscopy, the procedure indications recorded in the endoscopic database and the patient's self-reported symptom scores were compared. RESULTS: Ninety-three patients were evaluated. The symptoms of nausea/vomiting, diarrhea, and anorexia were highly predictive for the presence of these symptoms on the patient questionnaires. The symptoms of dyspepsia/abdominal pain did not predict well the presence of these symptoms on the questionnaire. Patients reported frequent and severe symptoms that were not recorded as indications for the procedures. The overall agreement (kappa statistic) was highly variable, from slight (kappa = 0.07 for anorexia) to moderate (kappa = 0.44 for diarrhea). CONCLUSIONS: Endoscopic indications are variably associated with self-reported symptom scores. These findings raise concerns about using some endoscopic database indications as accurate representations of patients' symptoms. Until performance characteristics of a given database are known, symptom-oriented research should use validated questionnaires whenever possible.

Databases, Factual↗

Construction and use of a computerized DNA fingerprint database for lactic acid bacteria from silage.

Efficient selection of new silage inoculant strains from a collection of over 10,000 isolates of lactic acid bacteria (LAB) requires excellent strain discrimination. Toward that end, we constructed a GelCompar II database of DNA fingerprint patterns of ethidium bromide-stained EcoRI fragments of total LAB DNA separated by conventional agarose gel electrophoresis. We found that the total DNA patterns were strain-specific; 56/60 American Type Culture Collection strains of 33 species of LAB could be distinguished. Enterococcus faecium strains ATCC19434 and ATCC35667 had identical total DNA patterns and RiboPrints. Lactobacillus rhamnosus strains ATCC7469 and ATCC27773 also had identical total DNA patterns, but different RiboPrints. EcoRI RiboPrint patterns could distinguish only about 9/23 Lactobacillus plantarum strains and about 6/10 Lactobacillus buchneri strains, whereas all 33 strains could be distinguished by EcoRI total DNA patterns. Despite gel-to-gel variation, new DNA patterns can be readily grouped with existing patterns using GelCompar II. The database contains large homogenous clusters of L. plantarum, E. faecium, L. buchneri, Lactobacillus brevis and Pediococcus species that can be used for tentative taxonomic assignment. We routinely use the DNA fingerprint database to identify and characterize new strains, eliminate duplicate isolates and for quality control of inoculant product strains. The GelCompar II database has been in continuous use for 7 years and contains more than 3600 patterns representing approximately 700 unique patterns from over 300 gels and is the largest computerized DNA fingerprint database for LAB yet reported.

DNA Fingerprinting↗

Identifying clinical trials in the medical literature with electronic databases: MEDLINE alone is not enough.

The objective of this study was to compare the performance of MEDLINE and EMBASE for the identification of articles regarding controlled clinical trials (CCTs) published in English and related to selected topics: rheumatoid arthritis (RA), osteoporosis (OP), and low back pain (LBP). MEDLINE and EMBASE were searched for literature published in 1988 and 1994. The initial selection of papers was then reviewed to confirm that the articles were about CCTs and to assess the quality of the studies. Selected journals were also hand searched to identify CCTs not retrieved by either database. Overall, 4111 different references were retrieved (2253 for RA, 978 for OP, and 880 for LBP); 3418 (83%) of the papers were in English. EMBASE retrieved 78% more references than MEDLINE (2895 versus 1625). Overall, 1217 (30%) of the papers were retrieved by both databases. Two hundred forty-three papers were about CCTs. Two-thirds of these were retrieved by both databases, and one-third by only one. An additional 16 CCTs not retrieved by either database were identified through hand searching. Taking these into account, EMBASE retrieved 16% more CCTs than MEDLINE (220 versus 188); the EMBASE search identified 85% of the CCTs compared to 73% by MEDLINE. No significant differences were observed in the mean quality scores and sample size of the CCTs missed by MEDLINE compared to those missed by EMBASE. Our findings suggest that the use of MEDLINE alone to identify CCTs is inadequate. The use of two or more databases and hand searching of selected journals are needed to perform a comprehensive search.

Controlled Clinical Trials as Topic↗

The development and use of a computerized database for the evaluation of facial fractures incorporating aspects of the AAOMS Parameters of Care.

PURPOSE: This article discusses the development and use of a computerized database to evaluate facial fracture patients. Examples of epidemiologic and treatment outcome analyses that can be performed are also discussed. MATERIALS AND METHODS: FileMaker Pro 2.1 and 3.0 (Claris Corporation, Santa Clara, CA) for the Macintosh (Apple Computer, Inc, Cupertino, CA) was used for the development of the database. The database contained information on the facial fracture patients treated at The University of Oklahoma Health Sciences Center by the Oral and Maxillofacial Surgery service between January 1, 1994 and December 31, 1996. Eight evaluation forms were used: general information, and mandibular, maxillary, zygomatic, nasal, naso-orbital-ethmoid, orbital, and frontal sinus fractures. Indications for therapy and postoperative complications from the AAOMS Parameters of Care, Section on Trauma Surgery, were also included. RESULTS: This database allowed collection of a vast amount of data on 265 patients. Some of the analyses done on patients with mandibular fractures are described. CONCLUSION: This computerized database provides a quick and systematic method of obtaining and retrieving information on facial fracture patients. Numerous epidemiologic and treatment outcome analyses can be performed. Overall complication rates based on the AAOMS Parameters of Care are higher than previously published rates because of the longer list of complications being evaluated.

Adolescent↗

The electric dipole moment of DNA-binding HU protein calculated by the use of an NMR database.

Electric birefringence measurements indicated the presence of a large permanent dipole moment in HU protein-DNA complex. In order to substantiate this observation, numerical computation of the dipole moment of HU protein homodimer was carried out by using NMR protein databases. The dipole moments of globular proteins have hitherto been calculated with X-ray databases and NMR data have never been used before. The advantages of NMR databases are: (a) NMR data are obtained, unlike X-ray databases, using protein solutions. Accordingly, this method eliminates the bothersome question as to the possible alteration of the protein structure due to the transition from the crystalline state to the solution state. This question is particularly important for proteins such as HU protein which has some degree of internal flexibility; (b) the three-dimensional coordinates of hydrogen atoms in protein molecules can be determined with a sufficient resolution and this enables the N-H as well as C = O bond moments to be calculated. Since the NMR database of HU protein from Bacillus stearothermophilus consists of 25 models, the surface charge as well as the core dipole moments were computed for each of these structures. The results of these calculations show that the net permanent dipole moments of HU protein homodimer is approximately 500-530 D (1 D = 3.33 x 10(-30) Cm) at pH 7.5 and 600-630 D at the isoelectric point (pH 10.5). These permanent dipole moments are unusually large for a small protein of the size of 19.5 kDa. Nevertheless, the result of numerical calculations is compatible with the electro-optical observation, confirming a very large dipole moment in this protein.

Bacterial Proteins↗

National occurrence reporting of Trichinella and trichinellosis using a computerized database.

Historical and contemporary national data on the occurrence and distribution of Trichinella and trichinellosis in people, domestic animals and wildlife in Canada have been incorporated into a computerized database, the Canada Database of Animal Parasites (CDAP). This database was established in 1998, and contains similar information on several other helminth and protozoan parasites which are of importance in animal health, human health, food safety or trade. The CDAP is a unique assemblage of national information on parasite occurrences, not available from any other single source. This paper describes the CDAP, with emphasis on sources of data, database structure and outputs, logistical issues associated with database development and maintenance, and the application of the CDAP to Trichinella surveillance at a national level. It is suggested that the CDAP, or a similar approach, could be applied in other countries for assembling data on Trichinella and trichinellosis.

Animals↗

Composition of the peptide fraction in human blood plasma: database of circulating human peptides.

A database was established from human hemofiltrate (HF) that consisted of a mass database and a sequence database, with the aim of analyzing the composition of the peptide fraction in human blood. To establish a mass database, all 480 fractions of a peptide bank generated from HF were analyzed by MALDI-TOF mass spectrometry. Using this method, over 20000 molecular masses representing native, circulating peptides were detected. Estimation of repeatedly detected masses suggests that approximately 5000 different peptides were recorded. More than 95% of the detected masses are smaller than 15000, indicating that HF predominantly contains peptides. The sequence database contains over 340 entries from 75 different protein and peptide precursors. 55% of the entries are fragments from plasma proteins (fibrinogen A 13%, albumin 10%, beta2-microglobulin 8.5%, cystatin C 7%, and fibrinogen B 6%). Seven percent of the entries represent peptide hormones, growth factors and cytokines. Thirty-three percent belong to protein families such as complement factors, enzymes, enzyme inhibitors and transport proteins. Five percent represent novel peptides of which some show homology to known peptide and protein families. The coexistence of processed peptide fragments, biologically active peptides and peptide precursors suggests that HF reflects the peptide composition of plasma. Interestingly, protein modules such as EGF domains (meprin Aalpha-fragments), somatomedin-B domains (vitronectin fragments), thyroglobulin domains (insulin like growth factor-binding proteins), and Kazal-type inhibitor domains were identified. Alignment of sequenced fragments to their precursor proteins and the analysis of their cleavage sites revealed that there are different processing pathways of plasma proteins in vivo.

Amino Acid Sequence↗

Efficient DNA database laboratory strategy for high through-put STR typing of reference samples.

DNA intelligence databases were installed successfully in various countries during the past few years. It is a general trend that laboratories performing STR analysis for DNA databases have to adjust to increased sample through-put, especially when dealing with a high number of reference samples. In contrast to routine forensic casework analysis, where samples of suspects and unknown samples are interpreted with regard to the specific circumstances of the case and are kept distinctly apart from other cases, DNA databases consist of single, primarily unlinked DNA profiles. Problems areas associated with the high number of anonymous DNA profiles are the risk of logistic errors, such as sample mix-up during the laboratory procedure, and the risk of typing errors during manual transcription of data and/or results. Thus, DNA databases clearly require new laboratory strategies to rise to the challenge. This paper presents an efficient automated laboratory strategy on the platform of a laboratory management information system (LIMS) with the Austrian DNA Intelligence Database as example. Two goals were tackled in particular: first, data safety by avoiding both manual interaction during critical laboratory steps (i.e. when DNA is transferred form one tube into another), and errors due to manual transcription of sample information and results. Secondly, efficient sample processing by automizing the laboratory procedure with the help of robotic instruments, thus, giving the DNA staff more time to analyze data.

Austria↗

Accuracy of on-line databases in determining vital status.

Ascertainment of vital status is important for epidemiological and clinical trial research. Two free databases based on the Social Security Administration Death Master File have become available on the Internet. The accuracy of these databases is unknown. A cohort of 124 patients known to be dead and a cohort of 203 patients not known to be dead were identified. The on-line databases were searched with both of these cohorts following a specific search algorithm. The results for both on-line databases were identical. The optimal algorithm had a sensitivity of 0.82 (95% confidence interval 0.74-0.89) for identification of deaths. The sensitivity excluding deaths that occurred during the first year of life (n = 118) was 0.86 (95% CI 0.79-0.92). The specificity was 1.00. This study found that free and convenient on-line databases based on the Social Security Administration Death Master File can be useful in the accurate ascertainment of vital status.

Algorithms↗

Method for screening peptide fragment ion mass spectra prior to database searching.

A methodology is described for screening fragment ion spectra of peptides prior to database searching for protein identification. A software routine written in the Perl programming language was used to analyze data from previous Sequest database searches and develop a set of statistical descriptors that could be used to identify spectra not likely to yield useful results in a database search. A second Perl program used an evolutionary algorithm to optimize the criteria for each statistical descriptor and generate a formula for determining spectral quality. This formula was used by a third Perl program to screen data sets from four independent liquid chromatography tandem mass spectrometry runs. On the average, use of the screening program reduced the time required for a database search by 1/2 with little loss of useful information from the database search results.

Algorithms↗

Assessment methodologies and statistical issues for computer-aided diagnosis of lung nodules in computed tomography: contemporary research topics relevant to the lung image database consortium.

Cancer of the lung and bronchus is the leading fatal malignancy in the United States. Five-year survival is low, but treatment of early stage disease considerably improves chances of survival. Advances in multidetector-row computed tomography technology provide detection of smaller lung nodules and offer a potentially effective screening tool. The large number of images per exam, however, requires considerable radiologist time for interpretation and is an impediment to clinical throughput. Thus, computer-aided diagnosis (CAD) methods are needed to assist radiologists with their decision making. To promote the development of CAD methods, the National Cancer Institute formed the Lung Image Database Consortium (LIDC). The LIDC is charged with developing the consensus and standards necessary to create an image database of multidetector-row computed tomography lung images as a resource for CAD researchers. To develop such a prospective database, its potential uses must be anticipated. The ultimate applications will influence the information that must be included along with the images, the relevant measures of algorithm performance, and the number of required images. In this article we outline assessment methodologies and statistical issues as they relate to several potential uses of the LIDC database. We review methods for performance assessment and discuss issues of defining "truth" as well as the complications that arise when truth information is not available. We also discuss issues about sizing and populating a database.

Algorithms↗

Database-assisted promoter analysis.

The analysis of regulatory sequences is greatly facilitated by database-assisted bioinformatic approaches. The TRANSFAC database contains information on transcription factors and their origins, functional properties and sequence-specific binding activities. Software tools enable us to screen the database with a given DNA sequence for interacting transcription factors. If a regulatory function is already attributed to this sequence then the database-assisted identification of binding sites for proteins or protein classes and subsequent experimental verification might establish functionally relevant sites within this sequence. The binding transcription factors and interacting factors might already be present in the database.

Binding Sites↗

Postoperative pain management on surgical wards-impact of database documentation of anesthesia organized services.

Postoperative pain management (POPM) should be based on an organization exploiting existing expertise and documenting the outcome of the POPM in each individual patient. The aims of the present study were to evaluate the adequacy of database documentation of POPM of an anesthesia organized, nurse-based, anesthesiologist-supervised acute pain service (APS) on surgical wards and to assess to what extent the information obtained was continuously used to improve practice. From 2890 registered cases in the database (patient controlled analgesia, n = 1975; epidural analgesia [EDA], n = 915), a homogeneous two-year sample of documentation charts from use of EDA for POPM in connection with major, open, abdominal surgical procedures (n = 381) was chosen for detailed analysis. The data charts contained information on patient data, drug dosage, total amount of infused drug, duration of EDA treatment, occurrence of side effects, and patient's level of satisfaction. The database information was easily accessible making assessment of relevant aspects of the routines, including associations between analgesic technique, patient related factors, and satisfaction with the services, immediately available. Only 58% of the data charts were properly completed and fed into the database but the clinical safety of the missing nondatabase documented sample was not found jeopardized. Although the database documentation routines were considered to fulfill basic requirements of data collection and monitoring of the appropriateness of POPM, they were not found to function optimally. The reason seemed to be inadequate feedback of information between the parties involved in the POPM services. The present study stresses the importance of establishing routines for adequate, continuous feedback of recorded audit data from the APS team to the surgical wards for the maintenance of a high level of compliance with accepted guidelines.

Adolescent↗

Object oriented database and electronic notebook for transmission electron microscopy.

As high-resolution biological transmission electron microscopy (TEM) has increased in popularity over recent years, the volume of data and number of projects underway has risen dramatically. A robust tool for effective data management is essential to efficiently process large data sets and extract maximum information from the available data. We present the Electron Microscopy Electronic Notebook (EMEN), a portable, object-oriented, web-based tool for TEM data archival and project management. EMEN has several unique features. First, the database is logically organized and annotated so multiple collaborators at different geographical locations can easily access and interpret the data without assistance. Second, the database was designed to provide flexibility to the user, so it can be used much as a lab notebook would be, while maintaining a structure suitable for data mining and direct interaction with data-processing software. Finally, as an object-oriented database, the database structure is dynamic and can be easily extended to incorporate information not defined in the original database specification.

Databases, Factual↗

Role of accurate mass measurement (+/- 10 ppm) in protein identification strategies employing MS or MS/MS and database searching.

We describe the impact of advances in mass measurement accuracy, +/- 10 ppm (internally calibrated), on protein identification experiments. This capability was brought about by delayed extraction techniques used in conjunction with matrix-assisted laser desorption ionization (MALDI) on a reflectron time-of-flight (TOF) mass spectrometer. This work explores the advantage of using accurate mass measurement (and thus constraint on the possible elemental composition of components in a protein digest) in strategies for searching protein, gene, and EST databases that employ (a) mass values alone, (b) fragment-ion tagging derived from MS/MS spectra, and (c) de novo interpretation of MS/MS spectra. Significant improvement in the discriminating power of database searches has been found using only molecular weight values (i.e., measured mass) of > 10 peptide masses. When MALDI-TOF instruments are able to achieve the +/- 0.5-5 ppm mass accuracy necessary to distinguish peptide elemental compositions, it is possible to match homologous proteins having > 70% sequence identity to the protein being analyzed. The combination of a +/- 10 ppm measured parent mass of a single tryptic peptide and the near-complete amino acid (AA) composition information from immonium ions generated by MS/MS is capable of tagging a peptide in a database because only a few sequence permutations > 11 AA's in length for an AA composition can ever be found in a proteome. De novo interpretation of peptide MS/MS spectra may be accomplished by altering our MS-Tag program to replace an entire database with calculation of only the sequence permutations possible from the accurate parent mass and immonium ion limited AA compositions. A hybrid strategy is employed using de novo MS/MS interpretation followed by text-based sequence similarity searching of a database.

Animals↗

Consideration of molecular weight during compound selection in virtual target-based database screening.

Virtual database screening allows for millions of chemical compounds to be computationally selected based on structural complimentary to known inhibitors or to a target binding site on a biological macromolecule. Compound selection in virtual database screening when targeting a biological macromolecule is typically based on the interaction energy between the chemical compound and the target macromolecule. In the present study it is shown that this approach is biased toward the selection of high molecular weight compounds due to the contribution of the compound size to the energy score. To account for molecular weight during energy based screening, we propose normalization strategies based on the total number of heavy atoms in the chemical compounds being screened. This approach is computationally efficient and produces molecular weight distributions of selected compounds that can be selected to be (1) lower than that of the original database used in the virtual screening, which may be desirable for selection of leadlike compounds or (2) similar to that of the original database, which may be desirable for the selection of drug-like compounds. By eliminating the bias in target-based database screening toward higher molecular weight compounds it is anticipated that the proposed procedure will enhance the success rate of computer-aided drug design.

Computer-Aided Design↗

Advanced exact structure searching in large databases of chemical compounds.

Efficient recognition of tautomeric compound forms in large corporate or commercially available compound databases is a difficult and labor intensive task. Our data indicate that up to 0.5% of commercially available compound collections for bioscreening contain tautomers. Though in the large registry databases, such as Beilstein and CAS, the tautomers are found in an automated fashion using high-performance computational technologies, their real-time recognition in the nonregistry corporate databases, as a rule, remains problematic. We have developed an effective algorithm for tautomer searching based on the proprietary chemoinformatics platform. This algorithm reduces the compound to a canonical structure. This feature enables rapid, automated computer searching of most of the known tautomeric transformations that occur in databases of organic compounds. Another useful extension of this methodology is related to the ability to effectively search for different forms of compounds that contain ionic and semipolar bonds. The computations are performed in the Windows environment on a standard personal computer, a very useful feature. The practical application of the proposed methodology is illustrated by several examples of successful recovery of tautomers and different forms of ionic compounds from real commercially available nonregistry databases.

Algorithms↗

Considerations in compound database preparation--"hidden" impact on virtual screening results.

Structure-based virtual screening (SBVS) utilizing docking algorithms has become an essential tool in the drug discovery process, and significant progress has been made in successfully applying the technique to a wide range of receptor targets. In silico validation of virtual screening protocols before application to a receptor target using a corporate or commercially available compound collection is key to establishing a successful process. Ultimately, retrieval of a set of active compounds from a database of inactives is required, and the metric of enrichment (E) is habitually used to discern the quality of separation of the two. Numerous reports have addressed the performance of docking algorithms with regard to the quality of binding mode prediction and the issue of postprocessing "hit lists" of docked ligands. However, the impact of ligand database preprocessing has yet to be examined in the context of virtual screening and prioritization of compounds for biological evaluation. We provide an insight into the implications of cheminformatic preprocessing of a validation database of compounds where multiple protonated, tautomeric, stereochemical, and conformational states have been enumerated. Several commonly used methods for the generation of ligand conformations and conformational ensembles are examined, paired with an exhaustive rigid-body algorithm for the docking of different "multimeric" compound representations to the ligand binding site of the human estrogen receptor alpha. Chemgauss, a shapegaussian scoring function with intrinsic chemical knowledge, was combined with PLP as a consensus-scoring scheme to rank output from the docking protocol and enrichment rates calculated for each screen. The overheads of CPU consumption and the effect on relative database size (disk requirement) for each of the protocols employed are considered. Assessment of these parameters indicates that SBVS enrichments are highly dependent on the initial cheminformatic treatment(s) used in database construction. The interplay of SMILES representations, stereochemical information, protonation state enumeration, and ligand conformation ensembles are critical in achieving optimum enrichment rates in such screening.

Computer Simulation↗