Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Advanced database methodology for the Collation of Connectivity data on the Macaque brain (CoCoMac).

The need to integrate massively increasing amounts of data on the mammalian brain has driven several ambitious neuroscientific database projects that were started during the last decade. Databasing the brain's anatomical connectivity as delivered by tracing studies is of particular importance as these data characterize fundamental structural constraints of the complex and poorly understood functional interactions between the components of real neural systems. Previous connectivity databases have been crucial for analysing anatomical brain circuitry in various species and have opened exciting new ways to interpret functional data, both from electrophysiological and from functional imaging studies. The eventual impact and success of connectivity databases, however, will require the resolution of several methodological problems that currently limit their use. These problems comprise four main points: (i) objective representation of coordinate-free, parcellation-based data, (ii) assessment of the reliability and precision of individual data, especially in the presence of contradictory reports, (iii) data mining and integration of large sets of partially redundant and contradictory data, and (iv) automatic and reproducible transformation of data between incongruent brain maps. Here, we present the specific implementation of the 'collation of connectivity data on the macaque brain' (CoCoMac) database (http://www.cocomac.org). The design of this database addresses the methodological challenges listed above, and focuses on experimental and computational neuroscientists' needs to flexibly analyse and process the large amount of published experimental data from tracing studies. In this article, we explain step-by-step the conceptual rationale and methodology of CoCoMac and demonstrate its practical use by an analysis of connectivity in the prefrontal cortex.

Animals↗

The physiology constant database of teen-agers in Beijing.

Physiology constants of adolescents are important to understand growing living systems and are a useful reference in clinical and epidemiological research. Until recently, physiology constants were not available in China and therefore most physiologists, physicians, and nutritionists had to use data from abroad for reference. However, the very difference between the Eastern and Western races casts doubt on the usefulness of overseas data. We have therefore created a database system to provide a repository for the storage of physiology constants of teen-agers in Beijing. The several thousands of pieces of data are now divided into hematological biochemistry, lung function, and cardiac function with all data manually checked before being transferred into the database. The database was accomplished through the development of a web interface, scripts, and a relational database. The physiology data were integrated into the relational database system to provide flexible facilities by using combinations of various terms and parameters. A web browser interface was designed for the users to facilitate their searching. The database is available on the web. The statistical table, scatter diagram, and histogram of the data are available for both anonym and user according to queries, while only the user can achieve detail, including download data and advanced search.

Adolescent↗

LySDB - Lysozyme Structural DataBase.

LySDB (Lysozyme Structural DataBase) is an integrated database containing 740 three-dimensional structures of lysozyme available in the Protein Data Bank. The database can be used to visualize the three-dimensional structure of the entire protein model or the substructures in which the user is interested (for example, insertions and deletions of amino acids) using the three-dimensional atomic coordinates. The database is provided with a search engine with several useful built-in facilities. The public domain graphics program RASMOL has been deployed for visualization. The three-dimensional structures used to create the database are updated at regular intervals and hence the users are provided with the current information available in the literature. The database LySDB is available over the World Wide Web and can be accessed at the URL http://iris.physics.iisc.ernet.in/lysdb/ or http://144.16.71.2/lysdb/.

Computational Biology↗

Toward a global bat-signal database.

We propose a scheme for a new database using standardized protocol for recording and analysis of bat calls. The proposed database will describe and archive echolocation signals to create a reference library of bat calls. This information should be accessible to the public, thus encouraging continuous feedback from a broad audience. Because it is essential to evaluate the quality and reliability of such data, detailed information of recording and analysis procedures as well as the resulting species identification is required. A standardized and growing database on bat calls would be a potentially invaluable tool for global species identification, comparison, and distribution of microchiropterans. Currently, apart from a few websites with local call libraries, there is no "global" database established yet. We hope that researchers, amateurs, and wildlife and management authorities will adopt and further modify our suggestions for a standardized database for bat calls. We also hope that this database will stimulate new directions in bat research.

Animals↗

Creation and application of a simulated database of dynamic [18f]MPPF PET acquisitions incorporating inter-individual anatomical and biological variability.

During the process of validation of a new tracer, estimation of performance and validation of processing algorithms have to be investigated with data sets representative of the ground truth. Because this ground truth is hardly accessible in positron emission tomography (PET), validations of processing algorithms often rely on the use of simulated data sets. Considering that Monte Carlo simulators are very time consuming and are not very easy to use, the building of publicly available databases of simulated PET volumes are becoming highly desirable. We present here the methodology employed for the creation of a database of simulated dynamic [18F]MPPF-PET data, including inter-individual anatomical and biological variability which meets the criteria of a gold standard database as defined by Lehmann: reliance, equivalence, independence, relevance, significance. The assessment of the realism of the built database against actual MPPF PET data is also presented here. Whereas the database was specifically created for the investigations of quantification of activity and binding of ligand-receptor with the [18F]MPPF PET tracer, it may serve the community with countless purposes. The full strength of this database, does not only stem from the knowledge of important information such as the true activity map and underlying anatomical data, but also from the possibility to fully control the biological difference between sets of simulated PET data. Indeed, time activity curves included in the simulated data sets are controlled by a multicompartmental model of ligand-receptor exchanges. This latter feature is of a great interest in the context of the improvement of the detectability of biological variation in PET.

Brain↗

Lessons from developing and running a clinical database for colorectal cancer.

BACKGROUND: Recent policy developments in the UK require the routine monitoring of the performance of cancer services. Developing and using clinical databases is one approach to meet this objective, but to date their implementation has been challenging. OBJECTIVE: To describe the development of the Thames Cancer Registry clinical database for colorectal cancer, and to present the lessons learnt in the first five years since its establishment. METHODS: Planning of this clinical database began in 1998. Detailed variables for the data set were derived by analysis of national standards and guidelines. Structured pro formas were designed to abstract data from clinical notes. A pilot study over 12 months collected 400 cases from seven hospital trusts in one cancer network. Data collection over the wider North Thames area began in 1999. RESULTS: The number of new records entered each year into the database rose from 747 in 1999 to 1107 in 2002. By 2004, it held a total of 8500. However, participation and completeness of data collection varied between trusts. Currently only 18 of 26 trusts in the area submit data and only 12 have done so every year. Overall completeness for key demographic and treatment variables has been between 80 and 100% but less so for more detailed diagnostic and treatment variables (40-60%). Barriers to implementation in trusts could be grouped as organizational, professional and data-related. Organizational barriers have included changes in the cancer networks, variability in trust commitment to different data sets and lack of personnel to enter data consistently. Professional barriers have included competing priorities and varying commitments within the multidisciplinary clinical teams. Data-related barriers include the wide range of database formats that are used in trusts, and a tendency for data to be collected at the end of the year rather than continuously. CONCLUSIONS: Creating and maintaining a clinical database is a time-consuming and complex undertaking. Completeness of ascertainment and quality are major issues of concern. Key lessons from this project have been that the commitment of clinicians and the ability of trusts to provide consistent support for data collection are crucial.

Colorectal Neoplasms↗

The reagent database at dbMHC.

The reagent database dbMHC was built by the National Center for Biotechnology Information (NCBI) as an open resource for registration and characterization of HLA DNA-typing kits and reagents. Each reagent is uniquely identified as sequence-specific oligonucleotide (SSO) or primer (SSP), SSO mix, or SSP mix. Computerized prediction of allele reactivities, based on annealing stringency, is performed on all submissions to the reagent database. User-specified allele reactivities may be added or deleted independently of the prediction algorithm. Updates of allele reactivities are performed in synchronization with the IMGT/HLA database, in order to account for newly discovered alleles. Probe and primer sequences aligned to allelic sequences can be displayed at any time. Reagents registered in the reagent database are grouped in typing kits. Each kit or kit batch is uniquely identified. Group-specific amplification of alleles can be specified for an entire kit or for sections of each kit. Kits designed to test multiple loci are supported. Kits can be entered and updated via the web or submitted as batches in extensible markup language (XML) format. A tool for online interpretation of typing results is available. Both the reagent database and the typing kit database have been designed to facilitate the exchange of HLA typing based on raw typing data using the unique identifiers of kits or individual reagents. In addition, batch-wise reinterpretation of previous typing data can be performed either using the NCBI web site or by locally using downloaded allele-reactivity lists. Reinterpretation by the NCBI requires submission of raw typing data in XML format.

Computational Biology↗

Euroethics--a database network on biomedical ethics.

BACKGROUND: EUROETHICS is a database covering European literature on ethics in medicine. It is produced within Eurethnet, a European information network on ethics in medicine and biotechnology. OBJECTIVES: The aim of Euroethics is to disseminate information on European bioethical literature that may otherwise be difficult to find. METHODS: A collaboration model for pooling data from different centres was developed. The policy was to accomplish data uniformity, while still allowing for local differences in terms of software, indexing practices and resources. Records contributed to the database follow common standards in terms of data fields and indexing terms. The indexing terms derive from two thesauri, Thesaurus Ethics in the Life Sciences (TELS) and Medical Subject Headings (MeSH). Combining elements from search tools developed previously, the developers sought to find a technical solution optimized for this data model. An approach relying on a thesaurus database that is loaded along with the bibliographic database is described. RESULTS AND CONCLUSIONS: The present case study offers examples of possible approaches to several tasks often encountered in database development, such as: merging data from diverse sources, getting the most out of indexing terms used in a database, and handling more than one thesaurus in the same system.

Abstracting and Indexing↗

A database of cardiac arrhythmias.

OBJECTIVE: To describe a database of cardiac arrhythmia recordings, useful for the development and testing of ECG rhythm processing or monitoring algorithms and devices. METHODS: The raw data were acquired within the Wisconsin-Dane County emergency medical technician-defibrillation program and contained emergency rhythm recordings of an average length of 30 minutes. The raw data were integrated into a software platform designed for the annotation and visualization of the recordings. RESULTS: Currently the database contains the following arrhythmia episodes: ventricular fibrillation (56), asystole (65), electromechanical dissociation (31), and other arrhythmias (42). The software, resident on personal computers, also can transmit any of the database recordings, through a digital-to-analog converter board, to a device under test. CONCLUSIONS: The database technique described will provide a useful means of objectively assessing electronic devices for their ability to detect arrhythmias. The database is unique in that it contains lengthy episodes of arrhythmias. The database will be extended to include additional cases.

Arrhythmias, Cardiac↗

Linking large administrative databases: a method for conducting emergency medical services cohort studies using existing data.

OBJECTIVE: To evaluate probabilistic matching for linking a cohort of cardiac arrest (CA) patients identified in the Metro Toronto Ambulance (MTA) database in Toronto, Ontario, Canada, to their appropriate record in either the Vital Statistics Information System (VSIS) or the Canadian Institute of Health Information (CIHI) databases and thus establish their clinical outcomes. METHODS: A linkage of a large administrative database was performed. A cohort of patients who suffered an out-of-hospital CA during the calendar years 1988-1993 was identified. To determine the patients' outcomes, the cohort was probabilistically linked to patient records in the VSIS and CIHI databases. Identifying variables used during the process of linking records included: names (first and last); New York State Identification and Intelligence System (NYSIIS) code; date of event; date of death; city; admitting hospital number; mode of admission to hospital; age; and sex. RESULTS: A cohort of 7,079 CA patients was identified from the MTA database; 6,448 (91%) patients were accurately linked to records in 1 of the 2 outcome databases (CIHI, VSIS). Missing data for > or = 1 of the linking variables were responsible for unlinked records. Using these longitudinal data, it was possible to determine the number of patients surviving their out-of-hospital CAs to be admitted to hospital (n = 833) (16%). No differences in survival rates (p = 0.06) or median lengths of hospital stay among the survivors (p = 0.15) were observed between admitting hospitals. CONCLUSIONS: Probabilistic matching is an effective method by which researchers can use existing administrative data to determine outcomes of population cohorts. This is especially valuable in situations where controlled intervention studies are not feasible or may be inappropriate. In this analysis, in-hospital management of admitted CA patients, as determined by hospital-specific survival rates and length of stay, suggests no measurable differences in the care provided to these patients by hospitals in Toronto.

Algorithms↗

Multicenter evaluation of the updated and extended API (RAPID) Coryne database 2.0.

In a multicenter study, 407 strains of coryneform bacteria were tested with the updated and extended API (RAPID) Coryne system with database 2.0 (bioMérieux, La-Balme-les-Grottes, France) in order to evaluate the system's capability of identifying these bacteria. The design of the system was exactly the same as for the previous API (RAPID) Coryne strip with database 1.0, i.e., the 20 biochemical reactions covered were identical, but database 2.0 included both more taxa and additional differential tests. Three hundred ninety strains tested belonged to the 49 taxa covered by database 2.0, and 17 strains belonged to taxa not covered. Overall, the system correctly identified 90.5% of the strains belonging to taxa included, with additional tests needed for correct identification for 55.1% of all strains tested. Only 5.6% of all strains were not identified, and 3.8% were misidentified. Identification problems were observed in particular for Corynebacterium coyleae, Propionibacterium acnes, and Aureobacterium spp. The numerical profiles and corresponding identification results for the taxa not covered by the new database 2.0 were also given. In comparison to the results from published previous evaluations of the API (RAPID) Coryne database 1.0, more additional tests had to be performed with version 2.0 in order to completely identify the strains. This was the result of current changes in taxonomy and to provide for organisms described since the appearance of version 1.0. We conclude that the new API (RAPID) Coryne system 2.0 is a useful tool for identifying the diverse group of coryneform bacteria encountered in the routine clinical laboratory.

Actinomycetales↗

Necessity of quality-controlled 16S rRNA gene sequence databases: identifying nontuberculous Mycobacterium species.

The use of the 16S rRNA gene for identification of nontuberculous mycobacteria (NTM) provides a faster and better ability to accurately identify them in addition to contributing significantly in the discovery of new species. Despite their associated problems, many rely on the use of public sequence databases for sequence comparisons. To best evaluate the taxonomic status of NTM species submitted to our reference laboratory, we have created a 16S rRNA sequence database by sequencing 121 American Type Culture Collection strains encompassing 92 species of mycobacteria, and have also included chosen unique mycobacterial sequences from public sequence repositories. In addition, the Ribosomal Differentiation of Medical Microorganisms (RIDOM) service has made freely available on the Internet mycobacterial identification by 16S rRNA analysis. We have evaluated 122 clinical NTM species using our database, comparing >1,400 bp of the 16S gene, and the RIDOM database, comparing approximately 440 bp. The breakdown of analysis was as follows: 61 strains had a sequence with 100% similarity to the type strain of an established species, 19 strains showed a 1- to 5-bp divergence from an established species, 11 strains had sequences corresponding to uncharacterized strain sequences in public databases, and 31 strains represented unique sequences. Our experience with analysis of the 16S rRNA gene of patient strains has shown that clear-cut results are not the rule. As many clinical, research, and environmental laboratories currently employ 16S-based identification of bacteria, including mycobacteria, a freely available quality-controlled database such as that provided by RIDOM is essential to accurately identify species or detect true sequence variations leading to the discovery of new species.

Databases, Nucleic Acid↗

How to evaluate and improve the quality and credibility of an outcomes database: validation and feedback study on the UK Cardiac Surgery Experience.

OBJECTIVES: To assess the quality and completeness of a database of clinical outcomes after cardiac surgery and to determine whether a process of validation, monitoring, and feedback could improve the quality of the database. DESIGN: Stratified sampling of retrospective data followed by prospective re-sampling of database after intervention of monitoring, validation, and feedback. SETTING: Ten tertiary care cardiac surgery centres in the United Kingdom. INTERVENTION: Validation of data derived from a stratified sample of case notes (recording of deaths cross checked with mortuary records), monitoring of completeness and accuracy of data entry, feedback to local data managers and lead surgeons. MAIN OUTCOME MEASURES: Average percentage missing data, average kappa coefficient, and reliability score by centre for 17 variables required for assignment of risk scores. Actual minus risk adjusted mortality in each centre. RESULTS: The database was incomplete, with a mean (SE) of 24.96% (0.09%) of essential data elements missing, whereas only 1.18% (0.06%) were missing in the patient records (P<0.0001). Intervention was associated with (a) significantly less missing data (9.33% (0.08%) P<0.0001); (b) marginal improvement in reliability of data and mean (SE) overall centre reliability score (0.53 (0.15) v 0.44 (0.17)); and (c) improved accuracy of assigned Parsonnet risk scores (kappa 0.84 v 0.70). Mortality scores (actual minus risk adjusted mortality) for all participating centres fell within two standard deviations of the mean score. CONCLUSION: A short period of independent validation, monitoring, and feedback improved the quality of an outcomes database and improved the process of risk adjustment, but with substantial room for further improvement. Wider application of this approach should increase the credibility of similar databases before their public release.

Cardiac Surgical Procedures↗

Information retrieved from a database and the augmentation of personal knowledge.

OBJECTIVE: To assess the degree to which information retrieved from a biomedical database can augment personal knowledge in addressing novel problems, and how the ability to retrieve information evolves over time. DESIGN: This longitudinal study comprised three assessments of two cohorts of medical students. The first assessment occurred just before student course experience in bacteriology, the second occurred just after the course, and the third occurred five months later. At each assessment, the students were initially given a set of bacteriology problems to solve using their personal knowledge only. Each student was then reassigned a sample of problems he or she had answered incorrectly, to work again with assistance from a database containing information about bacteria and bacteriologic concepts. The initial pass through the problems generated a "personal knowledge" score; the second pass generated a "database-assisted" score for each student at each assessment. RESULTS: Over two cohorts, students' personal knowledge scores were very low (approximately 12%) at the first assessment. They rose substantially at the second assessment (approximately 48%) but decreased six months later (approximately 25%). By contrast, database-assisted scores rose linearly: from approximately 44% at the first assessment to approximately 57% at the second assessment, to approximately 75% at the third assessment. CONCLUSION: The persistent increase in database-assisted scores, even when personal knowledge had attenuated, was the most remarkable finding of this study. While some of the increase may be attributed to artifacts of the design, the pattern seems to result from the retained ability to recognize problem-relevant information in a database even when it cannot be recalled.

Bacteriology↗

Development of a replicated database of DHCP data for evaluation of drug use.

This case report describes development and testing of a method to extract clinical information stored in the Veterans Affairs (VA) Decentralized Hospital Computer System (DHCP) for the purpose of analyzing data about groups of patients. The authors used a microcomputer-based, structured query language (SQL)-compatible, relational database system to replicate a subset of the Nashville VA Hospital's DHCP patient database. This replicated database contained the complete current Nashville DHCP prescription, provider, patient, and drug data sets, and a subset of the laboratory data. A pilot project employed this replicated database to answer questions that might arise in drug-use evaluation, such as identification of cases of polypharmacy, suboptimal drug regimens, and inadequate laboratory monitoring of drug therapy. These database queries included as candidates for review all prescriptions for all outpatients. The queries demonstrated that specific drug-use events could be identified for any time interval represented in the replicated database.

Databases, Factual↗

UMLS-based conceptual queries to biomedical information databases: an overview of the project ARIANE. Unified Medical Language System.

OBJECTIVE: The aim of the project ARIANE is to model and implement seamless, natural, and easy-to-use interfaces with various kinds of heterogeneous biomedical information databases. DESIGN: A conceptual model of some of the Unified Medical Language System (UMLS) knowledge sources has been developed to help end users to query information databases. A query is represented by a conceptual graph that translates the deep structure of an end-user's interest in a topic. A computational model exploits this conceptual model to build a query interactively represented as query graph. A query graph is then matched to the data graph built with data issued from each record of a database by means of a pattern-matching (projection) rule that applies to conceptual graphs. RESULTS: Prototypes have been implemented to test the feasibility of the model with different kinds of information databases. Three cases are studied: 1) information in records is structured according to the UMLS knowledge sources; 2) information is able to be structured without error in the frame of the UMLS knowledge; 3) information cannot be structured. In each case the pattern-matching is processed by the projection rule according to the structure of information that has been implemented in the databases. CONCLUSION: The conceptual graphs theory provides with a homogeneous and powerful formalism able to represent both concepts, instances of concepts in medical contexts, and associations by means of relationships, and to represent data at different levels of details. The conceptual-graphs formalism allows powerful capabilities to operate a semantic integration of information databases using the UMLS knowledge sources.

Databases as Topic↗

ALFRED: a Web-accessible allele frequency database.

We present a Web-accessible database (ALFRED) that allows public access to gene frequency data for a diverse set of population samples and genetic systems. The data in ALFRED are modeled based on the experience and needs of a single laboratory, but with the expectation that the database will meet the needs of a much broader scientific community that needs population-specific gene frequency estimates. Our database currently contains data on more than 40 populations representing most major regions of the world and data on more than 150 genetic systems including SNPs, STRPs, and insertion-deletion polymorphisms. While data are not available for all population-genetic system combinations, over 2000 allele frequency tables already exist. In this paper, we enumerate the broad needs in the scientific domain, describe their significance, and describe how we have designed the database to meet those needs. We compare our database with dbSNP, the NCBI database that has a broader but overlapping purpose.

Alleles↗

Alternative to hand-tuning conductance-based models: construction and analysis of databases of model neurons.

Conventionally, the parameters of neuronal models are hand-tuned using trial-and-error searches to produce a desired behavior. Here, we present an alternative approach. We have generated a database of about 1.7 million single-compartment model neurons by independently varying 8 maximal membrane conductances based on measurements from lobster stomatogastric neurons. We classified the spontaneous electrical activity of each model neuron and its responsiveness to inputs during runtime with an adaptive algorithm and saved a reduced version of each neuron's activity pattern. Our analysis of the distribution of different activity types (silent, spiking, bursting, irregular) in the 8-dimensional conductance space indicates that the coarse grid of conductance values we chose is sufficient to capture the salient features of the distribution. The database can be searched for different combinations of neuron properties such as activity type, spike or burst frequency, resting potential, frequency-current relation, and phase-response curve. We demonstrate how the database can be screened for models that reproduce the behavior of a specific biological neuron and show that the contents of the database can give insight into the way a neuron's membrane conductances determine its activity pattern and response properties. Similar databases can be constructed to explore parameter spaces in multicompartmental models or small networks, or to examine the effects of changes in the voltage dependence of currents. In all cases, database searches can provide insight into how neuronal and network properties depend on the values of the parameters in the models.

Algorithms↗