Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Development of a data warehouse at an academic health system: knowing a place for the first time.

In 1998, the University of Michigan Health System embarked upon the design, development, and implementation of an enterprise-wide data warehouse, intending to use prioritized business questions to drive its design and implementation. Because of the decentralized nature of the academic health system and the development team's inability to identify and prioritize those institutional business questions, however, a bottom-up approach was used to develop the enterprise-wide data warehouse. Specific important data sets were identified for inclusion, and the technical team designed the system with an enterprise view and architecture rather than as a series of data marts. Using this incremental approach of adding data sets, institutional leaders were able to experience and then further define successful use of the integrated data made available to them. Even as requests for the use and expansion of the data warehouse outstrip the resources assigned for support, the data warehouse has become an integral component of the institution's information management strategy. The authors discuss the approach, process, current status, and successes and failures of the data warehouse.

Academic Medical Centers↗

Quality control methods for data entry in pathology using a computerized data management system based on an extended data dictionary.

In pathology, computerized data management systems have been used increasingly to facilitate a more efficient supply of information. Since data entry precedes data utilization, the reliability of the information stored strongly depends on the quality of data input. Despite its potential capability, most personal computer-based database software does not provide versatile and user-friendly data validation procedures. Therefore, we developed a data dictionary-driven data management system that enables the user to perform extensive validation routines without the need for hard programming. Using examples from an existing database for endometrial carcinomas, different types of data errors and their error traps are explained. It is pointed out that data type definitions, defaults, templates, or picture clauses are suitable means to avoid formal errors. Validations on data domains and ranges test whether data fall into a predefined scope. Relational checks control data validity within a context of different data items, whereas process routines provide automatic data computation, thereby circumventing user input. By exploiting the facilities of an extended data dictionary, a powerful tool is made available to secure various aspects of data integrity simultaneously with input. In this way, computerized data quality control can improve the efficiency and reliability of data management tasks in pathology.

Medical Informatics Computing↗

The Defense Medical Surveillance System and the Department of Defense serum repository: glimpses of the future of public health surveillance.

The Defense Medical Surveillance System (DMSS) is the central repository of medical surveillance data for the US armed forces. The DMSS integrates data from sources worldwide in a continuously expanding relational database that documents the military and medical experiences of service members throughout their careers. The Department of Defense Serum Repository (DoDSR) is a central archive of sera drawn from service members for medical surveillance purposes. Currently, the DMSS contains data relevant to more than 7 million individuals who have served in the armed forces since 1990, and the DoDSR contains more than 27 million specimens that are linkable to data in the DMSS. Recent applications of the DMSS and DoDSR provide glimpses of the capabilities and uses of comprehensive public health surveillance systems.

Confidentiality↗

Benefits and threats of new technologies.

In this paper first the effect of new hardware and software technologies on threats to data integrity and usage integrity is considered. Next the potential is considered of new technical facilities for improving the protection. It is concluded that the increasing risks for data and usage integrity are not counter balanced at present by new protection measures. A concerted action is proposed to face this problem. Especially action is proposed to develop methods for quality assurance of software, for access control in networks and to improve data/usage integrity around PC's.

Computer Security↗

An alternative method for the analysis of neuron passive electrical data which uses integrals of voltage transients.

The traditional method for analyzing passive electrical data from neurons when specific morphological data are unavailable consists of decomposing the voltage response of the cell into a series of exponential functions (the peeling method) and substituting the time constants of these exponential functions into equations derived from cable theory (Rall W, Core conductor theory and cable properties of neurons. In: Handbook of Physiology. The Nervous System. Cellular Biology of Neurons. Bethesda, MD. Am Physiol Soc. Section 1, Part 1, 1977;1(3):39-97). In the present report, an alternative method is examined for analyzing these kinds of data, the integrals of transients method (Eisenberg RS, Mathias RT. Structural analysis of electrical properties of cells and tissues. CRC Critical Reviews in Bioengineering 1980;4:203-232). The integrals required are easily obtained from input resistance data and any theoretical model that is appropriate for the neurons under study can be used, provided that the impedance function can be determined. In order to demonstrate this alternative method, a simple 3-compartment model with both dendritic taper and somatic shunt is used to model data obtained from fast-type alpha-motoneurons in the spinal cord of the cat. These results are compared with results obtained using the traditional peeling method. This comparison indicates that passive electrical data from fast-type motoneurons are best analyzed using a theoretical model that includes both dendritic taper and somatic shunt. Furthermore, our results show that the integrals of transients method can facilitate this analysis.

Animals↗

Care planning as a strategy to manage variation in practice: from care plan to integrated person-based record.

This article begins with a summary of the trend toward a person-based health record, and the need to integrate data from a variety of sources to achieve this. A project is described that demonstrated problems with the structure of nursing care plans. These problems affected the ability to integrate care plan data into a clinical database capable of analysis to link control of process with clinical outcome. A second project is described that focused on the development of data sets holding higher-level descriptions suitable for the maintenance of a person-based record, but at a summarized level and with no clinical detail. Finally, a prototype care planning system is described that, while maintaining the data required by the Nursing Process, was more flexibly structured to support analysis and hierarchical levels of description.

Community Health Nursing↗

Ligand Depot: a data warehouse for ligands bound to macromolecules.

UNLABELLED: Ligand Depot is an integrated data resource for finding information about small molecules bound to proteins and nucleic acids. The initial release (version 1.0, November, 2003) focuses on providing chemical and structural information for small molecules found as part of the structures deposited in the Protein Data Bank. Ligand Depot accepts keyword-based queries and also provides a graphical interface for performing chemical substructure searches. A wide variety of web resources that contain information on small molecules may also be accessed through Ligand Depot. AVAILABILITY: Ligand Depot is available at http://ligand-depot.rutgers.edu/. Version 1.0 supports multiple operating systems including Windows, Unix, Linux and the Macintosh operating system. The current drawing tool works in Internet Explorer, Netscape and Mozilla on Windows, Unix and Linux.

Binding Sites↗

The future of the perfusion record: automated data collection vs. manual recording.

The perfusion record, whether manually recorded or computer generated, is a legal representation of the procedure. The handwritten perfusion record has been the most common method of recording events that occur during cardiopulmonary bypass. This record is of significant contrast to the integrated data management systems available that provide continuous collection of data automatically or by means of a few keystrokes. Additionally, an increasing number of monitoring devices are available to assist in the management of patients on bypass. These devices are becoming more complex and provide more data for the perfusionist to monitor and record. Most of the data from these can be downloaded automatically into online data management systems, allowing more time for the perfusionist to concentrate on the patient while simultaneously producing a more accurate record. In this prospective report, we compared 17 cases that were recorded using both manual and electronic data collection techniques. The perfusionist in charge of the case recorded the perfusion using the manual technique while a second perfusionist entered relevant events on the electronic record generated by the Stockert S3 Data Management System/Data Bahn (Munich, Germany). Analysis of the two types of perfusion records showed significant variations in the recorded information. Areas that showed the most inconsistency included measurement of the perfusion pressures, flow, blood temperatures, cardioplegia delivery details, and the recording of events, with the electronic record superior in the integrity of the data. In addition, the limitations of the electronic system were also shown by the lack of electronic gas flow data in our hardware. Our results confirm the importance of accurate methods of recording of perfusion events. The use of an automated system provides the opportunity to minimize transcription error and bias. This study highlights the limitation of spot recording of perfusion events in the overall record keeping for perfusion management.

Cardiopulmonary Bypass↗

euGenes: a eukaryote genome information system.

euGenes is a genome information system and database that provides a common summary of eukaryote genes and genomes, at http://iubio.bio.indiana.edu/eugenes/. Seven popular genomes are included: human, mouse, fruitfly, Caenorhabditis elegans worm, Saccharomyces yeast, Arabidopsis mustard weed and zebrafish, with more planned. This information, automatically extracted and updated from several source databases, offers features not readily available through other genome databases to bioscientists looking for gene relationships across organisms. The database describes 150 000 known, predicted and orphan genes, using consistent gene names along with their homologies and associations with a standard vocabulary of molecular functions, cell locations and biological processes. Usable whole-genome maps including features, chromosome locations and molecular data integration are available, as are options to retrieve sequences from these genomes. Search and retrieval methods for these data are easy to use and efficient, allowing one to ask combined questions of sequence features, protein functions and other gene attributes, and fetch results in reports, computable tabular outputs or bulk database forms. These summarized data are useful for integration in other projects, such as gene expression databases. euGenes provides an extensible, flexible genome information system for many organisms.

Animals↗

iModMix: integrative module analysis for multi-omics data.

SUMMARY: Integrative Module Analysis for Multi-omics Data (iModMix) is a biology-agnostic framework that enables the discovery of novel associations across any type of quantitative abundance data, including but not limited to transcriptomics, proteomics, and metabolomics. Instead of relying on pathway annotations or prior biological knowledge, iModMix constructs data-driven modules using graphical lasso to estimate sparse networks from omics features. These modules are summarized into eigenfeatures and correlated across datasets for horizontal integration, while preserving the distinct feature sets and interpretability of each omics type. iModMix operates directly on matrices containing expression or abundances for a wide range of features, including but not limited to genes, proteins, and metabolites. Because it does not rely on annotations (e.g., KEGG identifiers), it can seamlessly incorporate both identified and unidentified metabolites, addressing a key limitation of many existing metabolomics tools. iModMix is available as a user-friendly R Shiny application requiring no programming expertise (https://imodmix.moffitt.org), and as a Bioconductor R package for advanced users (https://bioconductor.org/packages/release/bioc/html/iModMix.html). The tool includes several public and in-house datasets to illustrate its utility in identifying novel multi-omics relationships in diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: iModMix is freely available from Bioconductor (https://bioconductor.org/packages/release/bioc/html/iModMix.html), and the example dataset package (iModMixData) is also available from Bioconductor (https://bioconductor.org/packages/release/ data/experiment/html/iModMixData.html). The R package source code and Docker are available from GitHub: https://github.com/biodatalab/iModMix. Shiny application can be accessed at: https://imodmix.moffitt.org.

Multiomics↗

Data management and quality assurance for an International project: the Indo-US Cross-National Dementia Epidemiology Study.

BACKGROUND: Data management and quality assurance play a vital but often neglected role in ensuring high quality research, particularly in collaborative and international studies. OBJECTIVE: A data management and quality assurance program was set up for a cross-national epidemiological study of Alzheimer's disease, with centers in India and the United States. METHODS: The study involved (a) the development of instruments for the assessment of elderly illiterate Hindi-speaking individuals; and (b) the use of those instruments to carry out an epidemiological study in a population-based cohort of over 5000 persons. Responsibility for data management and quality assurance was shared between the two sites. A cooperative system was instituted for forms and edit development, data entry, checking, transmission, and further checking to ensure that quality data were available for timely analysis. A quality control software program (CHECKS) was written expressly for this project to ensure the highest possible level of data integrity. CONCLUSIONS: This report addresses issues particularly relevant to data management and quality assurance at developing country sites, and to collaborations between sites in developed and developing countries.

Alzheimer Disease↗

Design and implementation of a picture archiving and communication system: the second time.

This report describes the authors' experience in the design and implementation of two large scale picture archiving and communication systems (PACS) during the past 10 years. The first system, which is in daily clinical operation was developed at University of California, Los Angeles from 1983 to 1992. The second system, which continues evolving, has been in development at University of California, San Francisco (UCSF) since 1992. The report highlights the differences between the two systems and points out the gradual change in the PACS design concept during the past 10 years from a closed architecture to an open hospital-integrated system. Both systems focus on system reliability and data integrity, with 24-hour on-line service and no loss of images. The major difference between the two systems is that the UCSF PACS infrastructure design is a completely open architecture and the system implementation uses more advanced technologies in computer software, digital communication, system interface, and stable industry standards. Such a PACS can withstand future technology changes without rendering the system obsolete, an essential criterion in any PACS design.

Diagnostic Imaging↗

Predicting end-stage renal disease: Bayesian perspective of information transfer in the clinical decision-making process at the individual level.

BACKGROUND: Predicting outcomes such as end-stage renal disease (ESRD) by integration and better utilization at individual level of epidemiologic data may facilitate clinical decision-making processes. METHODS: To predict individual ESRD risk in an average patient in the United States, ESRD prevalence and levels of uncertainty and conditional risk factors independence were considered by population data (1998) and pooled analysis of 11 randomized trials. Data integration and input were by decision-tree simulation approach (simple, parallel, and sequential scenarios) and Bayes' theorem. Sensitivity analysis and risk profiles were employed to address uncertainty and assess different risk factor combinations. A health state values, associated with ESRD outcome levels, were taken from the literature. RESULTS: In this theoretical study, we provided a scholarly example about the use of two known risk factors (urinary protein >/=3 g/day and systolic blood pressure >/=140 mm Hg) to predict individual ESRD risk in an average patient in the United States. The highest posterior (decisional) probability of ESRD occurrence (risk of 3.61% to 5.07%) in the individual patient was associated with the worst health state, as assessed by multidimensional scenarios when both risk factors were present. CONCLUSION: Decision tree models through an empirical Bayesian approach may serve to predict the individual ESRD risk on the basis of simple epidemiologic, demographic, and clinical information that is easily available already at the first patient evaluation.

Algorithms↗

Implementing a data warehouse at Inglis Innovative Services.

Data warehouses, data marts, and data mining have been hot topics in the 1990s, offering the promise of a vault of corporate data ripe for decision making. As is true with all promising technologies, the key issue is how to get started. Implementation of a corporate data warehouse involves a lot more than spending a huge amount of money on hardware, software, and consultants. Successful implementation of a data warehouse involves a corporate treasure hunt--identifying and cataloging data. It involves data ownership, data integrity, and business process analysis to determine what the data are, who owns them, how reliable they are, and how they are processed. Finally, implementation of the warehouse drives the issue of how good the decisions are that are based on the information in the warehouse. This article presents a case study of how one healthcare facility dealt with the challenges of implementing a data warehouse.

Computer Communication Networks↗

Cost-effective health care: new data.

The key to health care programs that meet their goals is to integrate data, coordinate care and ensure a patient-centered not cost-centered, focus. Then the purchaser can achieve the desired decrease in cost of care, increase in quality of care, improvement in quality of life, improvement in job performance, decrease in disability and decrease in absenteeism.

Cost-Benefit Analysis↗

Towards an e-biology of ageing: integrating theory and data.

Ageing is a highly complex process; it involves interactions between numerous biochemical and cellular mechanisms that affect many tissues in an organism. Although work on the biology of ageing is now advancing quickly, this inherent complexity means that information remains highly fragmented. We describe how a new web-based modelling initiative is seeking to integrate data and hypotheses from diverse biological sources.

Aging↗

[Creutzfeldt-Jakob disease and other human forms of transmissible spongiform encephalopathy in Italy: a mortality study carried out from different data sources].

Creutzfeldt-Jakob Disease (CJD) is a rare pathology (about 1 case per million) but it has a great importance for Public Health; the Italian National CJD register has been established in the Istituto Superiore di Sanita (ISS) since 1993, and epidemiological studies on CJD have been carried out as well. This paper reports a mortality study carried out comparing and integrating data from the two available sources: the National CJD Register and the Italian Data Base on Mortality, processed by the ISS Statistics Unit, on the data collected by the Italian Census Bureau (ISTAT). The study allowed to estimate: the underreporting of CJD mortality to both sources, the misclassification of ISTAT data and the integrated mortality rates from CJD in Italy: 1.58 per million persons aged 25 or more, average rate during the period 1993-1999.

Academies and Institutes↗

Electronic messaging between primary and secondary care: a four-year case report.

OBJECTIVE: To observe how electronic messaging between a hospital consultant and general practitioners (GPs) in 15 practices about patients suffering from diabetes evolved over a 3-year period after an initial 1-year study. DESIGN: Case report. Electronic messages between a hospital consultant and GPs were counted. The authors determined whether a message sent by the consultant was integrated into the receiving GP's electronic medical record system. After the observation period, the GPs answered a questionnaire. MEASUREMENTS: The number of electronic messages and the percentage of messages integrated into the electronic medical record. RESULTS: The volume of messages was maintained during the 3 years after the original study. In the original study, the percentage of the messages integrated by the GPs increased during the year. After that study, however, seven GPs stopped integrating data from messages. The extent to which received messages were integrated varied widely among practices. CONCLUSION: The authors conclude that extrapolation of the results of the original study would have led to incorrect conclusions. Although the volume of messages remained stable after the original study, GPs changed their method of handling messages. Initially, all GPs used the opportunity to copy data from the messages into their own records. At the end of the observation period (that is, the 3 years after completion of the original study), more than 50 percent of GPs had ceased copying data from the messages into their own records. The majority of GPs, however, wanted to expand the use of electronic messaging.

Attitude of Health Personnel↗