Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Using case management systems to integrate clinical and financial data.

Integrated delivery systems can use automated case management information systems to better manage relevant clinical and financial data across the continuum of care. Effective case management systems can help caregivers track clinical and financial information; match appropriate resources to patient needs; and analyze populations to identify risk, enhance adherence to clinical guidelines, and understand provider treatment profiles.

Case Management↗

Analysis of gene ontology features in microarray data using the Proteome BioKnowledge Library.

Microarray technology has resulted in an explosion of complex, valuable data. Integrating data analysis tools with a comprehensive underlying database would allow efficient identification of common properties among differentially regulated genes. In this study we sought to compare the utility of various databases in microarray analysis. The Proteome BioKnowledge Library (BKL), a manually curated, proteome-wide compilation of the scientific literature, was used to generate a list of Gene Ontology (GO) Biological Process (BP) terms enriched among proteins involved in cardiovascular disease. Analysis of DNA microarray data generated in a study of rat vascular smooth muscle cell responses revealed significant enrichment in a number of GO BPs that were also enriched among cardiovascular disease-related proteins. Using annotation from LocusLink and chip annotation from the Gene Expression Omnibus yielded fewer enriched cardiovascular disease-associated GO BP terms. Data sets of orthologous genes from mouse and human were generated using the BKL Retriever. Analysis of these sets focusing on BKL Disease annotation, revealed a significant association of these genes with cardiovascular disease. These results and the extensive presence of experimental evidence for BKL GO and Disease features, underscore the benefits of using this database for microarray analysis.

Animals↗

An object-oriented model for the integration of knowledge-based systems.

This paper discusses functional integration, data integration, and knowledge integration as basic problems concerning the integration of knowledge-based systems into a hospital information system. A system model for an integrated knowledge-based system is introduced. Object-oriented models for the systems meta-database, patient database, and knowledge-base are presented. It is expected that the reader is familiar with the basic concepts of the object-oriented approach.

Artificial Intelligence↗

Integrity of small data bases in computer analysis of dietary data.

The integrity of data bases to support microcomputer-based dietary analysis programs has become increasingly important to developers and users of nutritional analysis software. This paper reviews critical issues in maintaining data integrity during development of small nutritional data bases. Because a limited number of large, source data bases provides the data for smaller, special-purpose data bases, this review initially focuses on factors that affect the quality and precision of methodologies used in establishing large data bases. Issues discussed are accuracy of source data as determined by analytical methodology and imputation procedures, and methods for insuring representativeness of data. The effect of data transfer procedures on small data base integrity are discussed, including use of multiple sources and standardization of naming and coding conventions. Also reviewed are procedures for selecting reduced numbers of foods and nutrients without sacrificing accuracy of analysis, and methods currently in use for validating small data bases.

Databases, Factual↗

Contextualizing heterogeneous data for integration and inference.

Systems that attempt to integrate and analyze data from multiple data sources are greatly aided by the addition of specific semantic and metadata "context" that explicitly describes what a data value means. In this paper, we describe a systematic approach to constructing models of data and their context. Our approach provides a generic "template" for constructing such models. For each data source, a developer creates a customized model by filling in the tem-plate with predefined attributes and value. This approach facilitates model construction and provides consistent syntax and semantics among models created with the template. Systems that can process the template structure and attribute values can reason about any model so described. We used the template to create a detailed knowledge base for syndromic surveillance data integration and analysis. The knowledge base provided support for data integration, translation, and analysis methods.

Database Management Systems↗

Oncopacket: integration of cancer research data using GA4GH phenopackets.

SUMMARY: Lack of data integration remains a significant impediment to cancer research, and many analyses still require customized software to transform and prepare cancer data. We describe a software package to harmonize genetic and clinical cancer data into the GA4GH Phenopacket schema, an ISO standard for representing clinical case data. We integrated demographic, mutation, morphology, diagnosis, intervention, and survival data using case data from the National Cancer Institute for 12 cancer types. The Phenopacket standard provides a foundation for downstream use, including sophisticated statistical and AI/ML analyses. We demonstrate fitness for purpose by using the integrated data to recapitulate a known association between mutations in the gene encoding isocitrate dehydrogenase 1 and survival time in brain cancer patients. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at: https://github.com/monarch-initiative/oncopacket (archived at 10.5281/zenodo.15353125).

Humans↗

Technologies for integrating biological data.

The process of building a new database relevant to some field of study in biomedicine involves transforming, integrating and cleansing multiple data sources, as well as adding new material and annotations. This paper reviews some of the requirements of a general solution to this data integration problem. Several representative technologies and approaches to data integration in biomedicine are surveyed. Then some interesting features that separate the more general data integration technologies from the more specialised ones are highlighted.

Database Management Systems↗

Integration of data for gene annotation using the BioMediator system.

Gene annotation requires integration of data from multiple sources in order to functionally classify genes. We are using BioMediator, a general purpose data-integration solution, to develop a gene annotation system to automate the process of collecting data from disparate genomic databases. Integration of annotation data from multiple sources into a single format will facilitate use of analytic tools for the proper functional classification of genes.

Base Sequence↗

Information system architectures for syndromic surveillance.

INTRODUCTION: Public health agencies are developing the capacity to automatically acquire, integrate, and analyze clinical information for disease surveillance. The design of such surveillance systems might benefit from the incorporation of advanced architectures developed for biomedical data integration. Data integration is not unique to public health, and both information technology and academic research should influence development of these systems. OBJECTIVES: The goal of this paper is to describe the essential architectural components of a syndromic surveillance information system and discuss existing and potential architectural approaches to data integration. METHODS: This paper examines the role of data elements, vocabulary standards, data extraction, transport and security, transformation and normalization, and analysis data sets in developing disease-surveillance systems. It then discusses automated surveillance systems in the context of biomedical and computer science research in data integration, both to characterize existing systems and to indicate potential avenues of investigation to build systems that support public health practice. RESULTS: The Public Health Information Network (PHIN) identifies best practices for essential architectural components of a syndromic surveillance system. A schema for classifying biomedical data-integration software is useful for classifying present approaches to syndromic surveillance and for describing architectural variation. CONCLUSIONS: Public health informatics and computer science research in data-integration systems can supplement approaches recommended by PHIN and provide information for future public health surveillance systems.

Bioterrorism↗

Monitoring the integration of hospital information systems: How it may ensure and improve the quality of data.

Integration of hospital departmental information systems (HDIS) has become a common but difficult issue. In May 2003, the Department of Biostatistics and Medical Informatics implemented a Virtual Electronic Patient Record (VEPR) for the Hospital S. João (HSJ), a university hospital with over 1350 beds. The system integrates clinical data from 10 legacy HDIS plus the Hospital Administrative Database (HAD), aiming to deliver all patient information to health professionals. Currently, around 500 medical doctors use the system on a regular basis and the HSJ-VEPR retrieves an average of 3,000 new reports per day, in PDF or HTML formats. This paper describes and discusses the role of monitoring in the assurance and improvement of data quality. Three approaches were put in place: (a) monitoring the HSJ-VEPR concerning the frequency of clinical records retrieved from the DIS by checking if the daily number of reports sent by the HDIS fell in the normal range from similar week days; (b) monitoring inconsistencies in the patient's identification by cross-checking between HDIS and HAD; and (c) monitoring the integrity of clinical records delivered to medical doctors through the HSJ-VEPR by checking their digital signature. During 2005, the monitoring system detected 53 unusual frequency patterns of which 44 corresponded to real problems. Over a 6 months period, more than 400 alerts were generated concerning inconsistencies in the patient's identification found in laboratory reports. Nevertheless, a significant reduction in the number of these inconsistencies occurred - from 116 in July to 10 in December 2005--due to implementation of preventive measures by the DIS. Finally, report's integrity was checked each time the report was asked to be visualized i.e. in more than one hundred thousand times during a one year period. In conclusion, all information available in hospital information systems can and should be used to trigger alerts of malfunctions and inconsistencies, in order to improve data quality and ensure a better health care.

Hospital Information Systems↗

Integration of data and management tools into the new york state medicaid managed care encounter data system.

The New York State Department of Health has created a data warehouse to analyze and evaluate the Medicaid managed care program. Online query tools and reports, grouping tools such as Diagnostic Related Groups, and measurement tools such as Health Plan Data and Information Set (HEDIS) measures have been incorporated into the data warehouse. Other public health data sets including birth certificate data have also been integrated. The result is a powerful data set that can analyze information quickly and efficiently, with built-in data intelligence. Developed over time, this system can provide states, health insurance companies, and health data consortiums a roadmap on how to implement an integrated data warehouse solution.

Databases, Factual↗

Regulatory perspectives on data safety monitoring boards: protecting the integrity of data.

The use of interim analyses and data safety monitoring boards (DSMBs) can assist greatly in the timely determination of whether or not a medicine has an acceptable benefit-risk profile. Regulatory authorities regard the appropriate use of interim analyses favourably, but will consider the extent to which the conduct of interim analyses and the involvement of DSMBs may have compromised the evidence of efficacy and safety from a clinical trial. Issues of particular concern, which may potentially introduce bias, include the dissemination of interim data and the rules by which a trial might be terminated early. If data from trials which employ a DSMB are to be considered reliable and scientifically valid, it is the responsibility of the trial sponsor to demonstrate that the DSMB is set up and run appropriately and to verify that any bias introduced has had no important effect on the conclusions.

Clinical Trials Data Monitoring Committees↗

Integrative genomics: in silico coupling of rat physiology and complex traits with mouse and human data.

Integration of the large variety of genome maps from several organisms provides the mechanism by which physiological knowledge obtained in model systems such as the rat can be projected onto the human genome to further the research on human disease. The release of the rat genome sequence provides new information for studies using the rat model and is a key reference against which existing and new rat physiological results can be aligned. Previously, we described comparative maps of the rat, mouse, and human based on EST sequence comparisons combined with radiation hybrid maps. Here, we use new data and introduce the Integrated Genomics Environment, an extensive database of curated and integrated maps, markers, and physiological results. These results are integrated by using VCMapview, a java-based map integration and visualization tool. This unique environment allows researchers to relate results from cytogenetic, genetic, and radiation hybrid studies to the genome sequence and compare regions of interest between human, mouse, and rat. Integrating rat physiology with mouse genetics and clinical results from human by using the respective genomes provides a novel route to capitalize on comparative genomics and the strengths of model organism biology.

Animals↗

JASMINE: A powerful representation learning method for enhanced analysis of incomplete multi-omics data.

Integrative analysis of multi-omics data provides a more comprehensive and nuanced view of a subject's biological state. However, high-dimensionality and ubiquitous modality missingness present significant analytical challenges. Existing methods for incomplete multi-omics data are scarce, do not fully leverage both modality-specific and shared information, and produce task-biased representations. We propose JASMINE, a self-supervised representation learning method for incomplete multi-omics data that preserves both modality-specific and joint information and enhances sample similarity structure. JASMINE produces embeddings that achieve superior performance across multiple tasks for two different incomplete multi-omics datasets while requiring only a single round of training per dataset.

missing data↗

Challenges in integrating biological data sources.

Scientific data of importance to biologists reside in a number of different data sources, such as GenBank, GSDB, SWISS-PROT, EMBL, and OMIM, among many others. Some of these data sources are conventional databases implemented using database management systems (DBMSs) and others are structured files maintained in a number of different formats (e.g., ASN.1 and ACE). In addition, software packages such as sequence analysis packages (e.g., BLAST and FASTA) produce data and can therefore be viewed as data sources. To counter the increasing dispersion and heterogeneity of data, different approaches to integrating these data sources are appearing throughout the bioinformatics community. This paper surveys the technical challenges to integration, classifies the approaches, and critiques the available tools and methodologies.

Chromosomes, Artificial, Yeast↗

From information to understanding: the role of model organism databases in comparative and functional genomics.

Data integration is key to functional and comparative genomics because integration allows diverse data types to be evaluated in new contexts. To achieve data integration in a scalable and sensible way, semantic standards are needed, both for naming things (standardized nomenclatures, use of key words) and also for knowledge representation. The Mouse Genome Informatics database and other model organism databases help to close the gap between information and understanding of biological processes because these resources enforce well-defined nomenclature and knowledge representation standards. Model organism databases have a critical role to play in ensuring that diverse kinds of data, especially genome-scale data sets and information, remain useful to the biological community in the long-term. The efforts of model organism database groups ensure not only that organism-specific data are integrated, curated and accessible but also that the information is structured in such a way that comparison of biological knowledge across model organisms is facilitated.

Animals↗

Integration and cross-validation of high-throughput gene expression data: comparing heterogeneous data sets.

Data analysis--not data production--is becoming the bottleneck in gene expression research. Data integration is necessary to cope with an ever increasing amount of data, to cross-validate noisy data sets, and to gain broad interdisciplinary views of large biological data sets. New Internet resources may help researchers to combine data sets across different gene expression platforms. However, noise and disparities in experimental protocols strongly limit data integration. A detailed review of four selected studies reveals how some of these limitations may be circumvented and illustrates what can be achieved through data integration.

Animals↗