Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Integrating data from natural language processing into a clinical information system.

Demographic data extracted from discharge summaries by natural language processing was compared to data gathered by a conventional hospital admitting system. Discrepancies in data were noted in names, age, sex, race, and ethnicity. Some differences are attributable to errors in collection: interaction with patient, dictation, transcription, and data entry. Very few differences were due to errors in natural language processing. Other differences can be used to critique existing data, or to enhance data with more detailed information. Discrepancies in data as elementary as patient demographics raise the issue of resolving conflicts when neither source of data is known to be more reliable. Clinical repositories can represent conflicting data from multiple sources, but clinical information systems must bear the cost of increased complexity in the application programs that will use the data.

Evaluation Studies as Topic↗

Mapping data elements to terminological resources for integrating biomedical data sources.

BACKGROUND: Data integration is a crucial task in the biomedical domain and integrating data sources is one approach to integrating data. Data elements (DEs) in particular play an important role in data integration. We combine schema- and instance-based approaches to mapping DEs to terminological resources in order to facilitate data sources integration. METHODS: We extracted DEs from eleven disparate biomedical sources. We compared these DEs to concepts and/or terms in biomedical controlled vocabularies and to reference DEs. We also exploited DE values to disambiguate underspecified DEs and to identify additional mappings. RESULTS: 82.5% of the 474 DEs studied are mapped to entries of a terminological resource and 74.7% of the whole set can be associated with reference DEs. Only 6.6% of the DEs had values that could be semantically typed. CONCLUSION: Our study suggests that the integration of biomedical sources can be achieved automatically with limited precision and largely facilitated by mapping DEs to terminological resources.

Abstracting and Indexing↗

Integrated patient data for optimal patient management: the value of laboratory data in quality improvement.

Managed care organizations are shifting from traditional utilization management programs to focus on initiatives that improve the health of an insured population. This strategy requires sophisticated data integration to identify at-risk individuals and track outcomes. Laboratory data are becoming increasingly valuable tools for managed care organizations and healthcare providers. The HEDIS Effectiveness of Care measures have incorporated laboratory data into several key performance indicators. By building a comprehensive repository of laboratory data that includes both procedure codes and laboratory values, managed care organizations can realize substantial savings by avoiding the costly medical record reviews required when administrative data are incomplete. In addition to tracking clinical outcomes, laboratory data provide the ability to risk-stratify a population to target high-risk individuals for case management and disease management interventions. Healthcare organizations face several challenges in the integration of laboratory data into medical databases and practice management software. Confidentiality is a key consideration in view of recent healthcare regulations. Providers of laboratory services should work collaboratively with organizations setting standards for healthcare informatics to facilitate the pooling of data for quality improvement and outcomes research. Health Level Seven, Inc. (HL7), Logical Observation Identifier Names and Codes (LOINC), and Systematized Nomenclature of Medicine (SNOMED) will likely play a key role in this process.

Clinical Laboratory Techniques↗

The model organism as a system: integrating 'omics' data sets.

Various technologies can be used to produce genome-scale, or 'omics', data sets that provide systems-level measurements for virtually all types of cellular components in a model organism. These data yield unprecedented views of the cellular inner workings. However, this abundance of information also presents many hurdles, the main one being the extraction of discernable biological meaning from multiple omics data sets. Nevertheless, researchers are rising to the challenge by using omics data integration to address fundamental biological questions that would increase our understanding of systems as a whole.

Animals↗

Biomarkers in psychotropic drug development: integration of data across multiple domains.

This review focuses on the current status of biomarkers and/or approaches critical to assessing novel neuroscience targets with an emphasis on new paradigms and challenges in this field of research. The importance of biomarker data integration for psychotropic drug development is illustrated with examples for clinically used medications and investigational drugs. The question remains how to verify access to the brain. Early imaging studies including micro-PET can help to overcome this. However, in case of delayed tracer development or because of no feasible application of brain imaging effects of the molecule, using CSF as a matrix could fill this gap. Proteomic research using CSF will hopefully have a major impact on the development of treatments for psychiatric disorders.

Animals↗

Integrating structured biological data by Kernel Maximum Mean Discrepancy.

MOTIVATION: Many problems in data integration in bioinformatics can be posed as one common question: Are two sets of observations generated by the same distribution? We propose a kernel-based statistical test for this problem, based on the fact that two distributions are different if and only if there exists at least one function having different expectation on the two distributions. Consequently we use the maximum discrepancy between function means as the basis of a test statistic. The Maximum Mean Discrepancy (MMD) can take advantage of the kernel trick, which allows us to apply it not only to vectors, but strings, sequences, graphs, and other common structured data types arising in molecular biology. RESULTS: We study the practical feasibility of an MMD-based test on three central data integration tasks: Testing cross-platform comparability of microarray data, cancer diagnosis, and data-content based schema matching for two different protein function classification schemas. In all of these experiments, including high-dimensional ones, MMD is very accurate in finding samples that were generated from the same distribution, and outperforms its best competitors. CONCLUSIONS: We have defined a novel statistical test of whether two samples are from the same distribution, compatible with both multivariate and structured data, that is fast, easy to implement, and works well, as confirmed by our experiments. AVAILABILITY: http://www.dbs.ifi.lmu.de/~borgward/MMD.

Algorithms↗