Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A pilot bridging data integration and analytics: BioMediator and R?

Biological research today involves aggregating and analyzing large amounts of data from disparate sources. Tools such as the University of Washington's BioMediator system integrate heterogeneous data. Analytic packages such as the R environment have a rich set of tools to analyze biomedical research data. Our pilot project bridged data integration and analytics in a general way by successfully incorporating the BioMediator system into the R platform for specific analyses on neurophysiologic research data.

Brain↗

Data integration and visualization system for enabling conceptual biology.

MOTIVATION: Integration of heterogeneous data in life sciences is a growing and recognized challenge. The problem is not only to enable the study of such data within the context of a biological question but also more fundamentally, how to represent the available knowledge and make it accessible for mining. RESULTS: Our integration approach is based on the premise that relationships between biological entities can be represented as a complex network. The context dependency is achieved by a judicious use of distance measures on these networks. The biological entities and the distances between them are mapped for the purpose of visualization into the lower dimensional space using the Sammon's mapping. The system implementation is based on a multi-tier architecture using a native XML database and a software tool for querying and visualizing complex biological networks. The functionality of our system is demonstrated with two examples: (1) A multiple pathway retrieval, in which, given a pathway name, the system finds all the relationships related to the query by checking available metabolic pathway, transcriptional, signaling, protein-protein interaction and ontology annotation resources and (2) A protein neighborhood search, in which given a protein name, the system finds all its connected entities within a specified depth. These two examples show that our system is able to conceptually traverse different databases to produce testable hypotheses and lead towards answers to complex biological questions.

Computational Biology↗

Disparate systems, disparate data: integration, interfaces, and standards in emergency medicine information technology.

As part of the broader informatics consensus initiative sponsored by Academic Emergency Medicine, this report addresses the issues of integration, interfaces, and data standards and how they are relevant to information management in emergency medicine. The purpose of this report, and the workgroup that contributed to its content, is to provide emergency physicians and other stakeholders in the emergency informatics community a sense of direction as they design, build, and/or choose systems. Problems are identified, strategies to address these problems are discussed, and consensus recommendations are provided.

Emergency Medicine↗

Computational framework for the prediction of transcription factor binding sites by multiple data integration.

Control of gene expression is essential to the establishment and maintenance of all cell types, and its dysregulation is involved in pathogenesis of several diseases. Accurate computational predictions of transcription factor regulation may thus help in understanding complex diseases, including mental disorders in which dysregulation of neural gene expression is thought to play a key role. However, biological mechanisms underlying the regulation of gene expression are not completely understood, and predictions via bioinformatics tools are typically poorly specific. We developed a bioinformatics workflow for the prediction of transcription factor binding sites from several independent datasets. We show the advantages of integrating information based on evolutionary conservation and gene expression, when tackling the problem of binding site prediction. Consistent results were obtained on a large simulated dataset consisting of 13050 in silico promoter sequences, on a set of 161 human gene promoters for which binding sites are known, and on a smaller set of promoters of Myc target genes. Our computational framework for binding site prediction can integrate multiple sources of data, and its performance was tested on different datasets. Our results show that integrating information from multiple data sources, such as genomic sequence of genes' promoters, conservation over multiple species, and gene expression data, indeed improves the accuracy of computational predictions.

Animals↗

Anatomy of data integration.

Producing reliable information is the ultimate goal of data processing. The ocean of data created with the advances of science and technologies calls for integration of data coming from heterogeneous sources that are diverse in their purposes, business rules, underlying models and enabling technologies. Reference models, Semantic Web, standards, ontology, and other technologies enable fast and efficient merging of heterogeneous data, while the reliability of produced information is largely defined by how well the data represent the reality. In this paper, we initiate a framework for assessing the informational value of data that includes data dimensions; aligning data quality with business practices; identifying authoritative sources and integration keys; merging models; uniting updates of varying frequency and overlapping or gapped data sets.

Algorithms↗

Software package for integrated data processing for internal dose assessment in nuclear medicine (SPRIND).

PURPOSE: Internal radiation dose calculations are normally carried out using the Medical Internal Radiation Dose (MIRD) schema. This requires residence times of radiopharmaceutical activity and S-values for all organs of interest. Residence times can be obtained by quantitative nuclear imaging modalities. For dealing with S-values, the freeware packages MIRDOSE and, more recently, OLINDA/EXM are available. However, these software packages do not calculate residence times from image data. METHODS AND RESULTS: For this purpose, we developed an IDL-based software package for integrated data processing for internal dose assessment in nuclear medicine (SPRIND). SPRIND allows reading and viewing of planar whole-body scintigrams. Organ and background regions of interest (ROIs) can be drawn and are automatically mirrored from the anterior to the posterior view. ROI statistics are used to obtain anterior-posterior averaged counts for each organ, corrected for background activity and attenuation. Residence times for each organ are calculated based on effective decay. The total body biological half-time is calculated for use in the voiding bladder model. Red bone marrow absorbed dose can be calculated using bone regions in the scintigrams or by a blood-derived method. Finally, the results are written to a file in MIRDOSE-OLINDA/EXM format. Using scintigrams in DICOM, the complete analysis is gamma camera vendor independent, and can be performed on any computer using an IDL virtual machine. CONCLUSION: SPRIND is an easy-to-use software package for radiation dose assessment studies. It has made these studies less time consuming and less error prone.

Algorithms↗

Kleisli: a new tool for data integration in biology.

One of the central problems in bioinformatics is data retrieval and integration. The existing biological databases are geographically distributed across the Internet, complex and heterogeneous in data types and data structures, and constantly changing. With the current rapid growth of biomedical data, the challenge is how large volumes of data retrieved from multiple databases can be transformed and integrated automatically and flexibly. This article describes a powerful new tool, the Kleisli system, for complex queries across multiple databases and data integration.

Computational Biology↗

Expression array annotation using the BioMediator biological data integration system and the BioConductor analytic platform.

This paper presents the implementation of a model for expression array annotation (EAA) using the BioMediator biological data integration system along with BioConductor, an analytic tools platform. The model presented addresses the need for annotation sources identified during BioConductor inverted exclamation mark s development. Annotation provides us with well-curated genomic background knowledge for expression array analysis and interpretation. Annotation requests are constructed and posted to the query interface of the EAA package (the EAA model implemented as a component of BioConductor). The software enumerates all possible annotation paths for queries. These are then transformed to PQL queries and processed by BioMediator. Annotation entities returned from the EAA package answer the annotation request.

Computational Biology↗

Integrating data from biological experiments into metabolic networks with the DBE information system.

Modern 'omics'-technologies result in huge amounts of data about life processes. For analysis and data mining purposes this data has to be considered in the context of the underlying biological networks. This work presents an approach for integrating data from biological experiments into metabolic networks by mapping the data onto network elements and visualising the data enriched networks automatically. This methodology is implemented in DBE, an information system that supports the analysis and visualisation of experimental data in the context of metabolic networks. It consists of five parts: (1) the DBE-Database for consistent data storage, (2) the Excel-Importer application for the data import, (3) the DBE-Website as the interface for the system, (4) the DBE-Pictures application for the up- and download of binary (e. g. image) files, and (5) DBE-Gravisto, a network analysis and graph visualisation system. The usability of this approach is demonstrated in two examples.

Computational Biology↗

Cancer risk assessment for 1,3-butadiene: data integration opportunities.

The US Environmental Protection Agency recently released its new guidelines for carcinogen risk assessment together with supplemental guidance for assessing susceptibility from early-life exposure to carcinogens. In particular, these guidelines encourage the use of mechanistic data in support of dose-response characterization at doses below those at which an increase in tumor frequency over background levels might be detected. In this context of the utility of mechanistic data for human cancer risk assessment, the International Life Sciences Institute (ILSI) has developed a human relevance framework (HRF) that can be used to assess the plausibility of a mode of action (MoA) described for animal models operating in humans. The MoA is described as a sequence of key events and processes that result in an adverse outcome. A key event is a measurable precursor step that is in itself a necessary element of the MoA or is a bioindicator for such an element. A number of cellular and molecular perturbations have been identified as key events whereby DNA-reactive chemicals can produce tumors. These include DNA adducts in target tissues, gene mutations and/or chromosomal alterations in target tissues and enhanced cell proliferation in target tissues. This type of data integration approach to quantitative cancer risk assessment can be applied to 1,3-butadiene, for example, using data on biomarkers in exposed Czech workers [1]. For this study, an extensive range of biomarkers of exposure and response was assessed, including: polymorphisms in metabolizing enzymes; urinary concentrations of several metabolites of 1,3-butadiene; hemoglobin adducts; HPRT mutations in T-lymphocytes; chromosomal aberrations by FISH and conventional staining procedures; sister chromatid exchanges. Exposure levels were monitored in a comprehensive fashion. For risk assessment purposes, these data need to be considered in the context of how they inform the MoA for leukemia, the tumor type reported to be increased in synthetic rubber workers exposed to 1,3-butadiene. Also, for the HRF it is necessary to establish key events for a MoA in rodents for the induction of tumors by 1,3-butadiene. There is clearly a species difference in sensitivity to tumor induction, with mice being much more sensitive than rats; key events need to explain this difference. For butadiene, the MoA is DNA-reactivity and subsequent mutagenicity and so following the EPA's cancer guidelines, a linear extrapolation is used from the point of departure (POD), unless additional data support a non-linear extrapolation. For the present case, the human bioindicator data are not informative as far as dose-response characterization is concerned. Mouse chromosome aberration data for in vivo exposures might be used for establishing a POD, with linear extrapolation from this POD. The available cytogenetic data from rodent studies appear to be sufficiently extensive and consistent for this to be a viable approach. This approach of using MoA and key events to establish the human relevance can lead to the development of specific informative bioindicators of response that can be used as surrogates to predict the shape of the tumor dose response curve at low doses. Truly informative predictors of tumor responses should be able to provide estimates of human tumor frequencies at low, environmental exposures to 1,3-butadiene.

Animals↗

An overview of data integration methods for regional assessment.

The U.S. Environmental Protections Agency's (U.S. EPA) Regional Vulnerability Assessment(ReVA) program has focused much of its research over the last five years on developing and evaluating integration methods for spatial data. An initial strategic priority was to use existing data from monitoring programs, model results, and other spatial data. Because most of these data were not collected with an intention of integrating into a regional assessment of conditions and vulnerabilities, issues exist that may preclude the use of some methods or require some sort of data preparation. Additionally, to support multi-criteria decision-making, methods need to be able to address a series of assessment questions that provide insights into where environmental risks are a priority. This paper provides an overview of twelve spatial integration methods that can be applied towards regional assessment, along with preliminary results as to how sensitive each method is to data issues that will likely be encountered with the use of existing data.

Animals↗

Metabonomics: its potential as a tool in toxicology for safety assessment and data integration.

The functional genomic techniques of transcriptomics and proteomics promise unparalleled global information during the drug development process. However, if these technologies are used in isolation the large multivariate data sets produced are often difficult to interpret, and have the potential of missing key metabolic events (e.g. as a result of experimental noise in the system). To better understand the significance of these megavariate data the temporal changes in phenotype must be described. High resolution 1H NMR spectroscopy used in conjunction with pattern recognition provides one such tool for defining the dynamic phenotype of a cell, organ or organism in terms of a metabolic phenotype. In this review the benefits of this metabonomics/metabolomics approach to problems in toxicology will be discussed. One of the major benefits of this approach is its high throughput nature and cost effectiveness on a per sample basis. Using such a method the consortium for metabonomic toxicology (COMET) are currently investigating approximately 150 model liver and kidney toxins. This investigation will allow the generation of expert systems where liver and kidney toxicity can be predicted for model drug compounds, providing a new research tool in the field of drug metabolism. The review will also include how metabonomics may be used to investigate co-responses with transcripts and proteins involved in metabolism and stress responses, such as during drug induced fatty liver disease. By using data integration to combine metabolite analysis and gene expression profiling key perturbed metabolic pathways can be identified and used as a tool to investigate drug function.

Animals↗

GIMS: an integrated data storage and analysis environment for genomic and functional data.

Effective analyses in functional genomics require access to many kinds of biological data. For example, the analysis of upregulated genes in a microarray experiment might be aided by information concerning protein interactions or proteins' cellular locations. However, such information is often stored in different formats at different sites, in ways that may not be amenable to integrated analysis. The Genome Information Management System (GIMS) is an object database that integrates genomic data with data on the transcriptome, protein-protein interactions, metabolic pathways and annotations, such as gene ontology terms and identifiers. The resulting system supports the running of analyses over this integrated data resource, and provides comprehensive facilities for handling and interrelating the results of these analyses. GIMS has been used to store Saccharomyces cerevisiae data, and we demonstrate how the integrated storage of diverse types of data can be beneficial for analysis, using combinations of complex queries. As an example, we describe how GIMS has been used to analyse a collection of aryl alcohol dehydrogenase gene deletion mutants. The GIMS database can be accessed remotely using a Java application that can be downloaded from http://img.cs.man.ac.uk/gims.

Computational Biology↗

Ontologies and semantic data integration.

The increased generation of data in the pharmaceutical R&D process has failed to generate the expected returns in terms of enhanced productivity and pipelines. The inability of existing integration strategies to organize and apply the available knowledge to the range of real scientific and business issues is impacting on not only productivity but also transparency of information in crucial safety and regulatory applications. The new range of semantic technologies based on ontologies enables the proper integration of knowledge in a way that is reusable by several applications across businesses, from discovery to corporate affairs.

Databases as Topic↗

Improving and integrating data systems for public health surveillance.

The National Center for Health Statistics (NCHS) is the nation's principal health statistics agency, with a primary mission to collect, disseminate, and analyze health data. NCHS has a clear commitment to a wide range of improvements in surveillance and public health information systems. Building on its long history of conducting multipurpose surveys where the needs and interests of a variety of programmatic interests have to be accommodated, NCHS is working on a number of fronts to improve and better integrate data systems so that they will be more useful for public health surveillance. Examples include the redesign of the National Health Interview Survey, the integration of the Department of Health and Human Services' health surveys, the retooling of the vital statistics system, and the movement to subnational data collection.

Humans↗

Integrated data acquisition system for medical device testing and physiology research in compliance with good laboratory practices.

In seeking approval from the US Food and Drug Administration (FDA) for clinical trial evaluation of an experimental medical device, a sponsor is required to submit experimental findings and support documentation to demonstrate device safety and efficacy that are in compliance with Good Laboratory Practices (GLP). The objective of this project was to develop an integrated data acquisition (DAQ) system and documentation strategy for monitoring and recording physiological data when testing medical devices in accordance with GLP guidelines mandated by the FDA. Data aquisition systems were developed as stand-alone instrumentation racks containing transducer amplifiers and signal processors, analog-to-digital converters for data storage, visual display and graphical user-interfaces, power conditioners, and test measurement devices. Engineering standard operating procedures (SOP) were developed to provide a written step-by-step process for calibrating, validating, and certifying each individual instrumentation unit and the integrated DAQ system. Engineering staff received GLP and SOP training and then completed the calibration, validation, and certification process for the individual instrumentation components and integrated DAQ system. Eight integrated DAQ systems have been successfully developed that were inspected by regulatory affairs consultants and determined to meet GLP guidelines. Two of these DAQ systems were used to support 40 of the pre-clinical animal studies evaluating the AbiCor artificial heart (ABIOMED, Danvers, MA). Based in part on these pre-clinical animal data, the AbioCor clinical trials began in July 2001. The process of developing integrated DAQ systems, SOP, and the validation and certification methods used to ensure GLP compliance are presented in this article.

Equipment and Supplies↗

Integrated data mining and network pharmacology to explore the prescription patterns from a senior TCM oncologist's clinical practice in treating chemotherapy-induced hand-foot syndrome.

Hand-foot syndrome (HFS) is a common and refractory adverse effect of chemotherapy lacking specific therapeutic strategies currently. Traditional Chinese medicine (TCM) has shown empirical efficacy in clinical HFS management. This study integrated data mining and network pharmacology to systematically elucidate the medication principles and molecular mechanisms underlying Professor Gang Xie's prescriptions for HFS. All medical records from Professor Xie's specialist clinic (January 2020 to March 2025) were retrospectively collected and standardized in Excel. Prescriptions were analyzed through frequency statistics, association and clustering. Active ingredients of core herb pairs and their disease-related targets were identified using TCMSP, HERB, GeneCards, PharmGKB and GEO databases. Protein-protein interaction (PPI) networks, gene ontology (GO), and Kyoto encyclopedia of genes and genomes (KEGG) pathway analyses were performed. Molecular docking validated interactions between key bioactive compounds and targets. This study involved 217 prescriptions containing 150 herbs. Core herb combinations comprised Radix Astragali (Huangqi), Poria (Fuling), and Radix Pseudostellariae (Taizishen), predominantly classified as spleen-tonifying agents with warm properties, targeting lung, spleen, and stomach meridians. Network analysis identified 67 bioactive compounds and 899 disease targets. Quercetin, kaempferol, acacetin and luteolin were identified the key ingredients. The core targets (TP53, STAT3, PIK3CA, HSP90AA1, AKT1, CTNNB1, PI3KR1, MAPK1) were enriched in MAPK and PI3K-Akt signaling pathways. Molecular docking confirmed strong binding affinity between key compounds and targets. Professor Xie's therapeutic strategy for HFS emphasizes "spleen fortification, phlegm elimination, and stasis resolution." The core herb combination likely exerts anti-HFS effects via modulation of MAPK and PI3K-Akt pathways, providing a pharmacological basis for TCM-driven HFS management.

Network Pharmacology↗