Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Drug target ontology to classify and integrate drug discovery data.

BACKGROUND: One of the most successful approaches to develop new small molecule therapeutics has been to start from a validated druggable protein target. However, only a small subset of potentially druggable targets has attracted significant research and development resources. The Illuminating the Druggable Genome (IDG) project develops resources to catalyze the development of likely targetable, yet currently understudied prospective drug targets. A central component of the IDG program is a comprehensive knowledge resource of the druggable genome. RESULTS: As part of that effort, we have developed a framework to integrate, navigate, and analyze drug discovery data based on formalized and standardized classifications and annotations of druggable protein targets, the Drug Target Ontology (DTO). DTO was constructed by extensive curation and consolidation of various resources. DTO classifies the four major drug target protein families, GPCRs, kinases, ion channels and nuclear receptors, based on phylogenecity, function, target development level, disease association, tissue expression, chemical ligand and substrate characteristics, and target-family specific characteristics. The formal ontology was built using a new software tool to auto-generate most axioms from a database while supporting manual knowledge acquisition. A modular, hierarchical implementation facilitate ontology development and maintenance and makes use of various external ontologies, thus integrating the DTO into the ecosystem of biomedical ontologies. As a formal OWL-DL ontology, DTO contains asserted and inferred axioms. Modeling data from the Library of Integrated Network-based Cellular Signatures (LINCS) program illustrates the potential of DTO for contextual data integration and nuanced definition of important drug target characteristics. DTO has been implemented in the IDG user interface Portal, Pharos and the TIN-X explorer of protein target disease relationships. CONCLUSIONS: DTO was built based on the need for a formal semantic model for druggable targets including various related information such as protein, gene, protein domain, protein structure, binding site, small molecule drug, mechanism of action, protein tissue localization, disease association, and many other types of information. DTO will further facilitate the otherwise challenging integration and formal linking to biological assays, phenotypes, disease models, drug poly-pharmacology, binding kinetics and many other processes, functions and qualities that are at the core of drug discovery. The first version of DTO is publically available via the website http://drugtargetontology.org/ , Github ( http://github.com/DrugTargetOntology/DTO ), and the NCBO Bioportal ( http://bioportal.bioontology.org/ontologies/DTO ). The long-term goal of DTO is to provide such an integrative framework and to populate the ontology with this information as a community resource.

Biological Ontologies↗

The challenge of integrating disparate high-content data: epidemiological, clinical and laboratory data collected during an in-hospital study of chronic fatigue syndrome.

Chronic fatigue syndrome (CFS) is a debilitating illness characterized by multiple unexplained symptoms including fatigue, cognitive impairment and pain. People with CFS have no characteristic physical signs or diagnostic laboratory abnormalities, and the etiology and pathophysiology remain unknown. CFS represents a complex illness that includes alterations in homeostatic systems, involves multiple body systems and results from the combined action of many genes, environmental factors and risk-conferring behavior. In order to achieve understanding of complex illnesses, such as CFS, studies must collect relevant epidemiological, clinical and laboratory data and then integrate, analyze and interpret the information so as to obtain meaningful clinical and biological insight. This issue of Pharmacogenomics represents such an approach to CFS. Data was collected during a 2-day in-hospital study of persons with CFS, other medically and psychiatrically unexplained fatiguing illnesses and nonfatigued controls identified from the general population of Wichita, KS, USA. While in the hospital, the participants' psychiatric status, sleep characteristics and cognitive functioning was evaluated, and biological samples were collected to measure neuroendocrine status, autonomic nervous system function, systemic cytokines and peripheral blood gene expression. The data generated from these assessments was made available to a multidisciplinary group of 20 investigators from around the world who were challenged with revealing new insight and algorithms for integration of this complex, high-content data and, if possible, identifying molecular markers and elucidating pathophysiology of chronic fatigue. The group was divided into four teams with representation from the disciplines of medicine, mathematics, biology, engineering and computer science. The papers in this issue are the culmination of this 6-month challenge, and demonstrate that data integration and multidisciplinary collaboration can indeed yield novel approaches for handling large, complex datasets, and reveal new insight and relevance to a complex illness such as CFS.

Adult↗

Integrating task and data flow analyses using the pentanalysis technique.

Pentanalysis is a new method that allows the integration of task analysis with data flow analysis. This is novel and desirable as it provides one means by which a human-computer interaction orientated method can provide input to the requirements specification stage of the software life cycle. It bridges the gulf that exists between this stage and one major representation, data flow diagrams, commonly used in the software design stage. The method is described and a worked example is provided to illustrate it.

Humans↗

[Integrated gait analysis: a complex expert system approach to detailed data evaluation].

Integrated Gait Analysis is still not accepted as an important tool in a routine clinical environment because of its lack of assessment procedures closely based on the raw data measurements. Furthermore, an adequate treatment of the complex process of human movement is missing. This contribution is intended to successively describe the scheme of a dedicated approach for gait analysis data assessment. Therein, different technologies already discussed in the literature have been used. Some examples are given for the purpose of illustration. To provide in advance sufficiently different and well-known measurement conditions, extreme manipulations in the alignment of below-knee prostheses have been carried out. It is shown that some specific data formats consisting of subgroups of gait patterns have turned out to be important for further partial processing and for interfacing the different subsystems. However, due to data characteristics of the raw data and the lack of a well-founded a priori knowledge, some principal questions have to be left open. Therefore, on the basis of the different results presented below and their discussion, a more detailed model of the process is proposed for further research following this concept.

Artificial Limbs↗

PowerMarker: an integrated analysis environment for genetic marker analysis.

SUMMARY: PowerMarker delivers a data-driven, integrated analysis environment (IAE) for genetic data. The IAE integrates data management, analysis and visualization in a user-friendly graphical user interface. It accelerates the analysis lifecycle and enables users to maintain data integrity throughout the process. An ever-growing list of more than 50 different statistical analyses for genetic markers has been implemented in PowerMarker. AVAILABILITY: www.powermarker.net

Algorithms↗

Comprehensive post-genomic data analysis approaches integrating biochemical pathway maps.

Post-genomic era research is focusing on studies to attribute functions to genes and their encoded proteins, and to describe the regulatory networks controlling metabolic, protein synthesis and signal transduction pathways. To facilitate the analysis of experiments using post-genomic technologies, new concepts for linking the vast amount of raw data to a biological context have to be developed. Visual representations of pathways help biologists to understand the complex relationships between components of metabolic networks, and provide an invaluable resource for the integration of transcriptomics, proteomics and metabolomics data sets. Besides providing an overview of currently available bioinformatic tools for plant scientists, we introduce BioPathAt, a newly developed visual interface that allows the knowledge-based analysis of genome-scale data by integrating biochemical pathway maps (BioPathAtMAPS module) with a manually scrutinized gene-function database (BioPathAtDB) for the model plant Arabidopsis thaliana. In addition, we discuss approaches for generating a biochemical pathway knowledge database for A. thaliana that includes, in addition to accurate annotation, condensed experimental information regarding in vitro and in vivo gene/protein function.

Computational Biology↗

Merging lab, drug data boosts disease management.

Merging lab and pharmacy data can benefit disease management. Integrating lab data with other data sets gave Blue Cross Blue Shield of Southeast Michigan a way to focus its diabetes disease management program, assess the effectiveness of therapy, and profile physicians. Learn how the plan's experience in data integration can apply to hospital systems and why pursuing a partner in such a project might be a wise move.

Clinical Laboratory Information Systems↗

A review of medical imaging informatics.

This review of medical imaging informatics is a survey of current developments in an exciting field. The focus is on informatics issues rather than traditional data processing and information systems, such as picture archiving and communications systems (PACS) and image processing and analysis systems. In this review, we address imaging informatics issues within the requirements of an informatics system defined by the American Medical Informatics Association. With these requirements as a framework, we review, in four sections: (1) Methods to present imaging and associated data without causing an overload, including image study summarization, content-based medical image retrieval, and natural language processing of text data. (2) Data modeling techniques to represent clinical data with focus on an image data model, including general-purpose time-based multimedia data models, health-care-specific data models, knowledge models, and problem-centric data models. (3) Methods to integrate medical data information from heterogeneous clinical data sources. Advances in centralized databases and mediated architectures are reviewed along with a discussion on our efforts at data integration based on peer-to-peer networking and shared file systems. (4) Visualization schemas to present imaging and clinical data: the large volume of medical data presents a daunting challenge for an efficient visualization paradigm. In this section we review current multimedia visualization methods including temporal modeling, problem-specific data organization, including our problem-centric, context and user-specific visualization interface.

Databases, Factual↗

The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data.

The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.

Journal Article↗

Integration of metabolome data with metabolic networks reveals reporter reactions.

Interpreting quantitative metabolome data is a difficult task owing to the high connectivity in metabolic networks and inherent interdependency between enzymatic regulation, metabolite levels and fluxes. Here we present a hypothesis-driven algorithm for the integration of such data with metabolic network topology. The algorithm thus enables identification of reporter reactions, which are reactions where there are significant coordinated changes in the level of surrounding metabolites following environmental/genetic perturbations. Applicability of the algorithm is demonstrated by using data from Saccharomyces cerevisiae. The algorithm includes preprocessing of a genome-scale yeast model such that the fraction of measured metabolites within the model is enhanced, and thus it is possible to map significant alterations associated with a perturbation even though a small fraction of the complete metabolome is measured. By combining the results with transcriptome data, we further show that it is possible to infer whether the reactions are hierarchically or metabolically regulated. Hereby, the reported approach represents an attempt to map different layers of regulation within metabolic networks through combination of metabolome and transcriptome data.

Algorithms↗

Universal application-specific integrated circuit for bioelectric data acquisition.

Use of highly integrated application specific circuits (ASICs) in bioelectric data acquisition systems promise important new insights into the origin of a large variety of health problems by providing light-weight, low-power, low-cost medical measurement devices that allow long-term studies. They also promise significant cost reduction in medical care, as patients in principle become mobile and do not have to be hospitalized for observation. We report on the development and successful implementation of a universal ASIC, designed to meet key characteristics of a broad variety of bioelectric signals in terms of their dynamic range, sampling rate and input referred noise; e.g. electrocardiogram (ECG), electroencephalogram (EEG) and, most constringently, evoked potentials (EPs). Our approach for the first time makes cost-effective use of state-of-the-art microelectronics in medical measurement equipment, thus offering to replace discrete, single application devices used at present.

Amplifiers, Electronic↗

Integrated system brings hospital data together.

Healthcare industry changes during the 1980s--increased competition and alterations in the Medicare payment methodology--place new and more complex demands on a hospital's information systems, which often fall short of meeting those demands. These systems were designed for financial reporting, billing, or providing clinical data, and few of them are capable of linking with other unrelated systems. Today's hospital manager needs timely and simultaneous access to data from a variety of sources within the hospital. All the elements to accomplish this are collected somewhere in the hospital, but finding them and bringing them together is difficult. The key to the efficient management and use of data bases is in understanding the fundamental concept of relational data bases, which is the capability of linking or joining separate data files through a common data element in each file. In this way, data files may be integrated into a "related" data base. Any number of separate files, or tables, may exist within a "relational" data base as long as a series of threads links them. A strategic management information data base includes the information necessary to analyze, understand, and manage the hospital's markets, products, resources, and profitability. The major components of this information system are the case mix and cost accounting, budgeting, and modeling systems. The case mix and cost accounting factors involve managing concrete pieces of data, whereas the budgeting and modeling factors manipulate data to create a scenario. The strategic management information data base is the foundation of a hospital's decision support system, which is rapidly moving into the category of a necessary tool of the hospital manager's trade.

Computer Systems↗

A CORBA-based object framework with patient identification translation and dynamic linking. Methods for exchanging patient data.

Exchanging and integration of patient data across heterogeneous databases and institutional boundaries offers many problems. We focused on two issues: (1) how to identify identical patients between different systems and institutions while lacking universal patient identifiers; and (2) how to link patient data across heterogeneous databases and institutional boundaries. To solve these problems, we created a patient identification (ID) translation model and a dynamic linking method in the Common Object Request Broker Architecture (CORBA) environment. The algorithm for the patient ID translation is based on patient attribute matching plus computer-based human checking; the method for dynamic linking is temporal mapping. By implementing these methods into computer systems with help of the distributed object computing technology, we built a prototype of a CORBA-based object framework in which the patient ID translation and dynamic linking methods were embedded. Our experiments with a Web-based user interface using the object framework and dynamic linking-through the object framework were successful. These methods are important for exchanging and integrating patient data across heterogeneous databases and institutional boundaries.

Database Management Systems↗

Integrating Genomic and Nongenomic Data to Stratify the Risk of Contralateral Breast Cancer After Radiation Therapy.

PURPOSE: Women treated with radiation therapy (RT) for breast cancer have an increased risk of developing radiation-associated contralateral breast cancer (CBC). Predicting CBC events is challenging because of the complex interplay of genomic, treatment, personal, and clinical factors. This study investigated computational methods that integrate genome-wide single-nucleotide polymorphisms and nongenomic data to develop a risk stratification model for developing CBC in women treated with RT for their first primary breast cancer. METHODS AND MATERIALS: This study used a subset of the population-based Women's Environmental Cancer and Radiation Epidemiology study that included 633 CBC cases and 1253 individually matched unilateral breast cancer controls who were treated with RT and had single-nucleotide polymorphism data available from a genome-wide association study. The study population was split into training, validation, and test sets for rigorous modeling and validation. Three data integration methods were compared in terms of their ability to stratify CBC risk: (1) naive integration; (2) sequential integration; and (3) sequential iterative integration. A biological analysis of the final model was performed using gene set enrichment analysis and protein-protein interaction analysis with gene annotation information informed by the model. RESULTS: The best-performing integration method was the sequential iterative integration equipped with the mixed-effect random forest algorithm. This approach achieved an area under the curve of 0.64 to stratify CBC risk in the test set, representing moderate predictive power. Calibration analysis showed good agreement between the lowest and highest risk bins stratified using sorted predicted values in the test set, resulting in an odds ratio of 3.27 for both predicted and observed CBC occurrence. Gene set enrichment analysis and protein-protein interaction analysis revealed that genes with high importance scores were associated with pathways relevant to lipid and fatty acid metabolism as well as breast cancer sensitivity to tamoxifen. CONCLUSIONS: The mixed-effect random forest approach demonstrated the potential for integrating high-dimensional genomic and low-dimensional nongenomic data to stratify CBC risk.

Humans↗

Reengineering outcomes management: an integrated approach to managing data, systems, and processes.

The integration of outcomes management into organizational reengineering projects is often overlooked or marginalized in proportion to the entire project. Incorporation of an integrated outcomes management program strengthens the overall quality of reengineering projects and enhances their sustainability. This article presents a case study in which data, systems, and processes were reengineered to form an effective Outcomes Management program as a component of the organization's overall project. The authors describe eight steps to develop and monitor an integrated outcomes management program. An example of an integrated report format is included.

Attitude of Health Personnel↗

Application of the correlation integral to respiratory data of infants during REM sleep.

Non-linear time sequence analysis has been performed on infant sleep measurement data in order to obtain more information about the respiratory processes. As a first step, respiration data during REM sleep were analysed with methods from non-linear dynamics, especially, the correlation integral and the slope of its log-log plot, representing the correlation dimension. Before calculation of the correlation integral, a special kind of filtering has to be applied to the data. This filtering algorithm is a state space and singular value decomposition-based noise reduction method, and it is used to separate the noise and signal subspaces. The dynamics of a signal (in our case data from the respiratory process) and its degrees of freedom can be characterised by the correlation integral and by the correlation dimension, respectively. The main result of this study is that the highly irregular-looking breathing patterns during REM sleep could be described by a deterministic system, and finally the physiological significance of this finding is discussed.

Electroencephalography↗

New methods in cross-cultural psychiatry: psychiatric illness in Taiwan and the United States.

OBJECTIVE: Cross-cultural psychiatric research has suffered from many methodological shortcomings. To answer some of these shortcomings, the present study compared rates of psychiatric disorders in Taiwan and the United States by combining data from both countries into a single data set. METHOD: Results from large, community-based surveys in the United States and Taiwan, the National Institute of Mental Health (NIMH) Epidemiologic Catchment Area survey and the Taiwan Psychiatric Epidemiological Project, were combined into a single data set. This integration of the data sets was possible because both surveys used the NIMH Diagnostic Interview Schedule to ascertain cases. The integrated data sets were then analyzed with identical algorithms to generate lifetime prevalence rates of psychiatric disorders according to DSM-III criteria for both the United States and Taiwan. RESULTS: Lifetime prevalence rates of psychiatric illness in Taiwan were generally lower than U.S. rates. The rates of any disorder were 21.56% in Taiwan and 35.55% in the United States (Z = 22.34, p less than 10(-109]. The rates of most specific disorders were lower in Taiwan, and none of the rates was higher in Taiwan. CONCLUSIONS: While a culturally determined response bias may have lowered the rates in Taiwan somewhat, the results appear to be valid. Implications for the future use of structured diagnostic interviews in cross-cultural research are discussed.

Algorithms↗