Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

[The principles of data formalization for building the genetic-morphological model of shoot development in flowering plants].

Shoot system of a plant can be divided into elementary molecules composed of phyllome, internode, and meristem of the lateral bud. The capacity of plants for open growth is manifested as multiple reproductions of the modules. These main principles of plant structural organization can be used to formalize and integrate the data from various disciplines studying the shoot development--genetics of development, morphology, etc. At the example of model species Arabidopsis thaliana we show that the data on genetic control of shoot development can be considered in terms of individual modules reorganization. Several variants of the modules structural reorganization are allowed: reduction or transformation of phyllome, change in the internode length, and triggering active/inactive status of the lateral shoot meristem. Each variant of the module structure corresponds to specific pattern of genes activity. Such integration of the data on genetic and structural aspects of morphogenesis can form a basis for mathematical modeling of plant development.

Arabidopsis↗

Jointly analyzing gene expression and copy number data in breast cancer using data reduction models.

With the growing surge of biological measurements, the problem of integrating and analyzing different types of genomic measurements has become an immediate challenge for elucidating events at the molecular level. In order to address the problem of integrating different data types, we present a framework that locates variation patterns in two biological inputs based on the generalized singular value decomposition (GSVD). In this work, we jointly examine gene expression and copy number data and iteratively project the data on different decomposition directions defined by the projection angle theta in the GSVD. With the proper choice of theta, we locate similar and dissimilar patterns of variation between both data types. We discuss the properties of our algorithm using simulated data and conduct a case study with biologically verified results. Ultimately, we demonstrate the efficacy of our method on two genome-wide breast cancer studies to identify genes with large variation in expression and copy number across numerous cell line and tumor samples. Our method identifies genes that are statistically significant in both input measurements. The proposed method is useful for a wide variety of joint copy number and expression-based studies. Supplementary information is available online, including software implementations and experimental data.

Biomarkers, Tumor↗

Resources and tools for investigating biomolecular networks in mammals.

Molecular databases serve as primary information resources for the analysis of biological networks providing an essential and invaluable treasure for information exploration. Tools for projecting experimental data sets onto known functional information are a major need to support the analysis of samples produced in clinical research. A new concept is the notation of functional modules, i.e. the characterisation of sets of proteins that perform a defined biological function in cooperation. The determination and analysis of functional modules overcome the limitations of the analysis of individual genes and their properties. Although functional modules are not suitable to fully capture systems properties, they have the potential to unify the information generated by different types of experiments. We describe advances related to the problem of integrating heterogeneous data sets into functional modules for mouse and/or human cellular networks based on publicly available data resources, including advances in the design of ontologies for functional classification, problems of automatic protein functional annotation and integration of microarray data.

Algorithms↗

Challenges in data management for functional genomics.

Biological databases face challenges in four main areas: (1). integration, interoperation and federation; (2). ontologies and definitions of semantics; (3). community annotation; and (4). integration of data analysis tools with databases. Each of these areas provides interesting targets for research and development.

Computational Biology↗

Revealing modularity and organization in the yeast molecular network by integrated analysis of highly heterogeneous genomewide data.

The dissection of complex biological systems is a challenging task, made difficult by the size of the underlying molecular network and the heterogeneous nature of the control mechanisms involved. Novel high-throughput techniques are generating massive data sets on various aspects of such systems. Here, we perform analysis of a highly diverse collection of genomewide data sets, including gene expression, protein interactions, growth phenotype data, and transcription factor binding, to reveal the modular organization of the yeast system. By integrating experimental data of heterogeneous sources and types, we are able to perform analysis on a much broader scope than previous studies. At the core of our methodology is the ability to identify modules, namely, groups of genes with statistically significant correlated behavior across diverse data sources. Numerous biological processes are revealed through these modules, which also obey global hierarchical organization. We use the identified modules to study the yeast transcriptional network and predict the function of >800 uncharacterized genes. Our analysis framework, SAMBA (Statistical-Algorithmic Method for Bicluster Analysis), enables the processing of current and future sources of biological information and is readily extendable to experimental techniques and higher organisms.

Amino Acids↗

[Treatment of depression. Presentation of an integrated therapeutic model].

In clinical work, research data must be integrated in the daily handling of patients. At the same time, symptoms must be treated and respect paid to the different personalities of the patients. Based on clinical experience and research data I present an integrated model for the treatment of patients with moderate to severe depression. I used the main principles in the model over 4 to 5 years, and where research has pointed to new, clinically relevant findings, these have also been included. The model can be used in psychiatric out-patient-clinics, on emergency wards and for general practitioners with a special interest in psychiatry. After 5 to 8 sessions spread over 8 to 12 weeks, the condition of most patients will have stabilized and improved, and the psychological part of the treatment can than be terminated.

Depression↗

Monitoring, modelling and environmental exposure assessment of industrial chemicals in the aquatic environment.

Monitoring and laboratory data play integral roles alongside fate and exposure models in comprehensive risk assessments. The principle in the European Union Technical Guidance Documents for risk assessment is that measured data may take precedence over model results but only after they are judged to be of adequate reliability and to be representative of the particular environmental compartments to which they are applied. In practice, laboratory and field data are used to provide parameters for the models, while monitoring data are used to validate the models' predictions. Thus, comprehensive risk assessments require the integration of laboratory and monitoring data with the model predictions. However, this interplay is often overlooked. Discrepancies between the results of models and monitoring should be investigated in terms of the representativeness of both. Certainly, in the context of the EU risk assessment of existing chemicals, the specific requirements for monitoring data have not been adequately addressed. The resources required for environmental monitoring, both in terms of manpower and equipment, can be very significant. The design of monitoring programmes to optimise the use of resources and the use of models as a cost-effective alternative are increasing in importance. Generic considerations and criteria for the design of new monitoring programmes to generate representative quality data for the aquatic compartment are outlined and the criteria for the use of existing data are discussed. In particular, there is a need to improve the accessibility to data sets, to standardise the data sets, to promote communication and harmonisation of programmes and to incorporate the flexibility to change monitoring protocols to amend the chemicals under investigation in line with changing needs and priorities.

Environmental Exposure↗

A framework of integrating gene relations from heterogeneous data sources: an experiment on Arabidopsis thaliana.

One of the most important goals of biological investigation is to uncover gene functional relations. In this study we propose a framework for extraction and integration of gene functional relations from diverse biological data sources, including gene expression data, biological literature and genomic sequence information. We introduce a two-layered Bayesian network approach to integrate relations from multiple sources into a genome-wide functional network. An experimental study was conducted on a test-bed of Arabidopsis thaliana. Evaluation of the integrated network demonstrated that relation integration could improve the reliability of relations by combining evidence from different data sources. Domain expert judgments on the gene functional clusters in the network confirmed the validity of our approach for relation integration and network inference.

Arabidopsis↗

Neurolitigation: a perspective on the elements of expert testimony for extending the Daubert challenge.

Scientific expert witness testimony has the potential for affecting most court decisions in civil and criminal proceedings. Since experts were first utilized in English courts beginning in the 14th century, most contemporary courts struggle with seeking a balance between plaintiff and defense counsel allowing each party its day in court while taking into account the work which other courts have done previously in determining the admissibility of expert witness testimony. When these challenges present themselves in the courtroom, often other courts have approached these identical issues, many in proceedings involving the same expert(s). Confronted with these challenges, trial judges want to understand whether a new Daubert hearing must be held, deal with the issue from a clean slate approach or whether they must reinvent the proverbial wheel. Given these dilemmas, this exposition is based within a heuristic approach that will focus on the consideration of comprehensive data inclusion from an evidentiary foundation as it applies to expert witness testimony admissibility in neurolitigation. While the evidential force of FRE 702 specifically applies to admissibility of scientific evidence, it makes sense that along with scientific, objective data, inclusion of non-medical and other data in forming and admitting expert opinions, have mutual bearing upon the validity of opinions arrived at through neuropsychological assessment. It is these multi-data that should be factored into account when applying the Federal Rule of Evidence 702 scientific admissibility standard. Data from other relevant sources is just as vital as data obtained from objective measures, and co-exists with objective data. Without the integration of this information into resulting diagnostic data and opinions, one's methodology is open to scrutiny and can willfully be characterized as engaging in "junk science". Specific, pragmatic issues are discussed in order to avoid the plausible "junk science" question and to ultimately arrive at a factual and evidenced-based admissibility and reliability determination for the courts. Given the current standard, this article proposes an inclusionary method in neurolitigation as it would necessarily apply to Federal Rule of Evidence 702 which would extend to the integration of data outside medical and scientific information bases to establish accurate opinions for the trier of fact. In so doing, neuropsychological test data, non-medical data and expert testimony would be strengthened through inter-data consistency.

Activities of Daily Living↗

[caCORE: core architecture of bioinformation on cancer research in America].

A critical factor in the advancement of biomedical research is the ease with which data can be integrated, redistributed and analyzed both within and across domains. This paper summarizes the Biomedical Information Core Infrastructure built by National Cancer Institute Center for Bioinformatics in America (NCICB). The main product from the Core Infrastructure is caCORE--cancer Common Ontologic Reference Environment, which is the infrastructure backbone supporting data management and application development at NCICB. The paper explains the structure and function of caCORE: (1) Enterprise Vocabulary Services (EVS). They provide controlled vocabulary, dictionary and thesaurus services, and EVS produces the NCI Thesaurus and the NCI Metathesaurus; (2) The Cancer Data Standards Repository (caDSR). It provides a metadata registry for common data elements. (3) Cancer Bioinformatics Infrastructure Objects (caBIO). They provide Java, Simple Object Access Protocol and HTTP-XML application programming interfaces. The vision for caCORE is to provide a common data management framework that will support the consistency, clarity, and comparability of biomedical research data and information. In addition to providing facilities for data management and redistribution, caCORE helps solve problems of data integration. All NCICB-developed caCORE components are distributed under open-source licenses that support unrestricted usage by both non-profit and commercial entities, and caCORE has laid the foundation for a number of scientific and clinical applications. Based on it, the paper expounds caCORE-base applications simply in several NCI projects, of which one is CMAP (Cancer Molecular Analysis Project), and the other is caBIG (Cancer Biomedical Informatics Grid). In the end, the paper also gives good prospects of caCORE, and while caCORE was born out of the needs of the cancer research community, it is intended to serve as a general resource. Cancer research has historically contributed to many areas beyond tumor biology. At the same time, the paper makes some suggestions about the study at the present time on biomedical informatics in China.

Computational Biology↗

The integration of empirically derived personality assessment data into a behavioral conceptualization and treatment plan. Rationale, guidelines, and caveats.

This article suggests that the appropriate integration of personality (trait) data with information gleaned from a traditional behavioral interview will enhance client-treatment matching. The achievement of this goal is predicated on the ability of behaviorists to challenge tightly held (mis)conceptions of the process of "personality" assessment and to apply empirical criteria when engaged in this integrative endeavor. To facilitate such integration, a rationale for the use of dispositionally based assessment data within a behavioral framework is presented. Guidelines are provided for the accurate measurement and integration of such data in addition to a review of caveats that should be considered if such an enterprise is ultimately to reach fruition.

Adaptation, Psychological↗

Case study: data management strategies in an integrated pathway tool.

This paper describes the development strategies for an integrated tool to support scientists in the creative exploration of data relating to biochemical pathways. The multiple user groups, diverse functionalities, and many types and sources of data demanded a flexible yet coherent approach. This paper summarises the software requirements and the implied modules and functions, and focuses on the design decisions relevant to the representation, management and flow of data. Finally, several case studies in the use of the software are described and evaluated, and recommendations are made for future work.

Animals↗

JXP4BIGI: a generalized, Java XML-based approach for biological information gathering and integration.

MOTIVATION: In the post-genomic era, biologists interested in systems biology often need to import data from public databases and construct their own system-specific or subject-oriented databases to support their complex analysis and knowledge discovery. To facilitate the analysis and data processing, customized and centralized databases are often created by extracting and integrating heterogeneous data retrieved from public databases. A generalized methodology for accessing, extracting, transforming and integrating the heterogeneous data is needed. RESULTS: This paper presents a new data integration approach named JXP4BIGI (Java XML Page for Biological Information Gathering and Integration). The approach provides a system-independent framework, which generalizes and streamlines the steps of accessing, extracting, transforming and integrating the data retrieved from heterogeneous data sources to build a customized data warehouse. It allows the data integrator of a biological database to define the desired bio-entities in XML templates (or Java XML pages), and use embedded extended SQL statements to extract structured, semi-structured and unstructured data from public databases. By running the templates in the JXP4BIGI framework and using a number of generalized wrappers, the required data from public databases can be efficiently extracted and integrated to construct the bio-entities in the XML format without having to hard-code the extraction logics for different data sources. The constructed XML bio-entities can then be imported into either a relational database system or a native XML database system to build a biological data warehouse. AVAILABILITY: JXP4BIGI has been integrated and tested in conjunction with the IKBAR system (http://www.ikbar.org/) in two integration efforts to collect and integrate data for about 200 human genes related to cell death from HUGO, Ensembl, and SWISS-PROT (Bairoch and Apweiler, 2000), and about 700 Drosophila genes from FlyBase (FlyBase Consortium, 2002). The integrated data has been used in comparative genomic analysis of x-ray induced cell death. Also, as explained later, JXP4BIGI is a middleware and framework to be integrated with biological database applications, and cannot run as a stand-alone software for end users. For demonstration purposes, a demonstration version is accessible at (http://www.ikbar.org/jxp4bigi/demo.html).

Database Management Systems↗

Is HRCT the best way to diagnose idiopathic interstitial fibrosis?

PURPOSE OF REVIEW: High-resolution computed tomography (HRCT) has been the major advance in the diagnosis of idiopathic interstitial pneumonias in the last two decades. In diffuse lung diseases, HRCT now has a central role in routine diagnostic evaluation, and a major impact on the utility of other diagnostic tests, especially bronchoalveolar lavage and surgical lung biopsy. RECENT FINDINGS: Numerous published studies have evaluated the accuracy of HRCT. The clinical information was not always utilized to generate a noninvasive diagnosis, however. Despite failure to identify idiopathic pulmonary fibrosis on HRCT in a significant minority of cases, given compatible clinical data, characteristic HRCT appearances justify noninvasive diagnosis in most patients. The limitations of the published studies highlight importance of integrating HRCT data with baseline clinical information and, in selected cases, histopathologic findings. SUMMARY: When HRCT and clinical findings are both typical of an individual diffuse lung disease, i.e. 'pathognomonic', it is generally appropriate to institute management based on a confident noninvasive diagnosis. When clinical and HRCT data are divergent, or when HRCT features are 'indeterminate', however, histologic evaluation continues to play an essential role. Integration of histology with radiologic and clinical data is the best way to formulate the final diagnosis in these cases.

Humans↗

An integrated surgical suite management information system.

The operational aspects, application areas, and results achieved from an integrated surgical suite management information system are described. The system, which has been operating within Henry Ford Hospital in Detroit, Michigan, for 4 years, captures comprehensive data for each surgical episode, performs extensive edits on these data to assure data base integrity, and utilizes this data base in multiple applications. These applications include fixed-format reporting for medical staff and management; ad hoc retrieval capabilities to support research, education, and decision making; and linkage to other hospital systems to reduce both data redundancy and paper flow.

Hospital Departments↗

Complex trait analysis of gene expression uncovers polygenic and pleiotropic networks that modulate nervous system function.

Patterns of gene expression in the central nervous system are highly variable and heritable. This genetic variation among normal individuals leads to considerable structural, functional and behavioral differences. We devised a general approach to dissect genetic networks systematically across biological scale, from base pairs to behavior, using a reference population of recombinant inbred strains. We profiled gene expression using Affymetrix oligonucleotide arrays in the BXD recombinant inbred strains, for which we have extensive SNP and haplotype data. We integrated a complementary database comprising 25 years of legacy phenotypic data on these strains. Covariance among gene expression and pharmacological and behavioral traits is often highly significant, corroborates known functional relations and is often generated by common quantitative trait loci. We found that a small number of major-effect quantitative trait loci jointly modulated large sets of transcripts and classical neural phenotypes in patterns specific to each tissue. We developed new analytic and graph theoretical approaches to study shared genetic modulation of networks of traits using gene sets involved in neural synapse function as an example. We built these tools into an open web resource called WebQTL that can be used to test a broad array of hypotheses.

Animals↗

Developing qualitative databases for multiple users.

Conventionally, anthropological data are collected and analyzed by individuals, and although researchers may use data managers to organize their information, there is little need to classify and code systems to be accessible to others. Recently, however, qualitative and quantitative data have been collected in projects with multiple researchers. Difficulties with the establishment, verification, and management of databases for multiple users, particularly in longitudinal studies, are considerable if the rules underlying coding schemes are difficult to identify or if the documentation is cumbersome. Drawing on the authors' experiences in Australia, the use of computer packages for data management is discussed, and the importance of preserving the integrity of data and maintaining context while facilitating its continued and varied use is emphasized.

Australia↗