Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis.

biomaRt is a new Bioconductor package that integrates BioMart data resources with data analysis software in Bioconductor. It can annotate a wide range of gene or gene product identifiers (e.g. Entrez-Gene and Affymetrix probe identifiers) with information such as gene symbol, chromosomal coordinates, Gene Ontology and OMIM annotation. Furthermore biomaRt enables retrieval of genomic sequences and single nucleotide polymorphism information, which can be used in data analysis. Fast and up-to-date data retrieval is possible as the package executes direct SQL queries to the BioMart databases (e.g. Ensembl). The biomaRt package provides a tight integration of large, public or locally installed BioMart databases with data analysis in Bioconductor creating a powerful environment for biological data mining.

Algorithms↗

Giotto Suite: a multiscale and technology-agnostic spatial multiomics analysis ecosystem.

Emerging spatial multiomics technologies provide an increasingly large amount of information content at multiple scales. However, it remains challenging to efficiently represent and harmonize diverse spatial datasets. Here we present Giotto Suite, a suite of modular packages that provides scalable and extensible end-to-end solutions for multiscale and multiomic data analysis, integration and visualization. At its core, Giotto Suite is centered around an innovative data framework, allowing the representation and integration of spatial omics data in a technology-agnostic manner. Giotto Suite integrates molecular, morphology, spatial and annotated feature information to create a responsive and flexible workflow, as demonstrated by applications to several state-of-the-art spatial technologies. Furthermore, Giotto Suite builds upon interoperable interfaces and data structures that bridge the established fields of genomics and spatial data science in R, thereby enabling independent developers to create custom-engineered pipelines. As such, Giotto Suite creates an immersive and multiscale ecosystem for spatial multiomic data analysis.

Genomics↗

Chemical effects in biological systems--data dictionary (CEBS-DD): a compendium of terms for the capture and integration of biological study design description, conventional phenotypes, and 'omics data.

A critical component in the design of the Chemical Effects in Biological Systems (CEBS) Knowledgebase is a strategy to capture toxicogenomics study protocols and the toxicity endpoint data (clinical pathology and histopathology). A Study is generally an experiment carried out during a period of time for the purpose of obtaining data, and the Study Design Description captures the methods, timing, and organization of the Study. The CEBS Data Dictionary (CEBS-DD) has been designed to define and organize terms in an attempt to standardize nomenclature needed to describe a toxicogenomics Study in a structured yet intuitive format and provide a flexible means to describe a Study as conceptualized by the investigator. The CEBS-DD will organize and annotate information from a variety of sources, thereby facilitating the capture and display of toxicogenomics data in biological context in CEBS, i.e., associating molecular events detected in highly-parallel data with the toxicology/pathology phenotype as observed in the individual Study Subjects and linked to the experimental treatments. The CEBS-DD has been developed with a focus on acute toxicity studies, but with a design that will permit it to be extended to other areas of toxicology and biology with the addition of domain-specific terms. To illustrate the utility of the CEBS-DD, we present an example of integrating data from two proteomics and transcriptomics studies of the response to acute acetaminophen toxicity (A. N. Heinloth et al., 2004, Toxicol. Sci. 80, 193-202).

Acetaminophen↗

VA develops integrated text and image data system.

A standard PC workstation allows physicians to review a patient's entire record--both text and images--with a new integrated HIS now installed at the Department of Veterans Affairs Medical Center in Washington, D.C. The system uses high-resolution video cameras and the latest in fiberoptic technology.

Data Display↗

Challenges for biomedical informatics and pharmacogenomics.

Pharmacogenomics requires the integration and analysis of genomic, molecular, cellular, and clinical data, and it thus offers a remarkable set of challenges to biomedical informatics. These include infrastructural challenges such as the creation of data models and databases for storing these data, the integration of these data with external databases, the extraction of information from natural language text, and the protection of databases with sensitive information. There are also scientific challenges in creating tools to support gene expression analysis, three-dimensional structural analysis, and comparative genomic analysis. In this review, we summarize the current uses of informatics within pharmacogenomics and show how the technical challenges that remain for biomedical informatics are typical of those that will be confronted in the postgenomic era.

Communication↗

Evaluation of descending aortic flow volumes and effective orifice area through aortic coarctation by spatiotemporal integration of color Doppler data: An in vitro study.

Flow volumes in an in vitro model of the aorta with 3 different degrees of stiffness (stiff, moderately stiff, and compliant) proximal to a coarctation were calculated by using a digital color Doppler echocardiography flow calculation method that semiautomatically integrates spatial and temporal color flow velocity data. These flow volumes were compared with those obtained by the conventional pulsed Doppler method with reference to ultrasonic flowmeter. Flow volumes determined by the automated method agreed well with those obtained by ultrasonic flowmeter, even in this compliant aorta model with vessel size changing with pulsation, whereas the pulsed Doppler method overestimated the reference data, especially for more compliant descending aortic segments. The combination of flow data with continuous wave Doppler allows definition of effective orifice area for coarctation.

Aorta↗

Discovery informatics: its evolving role in drug discovery.

Drug discovery and development is a highly complex process requiring the generation of very large amounts of data and information. Currently this is a largely unmet informatics challenge. The current approaches to building information and knowledge from large amounts of data has been addressed in cases where the types of data are largely homogeneous or at the very least well-defined. However, we are on the verge of an exciting new era of drug discovery informatics in which methods and approaches dealing with creating knowledge from information and information from data are undergoing a paradigm shift. The needs of this industry are clear: Large amounts of data are generated using a variety of innovative technologies and the limiting step is accessing, searching and integrating this data. Moreover, the tendency is to move crucial development decisions earlier in the discovery process. It is crucial to address these issues with all of the data at hand, not only from current projects but also from previous attempts at drug development. What is the future of drug discovery informatics? Inevitably, the integration of heterogeneous, distributed data are required. Mining and integration of domain specific information such as chemical and genomic data will continue to develop. Management and searching of textual, graphical and undefined data that are currently difficult, will become an integral part of data searching and an essential component of building information- and knowledge-bases.

Artificial Intelligence↗

PlasmoDB: the Plasmodium genome resource. A database integrating experimental and computational data.

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates the recently completed P. falciparum genome sequence and annotation, as well as draft sequence and annotation emerging from other Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for intra- and inter-species comparisons. Sequence information is integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects and proteomics studies. The relational schema used to build PlasmoDB, GUS (Genomics Unified Schema) employs a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically-based, queries of the database. A stand-alone version of the database is also available on CD-ROM (P. falciparum GenePlot), facilitating access to the data in situations where internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to facilitate utilization of the vast quantities of genomic-scale data produced by the global malaria research community. The software used to develop PlasmoDB has been used to create a second Apicomplexan parasite genome database, ToxoDB (http://ToxoDB.org).

Animals↗

Causal Mediation Analysis for Integrating Exposure, Genomic, and Phenotype Data.

Causal mediation analysis provides an attractive framework for integrating diverse types of exposure, genomic, and phenotype data. Recently, this field has seen a surge of interest, largely driven by the increasing need for causal mediation analyses in health and social sciences. This article aims to provide a review of recent developments in mediation analysis, encompassing mediation analysis of a single mediator and a large number of mediators, as well as mediation analysis with multiple exposures and mediators. Our review focuses on the recent advancements in statistical inference for causal mediation analysis, especially in the context of high-dimensional mediation analysis. We delve into the complexities of testing mediation effects, especially addressing the challenge of testing a large number of composite null hypotheses. Through extensive simulation studies, we compare the existing methods across a range of scenarios. We also include an analysis of data from the Normative Aging Study, which examines DNA methylation CpG sites as potential mediators of the effect of smoking status on lung function. We discuss the pros and cons of these methods and future research directions.

causal inference↗

High-resolution 3-D shape integration of dentition and face measured by new laser scanner.

Face and dentition were measured using a high-resolution three-dimensional laser scanner to circumvent problems of radiation exposure and metal-streak artifacts associated with X-ray computed tomography. The resulting range data were integrated in order to visualize the dentition relative to the face. The acquisition interval for dentition by laser scanner was 0.18 mm, and complicated morphologies of the occlusal surface could be sufficiently reproduced. Reproduction of occlusal condition of upper and lower dentitions was conducted by matching the surface of the occlusal impression record with upper dentition data. To integrate dentition and face, a marker plate interface was devised and adopted on the lower dental cast or by the subject directly. Integration was performed by matching both sets of interface data. Reproduction of the occlusal condition and integration of the dentition and face were accomplished and visualized satisfactorily by computer graphics. The integration accuracy was examined by changing the attachment angle of the marker plate, and the marker plate attached at 45 degrees showed the smallest error of 0.2 mm. The current noninvasive method is applicable to clinical examination, diagnosis and explanation to the patient when dealing with the physical relationship between face and dentition.

Adult↗

The Gaggle: an open-source software system for integrating bioinformatics software and data sources.

BACKGROUND: Systems biologists work with many kinds of data, from many different sources, using a variety of software tools. Each of these tools typically excels at one type of analysis, such as of microarrays, of metabolic networks and of predicted protein structure. A crucial challenge is to combine the capabilities of these (and other forthcoming) data resources and tools to create a data exploration and analysis environment that does justice to the variety and complexity of systems biology data sets. A solution to this problem should recognize that data types, formats and software in this high throughput age of biology are constantly changing. RESULTS: In this paper we describe the Gaggle -a simple, open-source Java software environment that helps to solve the problem of software and database integration. Guided by the classic software engineering strategy of separation of concerns and a policy of semantic flexibility, it integrates existing popular programs and web resources into a user-friendly, easily-extended environment. We demonstrate that four simple data types (names, matrices, networks, and associative arrays) are sufficient to bring together diverse databases and software. We highlight some capabilities of the Gaggle with an exploration of Helicobacter pylori pathogenesis genes, in which we identify a putative ricin-like protein -a discovery made possible by simultaneous data exploration using a wide range of publicly available data and a variety of popular bioinformatics software tools. CONCLUSION: We have integrated diverse databases (for example, KEGG, BioCyc, String) and software (Cytoscape, DataMatrixViewer, R statistical environment, and TIGR Microarray Expression Viewer). Through this loose coupling of diverse software and databases the Gaggle enables simultaneous exploration of experimental data (mRNA and protein abundance, protein-protein and protein-DNA interactions), functional associations (operon, chromosomal proximity, phylogenetic pattern), metabolic pathways (KEGG) and Pubmed abstracts (STRING web resource), creating an exploratory environment useful to 'web browser and spreadsheet biologists', to statistically savvy computational biologists, and those in between. The Gaggle uses Java RMI and Java Web Start technologies and can be found at http://gaggle.systemsbiology.net.

Computational Biology↗

Integrating baseline health status data collection into the process of care.

BACKGROUND: Health status data are an increasingly important component of outcomes assessment and can be used to facilitate quality assessment and improvement efforts. An enormous challenge to the use of health status data among hospitalized patients, however, is collecting baseline data at the time of treatment, an essential component for risk-adjusting subsequent outcomes. The Mid America Heart Institute of Saint Luke's Hospital (Kansas City, Mo), attempted to integrate the collection of health status assessments within the process of performing coronary revascularization. THE DATA COLLECTION STRATEGY: The data collection strategy was developed for each admission portalelective outpatients (admissions for same-day procedures), inpatients, and emergent cases. Health status data were collected on all patients with coronary artery disease who were receiving a percutaneous coronary intervention or coronary artery bypass graft with no disruption to physician scheduling or nursing staff. RESULTS: In general, patients were agreeable to completing the health status survey. Despite initial efforts to educate the hospital staff about the goal and purpose of health status assessment, staff members who were unaware of the uses of these data seemed to minimize their value. Providing examples of how to use these data relative to the staff member's specific occupational role facilitated buy-in for this project. EPILOGUE: After the pilot study, which lasted until June 1999, data were continually collected for 18 months, through August 2000, even with the cessation of external grant funding for this project. Baseline data collection finally stopped, primarily because of a failure to accommodate data collection into the routine flow of patient care by existing nursing staff.

Angioplasty, Balloon, Coronary↗

OmicBrowse: a browser of multidimensional omics annotations.

UNLABELLED: OmicBrowse is a browser to explore multiple datasets coordinated in the multidimensional omic space integrating omics knowledge ranging from genomes to phenomes and connecting evolutional correspondences among multiple species. OmicBrowse integrates multiple data servers into a single omic space through secure peer-to-peer server communications, so that a user can easily obtain an integrated view of distributed data servers, e.g. an integrated view of numerous whole-genome tiling-array data retrieved from a user's in-house private-data server, along with various genomic annotations from public internet servers. OmicBrowse is especially appropriate for positional-cloning purposes. It displays both genetic maps and genomic annotations within wide chromosomal intervals and assists a user to select candidate genes by filtering their annotations or associated documents against user-specified keywords or ontology terms. We also show that an omic-space chart effectively represents schemes for integrating multiple datasets of multiple species. AVAILABILITY: OmicBrowse is developed by the Genome-Phenome Superbrain Project and is released as free open-source software under the GNU General Public License at http://omicspace.riken.jp.

Chromosome Mapping↗

Computerized case history--an effective tool for management of patients and clinical trials.

UNLABELLED: Monitoring diagnostic procedures, treatment protocols and clinical outcome are key issues in maintaining quality medical care and in evaluating clinical trials. For these purposes, a user-friendly computerized method for monitoring all available information about a patient is needed. OBJECTIVE: To develop a real-time computerized data collection system for verification, analysis and storage of clinical information on an individual patient. METHODS: Data was integrated on a single time axis with normalized graphics. Laboratory data was set according to standard protocols selected by the user and diagnostic images were integrated as needed. The system automatically detects variables that fall outside established limits and violations of protocols, and generates alarm signals.Results. The system provided an effective tool for detection of medical errors, identification of discrepancies between therapeutic and diagnostic procedures, and protocol requirements. CONCLUSIONS: The computerized case history system allows collection of medical information from multiple sources and builds an integrated presentation of clinical data for analysis of clinical trials and for patient follow-up.

Journal Article↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources.

MOTIVATION: Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. RESULTS: We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is "task agnostic", in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer's disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. AVAILABILITY AND IMPLEMENTATION: miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Humans↗

[Mapping of EST- and STS-markers in the human genome using a panel fo radiation hybrids].

Ten DNA markers were localized in the human genome by a screening procedure against the radiation hybrid somatic cell panel (GeneBridge 4 RH Panel) using polymerase chain reaction (RH mapping method). DNA markers were developed to nucleotide sequences adjacent to NotI sites of human chromosome 3 (NotI-STS markers) and also to nucleotide sequences of human cDNA (EST markers). Three EST markers mapped (B10164, S16R and 18F5R) were localized in the human genome for the first time. Marker B10164 was found to be homologous to the nucleotide sequence of the BASP1 gene coding a major receptor protein. Markers S16R and 18F5R presumably tagged new genes, because no homologies were revealed among the nucleotide sequences presented in the databases. For four NotI-STS, more precise localization on human chromosome 3 was determined. On the basis of the data obtained, the NotI map may be integrated with other types of physical maps of human chromosome 3. RH mapping with a standard commercial panel of radiation hybrid somatic cells provided a chance to integrate the data obtained into international databases and existing integrated human chromosomal maps.

Base Sequence↗

New data and tools for integrating discrete and continuous population modeling strategies.

Realistic population models have interactions between individuals. Such interactions cause populations to behave as systems with nonlinear dynamics. Much population data analysis is done using linear models assuming no interactions between individuals. Such analyses miss strong influences on population behavior and can lead to serious errors--especially for infectious diseases. To promote more effective population system analyses, we present a flexible and intuitive modeling framework for infection transmission systems. This framework will help population scientists gain insight into population dynamics, develop theory about population processes, better analyze and interpret population data, design more powerful and informative studies, and better inform policy decisions. Our framework uses a hierarchy of infection transmission system models. Four levels are presented here: deterministic compartmental models using ordinary differential equations (DE); stochastic compartmental (SC) models that relax assumptions about population size and include stochastic effects; individual event history models (IEH) that relax the SC compartmental structure assumptions by allowing each individual to be unique. IEH models also track each individual's history, and thus, allow the simulation of field studies. Finally, dynamic network (DNW) models relax the assumption of the previous models that contacts between individuals are instantaneous events that do not affect subsequent contacts. Eventually it should be possible to transit between these model forms at the click of a mouse. An example is presented dealing with Cryptosporidium. It illustrates how transiting model forms helps assess water contamination effects, evaluate control options, and design studies of infection transmission systems using nucleotide sequences of infectious agents.

Data Interpretation, Statistical↗

Systems integration: requirements for a fully functioning electronic radiology department.

Disparate computer-based information systems such as hospital information systems (HIS), radiology information systems (RIS), and picture archiving and communication systems (PACS) have been introduced into radiology departments at various times to meet specific operational objectives. Typically, these systems are implemented without an integration strategy. Systems integration, which optimizes integrity of data and labor savings, can be achieved by two general approaches. The first links the HIS to the PACS; the second involves interlinking of the HIS, RIS, and PACS, with the RIS as the central controlling system. Standardization in hardware, operating systems, and data base formats--which will allow true integration--is being addressed nationally and worldwide. Operational issues to resolve include ways to increase network capacity, control of data flow, and strategies for dealing with downtime. In the future, systems integration will enable prefetching, two-way interfaces, interfaces with digital dictation systems, and improved linkages with external digital input devices.

Computer Communication Networks↗