Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Integrating profiling data: using linear correlation to reveal coregulation of transcript and metabolites.

Recent advances in the medical and biological sciences have been characterized by a major paradigm shift from reductionism to integrated and holistic systems approaches. Such approaches are characterized at the experimental level by the multiparallel analysis of a multitude of parameters of a given biological system at a range of different molecular levels, following the systematic perturbation of the system in question. Although a multitude of studies have been carried out to assess the transcript, protein, and metabolite complements of cells under various conditions, to date, few have been attempted that encompass the profiling of more than one of these entities. In this chapter, we describe combined analysis of data obtained from transcript and metabolic profiling, and detail advantages of using both approaches in parallel.

Gas Chromatography-Mass Spectrometry↗

QIS: A framework for biomedical database federation.

Query Integrator System (QIS) is a database mediator framework intended to address robust data integration from continuously changing heterogeneous data sources in the biosciences. Currently in the advanced prototype stage, it is being used on a production basis to integrate data from neuroscience databases developed for the SenseLab project at Yale University with external neuroscience and genomics databases. The QIS framework uses standard technologies and is intended to be deployable by administrators with a moderate level of technological expertise: It comes with various tools, such as interfaces for the design of distributed queries. The QIS architecture is based on a set of distributed network-based servers, data source servers, integration servers, and ontology servers, that exchange metadata as well as mappings of both metadata and data elements to elements in an ontology. Metadata version difference determination coupled with decomposition of stored queries is used as the basis for partial query recovery when the schema of data sources alters.

Computer Communication Networks↗

Simple and Patlak models for myocardial blood flow measurements with nitrogen-13-ammonia and PET in humans.

UNLABELLED: The Simple and Patlak models for estimating myocardial blood flow with 13N-ammonia have become attractive for clinical applications with PET because of their simplicity and ease of implementation. However, these models are sensitive to factors such as the data acquisition times and data integration times, which can cause errors in the estimation of myocardial blood flow, as demonstrated in this study. Limiting the application of these models to specific conditions can minimize the errors. METHODS: Dynamic PET images of the uptake of 13N-ammonia in the heart were obtained in seven humans under rest and dipyridamole stress. Myocardial blood flow was estimated using the Simple and Patlak models for different data acquisition times and data integration times. Blood flow values were compared to flow values computed with the two-compartment model as a reference. RESULTS: Blood flow values calculated with the Simple and Patlak models during the first 2 min of data acquisition were closely correlated to the two-compartment model values. Longer acquisition times resulted in significant underestimation of blood flow for the Simple model. Long integration times of greater than 60 sec also resulted in significant underestimation of blood flow for both models. CONCLUSION: The Simple and Patlak models produce estimates of myocardial blood flow that are well correlated with the two-compartment model estimated blood flows for the integration time of 60 sec from 60 to 120 sec postinjection. Because of the errors associated with longer data acquisition times and longer integration times, use of these models should be limited to a well-documented data acquisition paradigm.

Ammonia↗

Human Immunodeficiency Virus Reverse Transcriptase and Protease Sequence Database: an expanded data model integrating natural language text and sequence analysis programs.

The HIV Reverse Transcriptase and Protease Sequence Database is an on-line relational database that catalogs evolutionary and drug-related sequence variation in the human immunodeficiency virus (HIV) reverse transcriptase (RT) and protease enzymes, the molecular targets of anti-HIV therapy (http://hivdb.stanford.edu). The database contains a compilation of nearly all published HIV RT and protease sequences, including submissions from International Collaboration databases and sequences published in journal articles. Sequences are linked to data about the source of the sequence sample and the antiretroviral drug treatment history of the individual from whom the isolate was obtained. During the past year 3500 sequences have been added and the data model has been expanded to include drug susceptibility data on sequenced isolates. Database content has also been integrated with didactic text and the output of two sequence analysis programs.

Amino Acid Sequence↗

Health care fraud and abuse data collection program: technical revisions to healthcare integrity and protection data bank data collection activities. Interim final rule with comment period.

The rule makes technical changes to the Healthcare Integrity and Protection Data Bank (HIPDB) data collection reporting requirements set forth in 45 CFR part 61 by clarifying the types of personal numeric identifiers that may be reported to the data bank in connection with adverse actions. Specifically, the rule clarifies that in lieu of a Social Security Number (SSN), an individual taxpayer identification number (ITIN) may be reported to the data bank when, in those limited situations, an individual does not have an SSN.

Data Collection↗

A personalized and automated dbSNP surveillance system.

The development of high throughput techniques and large-scale studies in the biological sciences has given rise to an explosive growth in both the volume and types of data available to researchers. A surveillance system that monitors data repositories and reports changes helps manage the data overload. We developed a dbSNP surveillance system (URL: http://www.pharmgkb.org/do/serve?id=tools.surveillance.dbsnp) that performs surveillance on the dbSNP database and alerts users to new information. The system is notable because it is personalized and fully automated. Each registered user has a list of genes to follow and receives notification of new entries concerning these genes. The system integrates data from dbSNP, LocusLink, PharmGKB, and Genbank to position SNPs on reference sequences and classify SNPs into categories such as synonymous and non-synonymous SNPs. The system uses data warehousing, object model-based data integration, object-oriented programming, and a platform-neutral data access mechanism.

DNA↗

Protein cross-linking analysis using mass spectrometry, isotope-coded cross-linkers, and integrated computational data processing.

Distance constraints in proteins and protein complexes provide invaluable information for calculation of 3D structures, identification of protein binding partners and localization of protein-protein contact sites. We have developed an integrative approach to identify and characterize such sites through the analysis of proteolytic products derived from proteins chemically cross-linked by isotopically coded cross-linkers using LC-MALDI tandem mass spectrometry and computer software. This method is specifically tailored toward the rapid analysis of low microgram amounts of proteins or multimeric protein complexes cross-linked with nonlabeled and deuterium-labeled bis-NHS ester cross-linking reagents (both commercially available and readily synthesized). Through labeling with [18O]water solvent and LC-MALDI analysis, the method further allows the possible distinction between Type 0 and Type 1 or Type 2 modified peptides (monolinks and looplinks or cross-links), although such a distinction is more readily made from analysis of tandem mass spectrometry data. When applied to the bacterial Colicin E7 DNAse/Im7 heterodimeric protein complex, 23 cross-links were identified including six intersubunit cross-links, all between residues that are close in space when examined in the context of the X-ray structure of the heterodimer. In addition, cross-links were successfully identified in five single subunit proteins, beta-lactoglobulin, cytochrome c, lysozyme, myoglobin, and ribonuclease A, establishing the generality of the approach.

Chromatography, Liquid↗

Uniform integration of genome mapping data using intersection graphs.

MOTIVATION: The methods for analyzing overlap data are distinct from those for analyzing probe data, making integration of the two forms awkward. Conversion of overlap data to probe-like data elements would facilitate comparison and uniform integration of overlap data and probe data using software developed for analysis of STS data. RESULTS: We show that overlap data can be effectively converted to probe-like data elements by extracting maximal sets of mutually overlapping clones. We call these sets virtual probes, since each set determines a site in the genome corresponding to the region which is common among the clones of the set. Finding the virtual probes is equivalent to finding the maximal cliques of a graph. We modify a known maximal-clique algorithm such that it finds all virtual probes in a large dataset within minutes. We illustrate the algorithm by converting fingerprint and Alu-PCR overlap data to virtual probes. The virtual probes are then analyzed using double-linkage intersection graphs and structure graphs to show that methods designed for STS data are also applicable to overlap data represented as virtual probes. Next we show that virtual probes can produce a uniform integration of different kinds of mapping data, in particular STS probe data and fingerprint and Alu-PCR overlap data. The integrated virtual probes produce longer double-linkage contigs than STS probes alone, and in conjunction with structure graphs they facilitate the identification and elimination of anomalies. Thus, the virtual-probe technique provides: (i) a new way to examine overlap data; (ii) a basis on which to compare overlap data and probe data using the same systems and standards; and (iii) a unique and useful way to uniformly integrate overlap data with probe data.

Algorithms↗

Integrating voice, data, and paging technologies to enhance information services.

The University of California at San Francisco Medical Center has made a commitment to upgrade its information and telecommunications systems infrastructure. One of the several projects being undertaken by the Medical Center, the Intelligent Console Project, demonstrates how integrating different systems, databases and technologies can improve the quality and accessibility of information, while reducing costs and stream-lining administrative activities. The Intelligent Console acts as an interface mechanism for the several constituent systems and data-bases of the Medical Center and provides a single, front-end control console by which operators can support communications using standardized procedures. Much paperwork has been eliminated and operator training and scheduling streamlined. Equipment consolidation has also freed up space at the Medical Center.

Academic Medical Centers↗

A system-based approach to interpret dose- and time-dependent microarray data: quantitative integration of gene ontology analysis for risk assessment.

Although microarray technology has emerged as a powerful tool to explore expression levels of thousands of genes or even complete genomes after exposure to toxicants, the functional interpretation of microarray data sets still represents a time-consuming and challenging task. Gene ontology (GO) and pathway mapping have both been shown to be powerful approaches to generate a global view of biological processes and cellular components impacted by toxicants. However, current methods only allow for comparisons across two experimental settings at one particular time point. In addition, the resulting annotations are presented in extensive gene lists with minimal or limited quantitative information, data that are crucial in the application of toxicogenomic data for risk assessment. To facilitate quantitative interpretation of dose- or time-dependent genomic data, we propose to use combined average raw gene expression values (e.g., intensity or ratio) of genes associated with specific functional categories derived from the GO database. We developed an extended program (GO-Quant) to extract quantitative gene expression values and to calculate the average intensity or ratio for those significantly altered by functional gene category based on MAPPFinder results. To demonstrate its application, we applied this approach to a previously published dose- and time-dependent toxicogenomic data set (J. F. Dillman et al., 2005, Chem. Res. Toxicol. 18, 28-34). Our results indicate that the above systems approach can describe quantitatively the degree to which functional gene systems change across dose or time. Additionally, this approach provides a robust measurement to illustrate results compared to single-gene assessments and enables the user to calculate the corresponding ED(50) for each specific functional GO term, important for risk assessment.

Animals↗

The iProClass integrated database for protein functional analysis.

Increasingly, scientists have begun to tackle gene functions and other complex regulatory processes by studying organisms at the global scales for various levels of biological organization, ranging from genomes to metabolomes and physiomes. Meanwhile, new bioinformatics methods have been developed for inferring protein function using associative analysis of functional properties to complement the traditional sequence homology-based methods. To fully exploit the value of the high-throughput system biology data and to facilitate protein functional studies requires bioinformatics infrastructures that support both data integration and associative analysis. The iProClass database, designed to serve as a framework for data integration in a distributed networking environment, provides comprehensive descriptions of all proteins, with rich links to over 50 databases of protein family, function, pathway, interaction, modification, structure, genome, ontology, literature, and taxonomy. In particular, the database is organized with PIRSF family classification and maps to other family, function, and structure classification schemes. Coupled with the underlying taxonomic information for complete genomes, the iProClass system (http://pir.georgetown.edu/iproclass/) supports associative studies of protein family, domain, function, and structure. A case study of the phosphoglycerate mutases illustrates a systematic approach for protein family and phylogenetic analysis. Such studies may serve as a basis for further analysis of protein functional evolution, and its relationship to the co-evolution of metabolic pathways, cellular networks, and organisms.

Amino Acid Sequence↗

A standardized data collection tool.

To successfully integrate data collection into the staff's responsibilities, the process must be simple, concise, and easy to use. The data collection tool described in this article includes all the important information at a glance, permits easy comparison with projected and actual thresholds, and analysis of data with follow-up action. It reduces the volumes of paperwork required in many systems. Since the indicators are written on only one paper, there is a reduction in transcription time. Additionally, one paper that contains a sample of twenty should be adequate for a unit-based indicator. Use of this tool reduces the number of papers that must be handled by the QA coordinator as well. Finally, the tool is flexible enough to use in a variety of settings.

Data Collection↗

Information system interoperability in a regional health care system infrastructure: a pilot study using health care information standards.

The 1st and 2nd Regional Health Care System Authority of Central Macedonia (1st and 2nd PeSY) are two of the seventeen Regional Healthcare System Authorities in Greece. Every single PeSY aims to improve the level of quality that health care organisations offer as well as to control the expenditure of health care services provided by the health care organisations, Hospitals and Primary Care Health units. There is currently an urgent need for Regional Health Authorities to deploy integrated healthcare information system, based on secure networks. The limited interoperability of current hospital information systems (HIS) poses a risk for the management of patient related information since there is a difficulty to transform processed data into useful information and knowledge. Thus, a pilot system was developed to achieve data integration record synchronisation using the Health Level 7 protocol between the existing HIS of two Hospitals of Thessaloniki and the central Offices of the PeSY. The pilot was funded by the Third Community Support Framework (jointly funded by EU and Greece) funds in order to prepare the forthcoming major healthcare IT projects in Greece. It is shown that such a system is pragmatic, achieves data integration and provides acceptable integration costs.

Greece↗

Analysis of longitudinal data: the integration of theoretical model, temporal design, and statistical model.

This article argues that ideal longitudinal research is characterized by the seamless integration of three elements: (a) a well-articulated theoretical model of change observed using (b) a temporal design that affords a clear and detailed view of the process, with the resulting data analyzed by means of (c) a statistical model that is an operationalization of the theoretical model. Two general varieties of theoretical models are considered: models in which the time-related change of primary interest is continuous, and those in which it is characterized by movement between discrete states. In addition, two general types of temporal designs are considered: the longitudinal panel design and the intensive longitudinal design. For each general category of theoretical models, some of the analytic possibilities available for longitudinal panel designs and for intensive longitudinal designs are discussed. The article concludes with brief discussions of two issues particularly relevant to longitudinal research--missing data and measurement--and a few words about exploratory research.

Humans↗

VIJB: a companion of the JBROWSE genome browser for the visually impaired people.

MOTIVATION: The availability of touch-sensitive and haptic devices has been a keystone development for the inclusion of visually impaired people (VIPs) in modern, highly digitized work environments. Braille displays have proven efficient and versatile enough to parse large and complex text files, making bioinformatics and text-heavy programming accessible to VIPs. However, the complex graphical objects -combining numerous datasets- typically generated during data integration remain challenging, even with the aid of descriptive AI. This is particularly true in functional genomics. Here, we present VIJB, a simple application that displays the multilayered output of the JBROWSE genome browser on a Braille reader, enabling VIPs to fully participate in data integration in functional genomics. AVAILABILITY AND IMPLEMENTATION: VIJB is programmed in Python and relies on the scientific library NumPy, the braillegraph and pyBigWig libraries, and the TABIX software. The architecture is summarized in Supplementary Material 1, available as supplementary data at Bioinformatics online. VIJB is available for download at the GitHub repository https://GitHub.com/NiBuMNHN/VIJB and is licenced under the GPL 3.0.

Persons with Visual Disabilities↗

DIVAS: an R package for identifying shared and individual variations of multiomics data.

MOTIVATION: Multiomics data integration aims to identify biological patterns shared across molecular modalities. Most existing methods detect either jointly shared variation, across all modalities, or individual variation, unique to a single modality, but overlook partially shared variation, shared by only a subset of modalities. This is a critical limitation, because many biological mechanisms manifest in some but not all molecular modalities. RESULTS: We present an open-source R package implementing data integration via analysis of subspaces (DIVAS), a framework for systematically identifying jointly shared, partially shared and individual variations across multiple data types. DIVAS combines angle-based subspace analysis with inference through rotational bootstrap, hierarchically searching all combinations of modalities to decompose multiomics data into interpretable components with scores and loadings. In simulations with a known sharing structure, DIVAS recovered every component across a wide range of noise levels, whereas existing methods did not. Applied to multi-modal COVID-19 data, it reveals partially shared immune and metabolic dysregulation patterns underpinning disease severity that conventional approaches would miss. AVAILABILITY AND IMPLEMENTATION: DIVAS is available at https://github.com/ByronSyun/DIVAS, with documentation and vignettes. The COVID-19 case study vignette is available at https://byronsyun.github.io/DIVAS_COVID19_CaseStudy/.

Multiomics↗

EpoDB: a database of genes expressed during vertebrate erythropoiesis.

EpoDB is a database designed for the study of gene regulation during differentiation and development of vertebrate red blood cells. In building EpoDB, we have taken the in advance approach to the data integration problem: we have extracted data relevant to red blood cells from GenBank, SWISS-PROT, TRRD (transcriptional regulation data) and GERD (expression levels data) to create a single integrated, highly curated view. Tools have been developed to automate data extraction from online resources, cleanse data of errors, enter information manually from the primary literature, generate a uniform, canonical representation of information and maintain data currency. The database is organized around biological features, e.g., genes, rather than sequences, which are supported by a controlled and consistent vocabulary for gene names and gene family names. Beyond the standard database queries, the functionality of EpoDB includes the ability to extract features and subsequences, display sequences and features graphically using bioWidget viewers and integrated analysis tools. EpoDB may be accessed at: http://cbil.humgen.upenn.edu/epodb/

Animals↗