Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Metabolic factors in the control of energy stores.

We have examined some of the factors involved in the control of food intake and integrated the data using methods of systems analysis from biomedical engineering. The central hypothesis is that energy stored in the body is a regulated variable. Alterations in the quantity of stored calories initiates changes designed to restore the store of calories to its original level. These responses are both short-term and long-term in nature. They involve integrating data on the quantities of protein, carbohydrate, and lipid stores in the body, probably through such feedback signals as amino acids, glucose, glycerol, and free fatty acids.

Adult↗

Integrated transcriptome and proteome data: the challenges ahead.

The recent availability of platform technologies for high throughput proteome analysis has led to the emergence of integrated messenger RNA and protein expression data. The Pearson correlation coefficients for these data range from 0.46 to 0.76. In these integrated studies, serial analyses of gene expression and DNA microarrays have been used to quantify the transcriptome, while proteome analysis has been based on two-dimensional gel electrophoresis, isotope coded affinity tags and multidimensional protein identification technology. This paper provides a comprehensive review of the analytical techniques used in these studies and explores the extent to which the choice of experimental methodology can bias the correlation or the ability to detect proteins.

Animals↗

A bioinformatics perspective on proteomics: data storage, analysis, and integration.

The field of proteomics is advancing rapidly as a result of powerful new technologies and proteomics experiments yield a vast and increasing amount of information. Data regarding protein occurrence, abundance, identity, sequence, structure, properties, and interactions need to be stored. Currently, a common standard has not yet been established and open access to results is needed for further development of robust analysis algorithms. Databases for proteomics will evolve from pure storage into knowledge resources, providing a repository for information (meta-data) which is mainly not stored in simple flat files. This review will shed light on recent steps towards the generation of a common standard in proteomics data storage and integration, but is not meant to be a comprehensive overview of all available databases and tools in the proteomics community.

Computational Biology↗

Integrating clinical and laboratory data in genetic studies of complex phenotypes: a network-based data management system.

The identification of genes underlying a complex phenotype can be a massive undertaking, and may require a much larger sample size than thought previously. The integration of such large volumes of clinical and laboratory data has become a major challenge. In this paper we describe a network-based data management system designed to address this challenge. Our system offers several advantages. Since the system uses commercial software, it obviates the acquisition, installation, and debugging of privately-available software, and is fully compatible with Windows and other commercial software. The system uses relational database architecture, which offers exceptional flexibility, facilitates complex data queries, and expedites extensive data quality control. The system is particularly designed to integrate clinical and laboratory data efficiently, producing summary reports, pedigrees, and exported files containing both phenotype and genotype data in a virtually unlimited range of formats. We describe a comprehensive system that manages clinical, DNA, cell line, and genotype data, but since the system is modular, researchers can set up only those elements which they need immediately, expanding later as needed.

Clinical Laboratory Information Systems↗

Phosphorus transfer in surface runoff from intensive pasture systems at various scales: a review.

Phosphorus transfer in runoff from intensive pasture systems has been extensively researched at a range of scales. However, integration of data from the range of scales has been limited. This paper presents a conceptual model of P transfer that incorporates landscape effects and reviews the research relating to P transfer at a range of scales in light of this model. The contribution of inorganic P sources to P transfer is relatively well understood, but the contribution of organic P to P transfer is still relatively poorly defined. Phosphorus transfer has been studied at laboratory, profile, plot, field, and watershed scales. The majority of research investigating the processes of P transfer (as distinct from merely quantifying P transfer) has been undertaken at the plot scale. However, there is a growing need to integrate data gathered at a range of scales so that more effective strategies to reduce P transfer can be identified. This has been hindered by the lack of a clear conceptual framework to describe differences in the processes of P transfer at the various scales. The interaction of hydrological (transport) factors with P source factors, and their relationship to scale, require further examination. Runoff-generating areas are highly variable, both temporally and spatially. Improvement in the understanding and identification of these areas will contribute to increased effectiveness of strategies aimed at reducing P transfers in runoff. A thorough consideration of scale effects using the conceptual model of P transfer outlined in this paper will facilitate the development of improved strategies for reducing P losses in runoff.

Animal Husbandry↗

Computer analysis and integration of animal pathology data.

Computer storage of data from toxicology, biochemistry, haematology and pathology has been found necessary in our Laboratory in order to handle the vast amount of information generated by animal toxicology studies. The value of the system to pathology is enormous and its potential has not been exhausted. All finding, from organ weights and macroscopic observations made at autopsy, to the final histopathological diagnosis made by the pathologist are computerized. A modified version of the American College of Pathologists' systematized nomenclature of pathology is used. The pathologist recordtor whose role in the system is indispensable. The designation of a pathologist with special responsibility for supervising the computerisation and its scientific validity ensures its smooth running. The integration of data from haematology and clinical chemistry as a profile for each animal is available to the pathologist when making the final diagnosis. The system has resulted in a standardisation of pathological terminology, greater speed and improved accuracy in report formulation, the establishment of a readily retrievable in-house data bank and an enormous saving in the time of pathologists and secretaries.

Humans↗

Integrated fossil and molecular data reconstruct bat echolocation.

Molecular and morphological data have important roles in illuminating evolutionary history. DNA data often yield well resolved phylogenies for living taxa, but are generally unattainable for fossils. A distinct advantage of morphology is that some types of morphological data may be collected for extinct and extant taxa. Fossils provide a unique window on evolutionary history and may preserve combinations of primitive and derived characters that are not found in extant taxa. Given their unique character complexes, fossils are critical in documenting sequences of character transformation over geologic time and may elucidate otherwise ambiguous patterns of evolution that are not revealed by molecular data alone. Here, we employ a methodological approach that allows for the integration of molecular and paleontological data in deciphering one of the most innovative features in the evolutionary history of mammals-laryngeal echolocation in bats. Molecular data alone, including an expanded data set that includes new sequences for the A2AB gene, suggest that microbats are paraphyletic but do not resolve whether laryngeal echolocation evolved independently in different microbat lineages or evolved in the common ancestor of bats and was subsequently lost in megabats. When scaffolds from molecular phylogenies are incorporated into parsimony analyses of morphological characters, including morphological characters for the Eocene taxa Icaronycteris, Archaeonycteris, Hassianycteris, and Palaeochiropteryx, the resulting trees suggest that laryngeal echolocation evolved in the common ancestor of fossil and extant bats and was subsequently lost in megabats. Molecular dating suggests that crown-group bats last shared a common ancestor 52 to 54 million years ago.

Animals↗

Pathway mapping tools for analysis of high content data.

The complexity of human biology requires a systems approach that uses computational approaches to integrate different data types. Systems biology encompasses the complete biological system of metabolic and signaling pathways, which can be assessed by measuring global gene expression, protein content, metabolic profiles, and individual genetic, clinical, and phenotypic data. High content screening assays can also be used to generate systems biology knowledge. In this review, we will summarize the pathway databases and describe biological network tools used predominantly with this genomics, proteomics, and metabolomics data but which are equally as applicable for high content screening data analysis. We describe in detail the integrated data-mining tools applicable to building biological networks developed by GeneGo, namely, MetaCore and MetaDrug.

Computational Biology↗

Integration of macromolecular diffraction data.

Diffraction intensities can be evaluated by two distinct procedures: summation integration and profile fitting. Equations are derived for evaluating the intensities and their standard errors for both cases, based on Poisson statistics. These equations highlight the importance of the contribution of the X-ray background to the standard error and give an estimate of the improvement which can be achieved by profile fitting. Profile fitting offers additional advantages in allowing estimation of saturated reflections and in dealing with incompletely resolved diffraction spots.

Crystallography, X-Ray↗

[Technical and methodologic aspects of ambulatory long-term blood pressure monitoring systems].

The development of noninvasive, portable blood pressure measuring units began in 1962 with semi-automatic devices and cassette recorders, followed by the first automatic unit in 1968 and the introduction of digital storage system in 1978. Systems in common use today consist of a portable, battery-driven blood pressure monitor and a print-out unit. In the following, the System 5200 from SpaceLabs, which has been in use for three years in our clinic, will be described. TECHNICAL ASPECTS: In a monitor unit, amplication and filtering of analogue data measured and differentiation between signal and noise is carried out. An A/D converter digitalizes the analogue data. An integrated microprocessor analyses measured data, regulates inflation and deflation of the cuff pressure, out-put of measured and calculated values on an LCD display and storage of data. Data from 200 measurements is stored in a 2K byte RAM CMOS system. A personal computer serves for programming the monitor and evaluation of the stored data. Blood pressure measurement is carried out auscultatory with a microphone or oscillometrically if Korotkoff sounds are not detected. If the signal is disturbed, measurement is repeated within two minutes. Blood pressure measurements are performed at freely-programmable intervals from six to 60 minutes; varying time intervals can also be chosen. AUSCULTATORY BLOOD PRESSURE MEASUREMENT: A miniature pump integrated in the monitor inflates the cuff within a few seconds to a pressure of 160 mmHg or, on subsequent measurements, to 25 mmHg above the last recorded systolic value. On registration of Korotkoff sounds, the cuff pressure is increased in steps of 25 mmHg until the sounds disappear and then deflated in steps of 3 to 5 mmHg. On detection of the first Korotkoff sound, the instantaneous cuff pressure (which is converted to an electric signal by a transducer) is stored as the systolic value. Further deflation then occurs rapidly to 90 mmHg or 10 mmHg above the last measured diastolic value. With higher diastolic values, again, there is an increase in cuff pressure in steps of 25 mmHg until the onset of Korotkoff sounds and then renewed deflation in steps of 3 to 5 mmHg. On disappearance of the Korotkoff sounds, the prevailing cuff pressure is recorded and stored as the diastolic value (Figure 1). To register the Korotkoff sounds optimally, the microphone is positioned above the brachial artery.(ABSTRACT TRUNCATED AT 400 WORDS)

Ambulatory Care↗

Integration of multimodality imaging data for radiotherapy treatment planning.

This paper describes computational techniques to permit the quantitative integration of magnetic resonance (MR), positron emission tomography (PET), and x-ray computed tomography (CT) imaging data sets. These methods are used to incorporate unique diagnostic information provided by PET and MR imaging into CT-based treatment planning for radiotherapy of intracranial tumors and vascular malformations. Integration of information from the different imaging modalities is treated as a two-step process. The first step is to determine the set of geometric parameters relating the coordinates of two imaging data sets. No universal method for determining these parameters is appropriate because of the diversity of contemporary imaging methods and data formats. Most situations can be handled by one of the four different techniques described. These four methods make use of specific geometric objects contained in the two data sets to determine the parameters. These objects are: (a) anatomical and/or fiducial points, (b) attached line markers, (c) anatomical surfaces, and (d) outlines of anatomical structures. The second step involves using the derived transformation to transfer outlines of treatment volumes and/or anatomical structures drawn on the images of one imaging study to the images of another study, usually the treatment planning CT. Solid modelling and image processing techniques have been adapted and developed further to accomplish this task. Clinical examples and phantom studies are presented which verify the different aspects of these techniques and demonstrate the accuracy with which they can be applied. Clinical use of these techniques for treatment planning has resulted in improvements in localization of treatment volumes and critical structures in the brain. These improvements have allowed greater sparing of normal tissues and more precise delivery of energy to the desired irradiation volume. It is believed that these improvements will have a positive impact on the outcome of radiation therapy.

Algorithms↗

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Development of polymer-based sensors for integration into a wireless data acquisition system suitable for monitoring environmental and physiological processes.

In this work, the pressure sensing properties of polyethylene (PE) and polyvinylidene fluoride (PVDF) polymer films were evaluated by integrating them with a wireless data acquisition system. Each device was connected to an integrated interface circuit, which includes a capacitance to frequency converter (C/F) and an internal voltage regulator to suppress supply voltage fluctuations on the transponder side. The system was tested under hydrostatic pressures ranging from 0 to 17 kPa. Results show PE to be the more sensitive to pressure changes, indicating that it is useful for the accurate measurement of pressure over a small range. On the other hand PVDF devices could be used for measurement over a wider range and should be considered due to the low hysteresis and good repeatability displayed during testing. It is thought that this arrangement could form the basis of a cost-effective wireless monitoring system for the evaluation of environmental or physiological processes.

Environmental Monitoring↗

Integration of single cell multiomics data by deep transfer hypergraph neural network.

Multi-omics characterization of individual cells offers remarkable potential for analyzing the dynamics and relationships of gene regulatory states across millions of cells. How to integrate multimodal data is an open problem, existing integration methods struggle with accuracy and modality-specific biological variation retention. In this paper, we present scHyper (scalable, interpretable machine learning for single cell integration), a low-code and data-efficient deep transfer model designed for integrating paired and unpaired single-cell multimodal data. We benchmark scHyper against datasets from different multimodal data. ScHyper learns a low-dimensional representation and aligns the covariance matrices of the measured modalities, achieving high accuracy even with large scale atlas-level datasets with low memory and computational time across different cell lines, shedding light on regulatory relationships between different types of omics. Altogether, we show that scHyper is a versatile and robust tool for cell-type label transfer and integration from multimodal single-cell datasets.

Single-Cell Analysis↗

Global sulfur emissions from 1850 to 2000.

The ASL database provides continuous time-series of sulfur emissions for most countries in the World from 1850 to 1990, but academic and official estimates for the 1990s either do not cover all years or countries. This paper develops continuous time series of sulfur emissions by country for the period 1850-2000 with a particular focus on developments in the 1990s. Global estimates for 1996-2000 are the first that are based on actual observed data. Raw estimates are obtained in two ways. For countries and years with existing published data I compile and integrate that data. Previously published data covers the majority of emissions and almost all countries have published emissions for at least 1995. For the remaining countries and for missing years for countries with some published data, I interpolate or extrapolate estimates using either an econometric emissions frontier model, an environmental Kuznets curve model, or a simple extrapolation, depending on the availability of data. Finally, I discuss the main movements in global and regional emissions in the 1990s and earlier decades and compare the results to other studies. Global emissions peaked in 1989 and declined rapidly thereafter. The locus of emissions shifted towards East and South Asia, but even this region peaked in 1996. My estimates for the 1990s show a much more rapid decline than other global studies, reflecting the view that technological progress in reducing sulfur based pollution has been rapid and is beginning to diffuse worldwide.

Acid Rain↗

Building a testis.

Specific cellular, subcellular and acellular components of the rat testis including the capsule, the peritubular tissue (tunica propria) and the lymphatic endothelium were analyzed using morphometric techniques at cellular and subcellular levels to yield volume and surface area data. These data were integrated with previously published data for other cellular components of the rat testis to provide information about the volumetric composition for virtually every component of this organ. For major cell types (Leydig, Sertoli, myoid cells and germ cells) the data are expressed to the subcellular level in terms of volume and, in some instances, surface area. Graphic portrayals of testis constituents are used for rapid visual understanding of testis structure. The data presented herein are useful in conjunction with biochemical data to describe physiological properties of cells and cell components and also for understanding how structure differs under experimental and in pathological situations.

Animals↗

Physical properties of the Escherichia coli transcription termination factor rho. 2. Quaternary structure of the rho hexamer.

Under approximately physiological conditions, the transcription termination factor rho from Escherichia coli is a hexamer of planar hexagonal geometry [Geiselmann, J., Yager, T. D., Gill, S. C., Calmettes, P., & von Hippel, P. H. (1992) Biochemistry (preceding paper in this issue)]. Here we describe studies that further define the quaternary structure of this hexamer. We use a combination of chemical cross-linking and treatment with mild denaturants to show that the fundamental unit within the rho hexamer is a dimer stabilized by an isologous (or pseudoisologous) bonding interface. Three identical dimers of rho interact via a second type of isologous bonding interface to yield a hexamer with C3 or D3 symmetry. Cross-linking and denaturation experiments definitely rule out C6 and C2 symmetry for the rho hexamer. Data from fluorescence quenching, lifetime, and energy transfer experiments also argue against C2 symmetry. The simplest symmetry assignment that is not contradicted by any experimental data is D3; thus we conclude that the rho hexamer has D3 symmetry. We also consider the positioning of the binding sites for RNA and ATP relative to the coordinate reference frame of the D3 hexamer. Fluorescence energy transfer data are presented and integrated with data from the literature to arrive at a self-consistent model for the quaternary structure of the rho hexamer.

Adenosine Triphosphate↗

A doubly robust framework for addressing outcome-dependent selection bias in multi-cohort EHR studies.

Selection bias can hinder accurate estimation of association parameters in binary disease risk models using non-probability samples like electronic health records (EHRs). The issue is compounded when participants are recruited from multiple clinics/centers with varying selection mechanisms that may depend on the disease/outcome of interest. Traditional inverse-probability-weighted (IPW) methods, based on constructed parametric selection models, often struggle with misspecifications when selection mechanisms vary across cohorts. This paper introduces a new Joint Augmented Inverse Probability Weighted (JAIPW) method, which integrates individual-level data from multiple cohorts collected under potentially outcome-dependent selection mechanisms, with data from an external probability sample. JAIPW offers double robustness by incorporating a flexible auxiliary score model to address potential misspecifications in the selection models. We outline the asymptotic properties of the JAIPW estimator, and our simulations reveal that JAIPW achieves up to 6 times lower relative bias and 5 times lower root mean square error (RMSE) compared to the best performing joint IPW methods under scenarios with misspecified selection models. Applying JAIPW to the Michigan Genomics Initiative (MGI), a multi-clinic EHR-linked biobank, combined with external national probability samples, resulted in cancer-sex association estimates closely aligned with national benchmark estimates. We also analyzed the association between cancer and polygenic risk scores (PRS) in MGI to illustrate a situation where the exposure variable is not measured in the external probability sample.

Selection Bias↗