Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

[Integrated and automated data analysis for neuronal activation studies using positron emission tomography: methodology and applications].

A data analysis method was developed for neuronal activation studies using [15O]water positron emission tomography (PET). The method consists of several procedures including intra-subject head motion correction (co-registration), detection of the mid-sagittal plane of the brain, detection of the intercommissural (AC-PC) line, linear scaling and non-linear warping for anatomical standardization, pixel-by-piexl statistical analysis, and data display. All steps are performed in three dimensions and are fully automated. Each step was validated using a brain phantom, computer simulations, and data from human subjects, demonstrating accuracy and reliability of the procedure. The method was applied to human neuronal activation studies using vibratory and visual stimulations. The method detected significant blood flow increases in the primary sensory cortices as well as in other regions such as the secondary sensory cortex and cerebellum. The proposed method should enhance application of PET neuronal activation studies to the investigation of higher-order human brain functions.

Adult↗

Evaluation of sister chromatid exchange and chromosome breaks in a cohort of untreated Hodgkin's disease patients.

Cytogenetic biomarkers, chromosomal breaks [spontaneous breaks (SB) and bleomycin-induced breaks (BIB)], and sister chromatid exchange (SCE) have been shown to be sensitive cytological assays to defect susceptibility to DNA-damaging effects. However, little information is available on how environmental factors and demographic and clinical characteristics influence variation among individuals. We sought to characterize interindividual variability in a cohort of 105 untreated adult Hodgkin's disease patients. SB, BIB, and SCE data were integrated with epidemiological data by using linear regression analysis. Age, sex, ethnicity, education, histology, history of mononucleosis, and family history of cancer showed no association with any biomarker. In univariate analysis, alcohol intake was significantly associated with high SCEs (P = 0.005) and SBs (P = 0.02). Current smoking was associated only with high frequencies of SCE (P = 0.05). Advanced stage of disease was related with high SBs (P = 0.01). BIBs were not associated with any of the variables studied. In multivariate modeling, current alcohol intake was associated with high SCEs (P = 0.04) and SBs (P = 0.01). Former smokers had higher SBs than nonsmokers did (P = 0.02). A small positive correlation was found among each pair of markers. The higher SCEs and SBs in patients who smoke and consume alcohol indicate the need for evaluating these exposures when interpreting these biomarkers.

Adult↗

Integrative investigation of metabolic and transcriptomic data.

BACKGROUND: New analysis methods are being developed to integrate data from transcriptome, proteome, interactome, metabolome, and other investigative approaches. At the same time, existing methods are being modified to serve the objectives of systems biology and permit the interpretation of the huge datasets currently being generated by high-throughput methods. RESULTS: Transcriptomic and metabolic data from chemostat fermentors were collected with the aim of investigating the relationship between these two data sets. The variation in transcriptome data in response to three physiological or genetic perturbations (medium composition, growth rate, and specific gene deletions) was investigated using linear modelling, and open reading-frames (ORFs) whose expression changed significantly in response to these perturbations were identified. Assuming that the metabolic profile is a function of the transcriptome profile, expression levels of the different ORFs were used to model the metabolic variables via Partial Least Squares (Projection to Latent Structures--PLS) using PLS toolbox in Matlab. CONCLUSION: The experimental design allowed the analyses to discriminate between the effects which the growth medium, dilution rate, and the deletion of specific genes had on the transcriptome and metabolite profiles. Metabolite data were modelled as a function of the transcriptome to determine their congruence. The genes that are involved in central carbon metabolism of yeast cells were found to be the ORFs with the most significant contribution to the model.

Algorithms↗

Trends in fertility and intermarriage among immigrant populations in Western Europe as measures of integration.

Demographic data on fertility and intermarriage are useful measures of integration and assimilation. This paper reviews trends in total fertility and intermarriage of foreign populations in Europe and compares them with the trends in fertility of the host population and the sending country. In almost all cases fertility has declined. The fertility of most European immigrant populations and of some West Indian and non-Muslim Asian populations has declined to a period level at or below that of the host society. Muslim populations from Turkey, North Africa and South Asia have shown the least decline. Intermarriage is proceeding faster than might be expected in immigrant populations which seemed in economic terms to be imperfectly integrated. Up to 40% of West Indians born in the UK, for example, appear to have white partners as do high proportions of young Maghrebians in France.

Acculturation↗

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models↗

Integration of in vitro data into allometric scaling to predict hepatic metabolic clearance in man: application to 10 extensively metabolized drugs.

In this study, we investigated rational and reliable methods of using animal data to predict in humans the clearance of drugs which are mainly eliminated through hepatic metabolism. For 10 extensively metabolized compounds, adjusting the in vivo clearance in the different animal species for the relative rates of metabolism in vitro dramatically improved the predictions of human clearance compared to the approach in which clearance is directly extrapolated using body weight. Using hepatocyte data to normalize the in vivo clearances led to lower median deviations between the observed and predicted clearances in man compared to the approach normalizing data with brain weight (30-40% vs 60-80%, respectively). In addition, the approach integrating in vitro data appeared to be superior with respect to the range of deviations: approximately 2-fold underestimation, in the worst case, was observed by using in vitro data, whereas normalizing data by brain weight led to up to 10-fold underestimation of clearance in man. In addition, the integration of in vitro data provides a more rational basis to predict the metabolic clearance in man and may be applicable to compounds undergoing phase I and phase II metabolism as well.

Animals↗

Functional cranial neuronavigation. Direct integration of fMRI and PET data.

OBJECTIVE: We report our first experiences with the direct integration of fMRI data into cranial neuronavigation. METHOD: For navigation we used the MKM system and thin-sliced T1 contrast enhanced images. As a first step 21 patients had fMRI for localization of the precentral gyrus, 2 patients for Broca area detection. By anatomical correlation, these functional data were indirectly compared to the intraoperative findings using cortical SSEP (n=20) or cortical stimulation (n=3). Encouraged by these preliminary results, we started the direct integration of fMRI into neuronavigation in June 1999, followed by PET in January 2000, enabling us to compare functional images with intraoperative findings directly. fMRI and PET data were integrated by landmark matching referring on skin fiducials. Meanwhile, fMRI data of 8 patients (6 motorcortex, 2 Broca) and PET images of 1 patient were directly integrated into neuronavigation. Six out of 8 patients had additional cortical monitoring, 2/8 were exclusively operated on by functional neuronavigation. RESULTS: Using indirect comparison between fMRI and intraoperative findings we observed a good correlation in every case for the motorcortex, but only in 1/2 for the speech area. In all 6 direct integrated fMRI cases, these findings corresponded well to the conventional ones. Both patients with sole functional navigation did not have any postoperative neurological deficit. The inaccuracy of the fMRI ifT1 matching was 2. 7 mm (sigma=0.9 mm) and 1.3 mm (sigma=0.4 mm) of the subsequent referenciation of the navigation. The tumor delinement shown by 11C-methionine PET could be proven by intraoperative biopsy outside its indicated tumor margin. The inaccuracy of the PET matching was 0. 8 mm. CONCLUSION: Functional neuronavigation enables to visualize and preserve relevant brain areas. Other functional areas like short-term memory, which solely can be detected by fMRI might also be monitored in the future. The integration of PET data expect to gain a better differentiation of tumor and edema.

Adult↗

A model system for studying the integration of molecular biology databases.

MOTIVATION: Integration of molecular biology databases remains limited in practice despite its practical importance and considerable research effort. The complexity of the problem is such that an experimental approach is mandatory, yet this very complexity makes it hard to design definitive experiments. This dilemma is common in science, and one tried-and-true strategy is to work with model systems. We propose a model system for this problem, namely a database of genes integrating diverse data across organisms, and describe an experiment using this model. RESULTS: We attempted to construct a database of human and mouse genes integrating data from GenBank and the human and mouse genome-databases. We discovered numerous errors in these well-respected databases: approximately 15% of genes are apparently missing from the genome-databases; links between the sequence and genome-databases are missing for another 5-10% of the cases; about a third of likely homology links are missing between the genome-databases; 10-20% of entries classified as 'genes' are apparently misclassified. By using a model system, we were able to study the problems caused by anomalous data without having to face all the hard problems of database integration. CONTACT: nat@jax.org

Animals↗

MPact: the MIPS protein interaction resource on yeast.

In recent years, the Munich Information Center for Protein Sequences (MIPS) yeast protein-protein interaction (PPI) dataset has been used in numerous analyses of protein networks and has been called a gold standard because of its quality and comprehensiveness [H. Yu, N. M. Luscombe, H. X. Lu, X. Zhu, Y. Xia, J. D. Han, N. Bertin, S. Chung, M. Vidal and M. Gerstein (2004) Genome Res., 14, 1107-1118]. MPact and the yeast protein localization catalog provide information related to the proximity of proteins in yeast. Beside the integration of high-throughput data, information about experimental evidence for PPIs in the literature was compiled by experts adding up to 4300 distinct PPIs connecting 1500 proteins in yeast. As the interaction data is a complementary part of CYGD, interactive mapping of data on other integrated data types such as the functional classification catalog [A. Ruepp, A. Zollner, D. Maier, K. Albermann, J. Hani, M. Mokrejs, I. Tetko, U. Güldener, G. Mannhaupt, M. Münsterkötter and H. W. Mewes (2004) Nucleic Acids Res., 32, 5539-5545] is possible. A survey of signaling proteins and comparison with pathway data from KEGG demonstrates that based on these manually annotated data only an extensive overview of the complexity of this functional network can be obtained in yeast. The implementation of a web-based PPI-analysis tool allows analysis and visualization of protein interaction networks and facilitates integration of our curated data with high-throughput datasets. The complete dataset as well as user-defined sub-networks can be retrieved easily in the standardized PSI-MI format. The resource can be accessed through http://mips.gsf.de/genre/proj/mpact.

Databases, Protein↗

transFusion: a novel comprehensive platform for integration analysis of single-cell and spatial transcriptomics.

MOTIVATION: Understanding spatial organization, intercellular interactions, and regulatory networks within the spatial context of tissues is crucial for uncovering complex biological processes and disease mechanisms. Spatial transcriptomics technologies have revolutionized this field by enabling the spatially resolved profiling of gene expression. 10× Visium has emerged as the predominant spatial technology, but its low resolution and the complexity of integrating multimodal datasets present significant analytical challenges, particularly for researchers with limited computational and statistical expertise. Current spatial transcriptomics analysis platforms generally fall short of effectively integrating multimodal data and maximizing the utility of spatial information-such as uncovering complex cellular spatial dependencies, multimodal gradient patterns, and spatial coexpression of ligand-receptor pairs and regulatory networks related to disease or biological states-thereby limiting their ability to provide comprehensive end-to-end analytical workflows when analyzing 10× Visium data. RESULTS: To address these limitations, we developed transFusion, a novel, advanced web-based platform specializing in the most comprehensive and effective integration analysis of scRNA-seq and 10× Visium spatial transcriptomics data. transFusion offers 12 key functions, from basic visualization to advanced analyses, including intercellular dependency analysis, ligand-receptor coexpression identification and visualization, and spatial multimodal gradient variation patterns. Two case studies were used to demonstrate transFusion's capabilities in exploring tissue architecture, intercellular communication, dependency networks, and multimodal gradient variation patterns with minimal computational skills and statistical expertise. transFusion provides a flexible and powerful framework for multimodal data integration analysis. AVAILABILITY AND IMPLEMENTATION: transFusion is freely available at https://github.com/WQLin8/transFusion.

Spatial Transcriptomics↗

Integrated analysis of multiple data sources reveals modular structure of biological networks.

It has been a challenging task to integrate high-throughput data into investigations of the systematic and dynamic organization of biological networks. Here, we presented a simple hierarchical clustering algorithm that goes a long way to achieve this aim. Our method effectively reveals the modular structure of the yeast protein-protein interaction network and distinguishes protein complexes from functional modules by integrating high-throughput protein-protein interaction data with the added subcellular localization and expression profile data. Furthermore, we take advantage of the detected modules to provide a reliably functional context for the uncharacterized components within modules. On the other hand, the integration of various protein-protein association information makes our method robust to false-positives, especially for derived protein complexes. More importantly, this simple method can be extended naturally to other types of data fusion and provides a framework for the study of more comprehensive properties of the biological network and other forms of complex networks.

Algorithms↗