Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Predicting co-complexed protein pairs using genomic and proteomic data integration.

BACKGROUND: Identifying all protein-protein interactions in an organism is a major objective of proteomics. A related goal is to know which protein pairs are present in the same protein complex. High-throughput methods such as yeast two-hybrid (Y2H) and affinity purification coupled with mass spectrometry (APMS) have been used to detect interacting proteins on a genomic scale. However, both Y2H and APMS methods have substantial false-positive rates. Aside from high-throughput interaction screens, other gene- or protein-pair characteristics may also be informative of physical interaction. Therefore it is desirable to integrate multiple datasets and utilize their different predictive value for more accurate prediction of co-complexed relationship. RESULTS: Using a supervised machine learning approach--probabilistic decision tree, we integrated high-throughput protein interaction datasets and other gene- and protein-pair characteristics to predict co-complexed pairs (CCP) of proteins. Our predictions proved more sensitive and specific than predictions based on Y2H or APMS methods alone or in combination. Among the top predictions not annotated as CCPs in our reference set (obtained from the MIPS complex catalogue), a significant fraction was found to physically interact according to a separate database (YPD, Yeast Proteome Database), and the remaining predictions may potentially represent unknown CCPs. CONCLUSIONS: We demonstrated that the probabilistic decision tree approach can be successfully used to predict co-complexed protein (CCP) pairs from other characteristics. Our top-scoring CCP predictions provide testable hypotheses for experimental validation.

Computational Biology↗

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning↗

The implanted electrical resistance strain gauge: in vitro studies on data integrity.

The aim of this research is to investigate, and thus counter, the adverse effects of tissue fluid ingress on the performance of the electrical resistance strain gauge when used in ascertaining in vivo loading on a spinal implant. Moisture absorption has been minimized by adopting maximum metallic coverage in a package comprising stainless steel foil on vacuum-injected pacemaker grade epoxide. In a simulation of the implanted environment, cyclic strain wet endurance testing in saline suggests that, in the body, the fall in indicated quasi-dynamic strain would be less than 1.5% at 24 weeks post-operation (the longevity needed to span adequately the bony fusion phase). This implies that stiffening of the fusion mass will be deducible to a similar accuracy (from stepped-load exercises), in which creep is a secondary effect. However, crucial information (from quasi-static (passive) studies) regarding remodelling and load-sharing processes would be subject to a total signal error (primarily due to grid corrosion) in excess of 16% by 24 weeks, since long-term drifts are not inherently cancelled. Signal compensation is therefore additionally required, and an approximate empirical characterization of total error versus time has been derived.

Animals↗

Advancing the state of data integration in healthcare.

There is growing consensus that clinical information systems will provide the bridge to advancing the integration of information systems in healthcare. In spite of developments in technology that have enabled some organizations to integrate clinical information with care delivery in ways that can promote safer, more efficient patient care, the majority of healthcare has yet to achieve this goal. Why aren't we there yet?

Diffusion of Innovation↗