Search PubMedSearch

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

AI-HOPE: an AI-driven conversational agent for enhanced clinical and genomic data integration in precision medicine research.

MOTIVATION: The growing complexity of clinical cancer research has fueled a surge in demand for automated bioinformatics tools capable of integrating clinical and genomic data to accelerate discovery efforts. RESULTS: We present the Artificial Intelligence Agent for High-Optimization and Precision Medicine (AI-HOPE), an AI-driven system that enables domain experts to conduct integrative data analyses through natural language interactions. Powered by Large Language Models, AI-HOPE interprets user instructions, converts them into executable code, and autonomously analyzes locally stored data. It supports flexible association studies, subset comparisons, clinical prevalence assessments and survival analyses. In addition, AI-HOPE enables global variable scans to identify features significantly associated with a user-defined outcome, making a powerful and intuitive tool for advancing precision medicine research. Importantly, its closed-system design prevents clinical data leakage. To demonstrate its utility, AI-HOPE was applied to The Cancer Genome Atlas data to address two clinical questions. First, it identified significant enrichment of TP53 mutations in late-stage colorectal cancer compared to early-stage cases. Second, it uncovered a strong association between KRAS mutations and poorer progression-free survival in FOLFOX-treated patients. These findings align with established literature and demonstrate AI-HOPE's ability to generate meaningful insights independently, without prior assumptions. By removing programming barriers and simplifying complex analyses, AI-HOPE bridges the gap between data complexity and research needs. With its scalable and adaptable framework, AI-HOPE has the potential to support diverse biomedical research fields, driving innovation and efficiency in translational studies. AVAILABILITY AND IMPLEMENTATION: The AI-HOPE software and demonstration data is available at https://github.com/Velazquez-Villarreal-Lab/AI-HOPE.

Precision Medicine

Scalable, open-access and multidisciplinary data integration pipeline for climate-sensitive diseases.

Climate-sensitive infectious diseases pose an important challenge for human, animal and environmental health and it has been estimated that over half of known human pathogenic diseases can be aggravated by climate change. While climatic and weather conditions are important drivers of transmission of vector-borne diseases, socio-economic, behavioural, and land-use factors as well as the interactions among them impact transmission dynamics. Analysis of drivers of climate-sensitive diseases require rapid integration of interdisciplinary data to be jointly analysed with epidemiological (including genomic and clinical) data. Current tools for the integration of multiple data sources are often limited to one data type or rely on proprietary data and software. To address this gap, we develop a scalable and open-access pipeline for the integration of multiple spatio-temporal datasets that requires only the declaration of the country and temporal range and resolution of the study. The tool is locally deployable and can easily be integrated into existing climate-disease-modelling applications. We demonstrate the utility of the tool for dengue modelling in Vietnam where epidemiological data are legally required to remain local. We include a pipeline for bias correction of climate data to enhance their quality for downstream modelling tasks. The Dengue Advanced Readiness Tools-Pipeline empowers users by simplifying complex download, correction, and aggregation steps, fostering data-driven discovery of relationships between infectious diseases and their drivers in space and time, and enhancing reproducibility in research. Additional modules and datasets can be added to the existing ones to make the pipeline extendable to use cases other than the ones presented here.

automated workflows

The organization engine: virtual data integration.

The Organization Engine is an early example of Virtual Data Integration--providing the appearance of integration at the desktop without modifying existing infrastructure. Starting with the Organization Engine, eight programming days were needed to provide uniform desktop access to a CODASYL-compliant hospital information system and to a MUMPS-based radiology information system (the technique is equally effective for relational and other data bases). The resulting tool provides a seamless integration of these two systems, image storage, pre-recorded audio, and document storage. In addition to providing uniform access, the tool allows healthcare providers to organize the data to suit their individual needs. The ease of this integration lies in two simple techniques: the transformation of data from all sources into a single, homogeneous representation, and the use of simple customization files to describe new object types and formats. The approach is sufficiently general to allow the integration of applications which present external interfaces of radically different forms. Two such forms are discussed here: data map publication and transactions.

Computer Communication Networks

A comparison of two methods for facilitating clinical data integration by medical students.

Two important charting strategies to help students organize patients' data are Weed's problem-oriented medical record (POMR) and Russell's condition diagram (CD). The authors conducted the present study in 1987 to determine whether either was superior for clinical data integration. Sophomore medical students at The University of Texas Medical School at San Antonio indicated whether they preferred the POMR, the CD, or neither. They were then divided into three study groups according to their preferences, with the POMR and CD groups receiving 80 hours of training and the control group receiving only the standard preclinical training. Each student then examined a standardized patient and wrote an open-ended report about the patient's medical problem. After examining a second patient, students were asked to write a structured report providing information about each of ten components of diagnosis. Both the CD and the POMR groups scored numerically higher on the structured type of report than did the controls, but only the CD group scored significantly higher. The CD group also scored higher than did the POMR group on both types of report, but the differences were not statistically significant. This study indicates that the clinical reasoning of medical students can be enhanced by focused training in either the CD or the POMR methods. It suggests that the CD format may be particularly helpful for students with lower academic achievement.

Algorithms

Integrated data mining and network pharmacology to explore the prescription patterns from a senior TCM oncologist's clinical practice in treating chemotherapy-induced hand-foot syndrome.

Hand-foot syndrome (HFS) is a common and refractory adverse effect of chemotherapy lacking specific therapeutic strategies currently. Traditional Chinese medicine (TCM) has shown empirical efficacy in clinical HFS management. This study integrated data mining and network pharmacology to systematically elucidate the medication principles and molecular mechanisms underlying Professor Gang Xie's prescriptions for HFS. All medical records from Professor Xie's specialist clinic (January 2020 to March 2025) were retrospectively collected and standardized in Excel. Prescriptions were analyzed through frequency statistics, association and clustering. Active ingredients of core herb pairs and their disease-related targets were identified using TCMSP, HERB, GeneCards, PharmGKB and GEO databases. Protein-protein interaction (PPI) networks, gene ontology (GO), and Kyoto encyclopedia of genes and genomes (KEGG) pathway analyses were performed. Molecular docking validated interactions between key bioactive compounds and targets. This study involved 217 prescriptions containing 150 herbs. Core herb combinations comprised Radix Astragali (Huangqi), Poria (Fuling), and Radix Pseudostellariae (Taizishen), predominantly classified as spleen-tonifying agents with warm properties, targeting lung, spleen, and stomach meridians. Network analysis identified 67 bioactive compounds and 899 disease targets. Quercetin, kaempferol, acacetin and luteolin were identified the key ingredients. The core targets (TP53, STAT3, PIK3CA, HSP90AA1, AKT1, CTNNB1, PI3KR1, MAPK1) were enriched in MAPK and PI3K-Akt signaling pathways. Molecular docking confirmed strong binding affinity between key compounds and targets. Professor Xie's therapeutic strategy for HFS emphasizes "spleen fortification, phlegm elimination, and stasis resolution." The core herb combination likely exerts anti-HFS effects via modulation of MAPK and PI3K-Akt pathways, providing a pharmacological basis for TCM-driven HFS management.

Network Pharmacology

The network approach: a means for the collection of integrated data following standardized protocols.

The understanding of the epidemiology of a vector borne disease, involving various vector and host species in a defined area requires a multidisciplinary approach. It is essential that specialists obtain data relevant to common objectives. In the case of African Trypanosomiasis this means that the observations being made on the health status of the human and domestic--and wild animal populations are made over the same period of time as those on the tsetse population to which these hosts are exposed. Standardization of methodology is a pre-condition for reliable comparison of observations. The creation of a network of research situations is one possibility for the fulfillment of this pre-condition; while at the same time it is a suitable means for the collection of integrated data. The African Trypanotolerant Livestock Network, created in 1982 is one example of such a network. Examples of conclusions which could be drawn after analysis of data collected over a two-year period through this Network are presented.

Animals

Global inequities in hepatitis B and C genomic surveillance revealed through an interactive data integration dashboard.

OBJECTIVES: To assess global disparities in hepatitis B virus (HBV) and hepatitis C virus (HCV) genomic surveillance and to develop an integrated platform that links genomic data with epidemiological burden. STUDY DESIGN: Retrospective observational analysis. METHODS: We reviewed existing viral genomic repositories to identify structural and analytical limitations. Subsequently, we integrated 10 996 HBV and 3533 HCV whole-genome sequences (WGS) from public databases with Global Burden of Disease (GBD) estimates to quantify inequities in genomic surveillance across countries and genotypes. Using these data, we developed the open-access Hepatitis Dashboard, incorporating >14 000 sequences from 141 countries with GBD metrics to evaluate representativeness and sequencing coverage relative to disease burden. RESULTS: Marked inequities in hepatitis genomic surveillance were identified. Despite increasing HBV- and HCV-associated mortality, virus sequence availability remains geographically and genotypically skewed-dominated by China and the United States, with substantial underrepresentation of HBV genotype E and HCV genotypes 5 and 8. Many high-endemic countries in Africa and the Western Pacific remain severely undersampled. We detected circulating antiviral drug-resistance mutations and developed a burden-adjusted sequencing coverage metric, revealing that several high-burden countries, including China, Nigeria and India, are among the least represented in global genomic datasets. Projections to 2030 indicate that neither HBV nor HCV are currently on track to meet WHO elimination targets. CONCLUSIONS: The Hepatitis Dashboard provides an integrated, continuously updated resource that links genomic and epidemiological data to quantify and visualise global surveillance gaps. This analysis highlights a critical disconnect between sequencing efforts and public health needs, which may limit the effectiveness of surveillance-informed strategies to support progress toward WHO 2030 elimination goals. By enabling burden-adjusted prioritisation and longitudinal tracking of genomic coverage, the platform supports evidence-based sampling strategies, equitable resource allocation, and monitoring of global progress toward hepatitis elimination.

Humans

Research data integrity: a result of an integrated information system.

The toxicologic problems of today frequently require long-term, multidisciplinary experimentation involving large numbers of animals. In order to provide the extensive safety evaluation necessary to produce data that can be reasonably extrapolated to humans, automated research support systems have transcended the position of useful tools and have become an integral part of the total design of experimental protocols. For an automated information system to fully represent the reality of the experiment, it must be able to assure integrity, as well as provide for the storage, calculation, and retrieval of data values of the quality and quantity necessary for fulfilling protocol requirements. Guarantees against error and loss of data, in addition to flexibility and easy access, must be an inherent part of the system if the acceptance and condifence of the investigator are to be obtained. This paper discusses the criteria, philosophies, and benefits of integrated data systems that ensure integrity of toxicologic research support.

Computers

Interpretable data integration for single-cell and spatial multi-omics.

Integrating single-cell or spatial transcriptomic and epigenomic data enables scrutinizing the transcriptional regulatory mechanisms controlling cell fate. Current integration methods usually align multi-omics data into a shared latent space but fail to reveal the underlying connections between genes and regulatory elements. The correlation- or regression-based regulatory inference methods cannot dissect different transcriptional regulation codes for cells under different spatial and temporal states. To address both problems, we develop a feature-guided optimal transport (FGOT) method, which simultaneously uncovers cellular heterogeneity and their associated transcriptional regulatory links. FGOT also provides post hoc interpretability for existing integration methods. FGOT is applicable for paired/unpaired single-cell multi-omics data and paired spatial multi-omics data. Benchmarking and validating via histone modification data or three-dimensional (3D) genomics data show good robustness and accuracy in integration and inference of regulatory links. The method allows systematic screening of cell-state and spatial-location-specific regulatory elements in diseases at the single-cell level. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis

Improving recombinant protein productivity in CHO cells via multi-omics data integration.

Chinese hamster ovary (CHO) cells represent the dominant host system for the production of recombinant therapeutic proteins. In recent decades, extensive research has focused on process/media optimization and cell line engineering to improve both the productivity and quality of biopharmaceutical proteins produced in CHO cells. Nevertheless, the inherent complexity of biological pathways and the heterogeneous cellular responses to different environmental conditions have posed substantial challenges to traditional methodologies. Recent advances in omics technologies have enabled comprehensive characterization of CHO cell physiology, providing multidimensional molecular and phenotypic insights that facilitate the enhancement of recombinant protein production. This review first summarizes the methodologies and advances in CHO omics research, including genomics, transcriptomics, proteomics, metabolomics, and epigenomics. It then examines contemporary approaches to integrate and analyze multi-omics data in CHO cells. The review further elucidates how these multi-omics datasets can be strategically applied across various developmental stages, including cell line selection, genetic engineering, expression vector design, and bioprocess optimization. Finally, we explore the transformative potential of integrating multi-omics with artificial intelligence and discuss promising future research directions in CHO cell studies. These emerging paradigms offer novel opportunities for data-driven cell engineering and bioprocess optimization in CHO-based biomanufacturing.

Bioprocessing

CrossAttOmics: multiomics data integration with cross-attention.

MOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal.

Humans

The phenomenology of spatial integration: data and models.

A briefly presented visual stimulus followed by darkness seems to persist beyond its physical offset. We are concerned here with the relation between two characteristics of this visible persistence: first, its phenomenological resemblance to the stimulus that spawned it and second, its usefulness as a basis for integrating visual stimuli that are separated in time. We describe two experiments using a task in which two halves of a visual stimulus were presented successively and observers reported how complete the stimulus appeared to be. Stimuli appeared less complete with increases in both the duration of the interval intervening between presentation of the two halves and the duration of the initially presented stimulus half. This data pattern is similar to that obtained in tasks in which spatial integration of two temporally disparate stimuli is necessary for correct responding. On the basis of this similarity, we argue that phenomenological appearance and ability to integrate stimuli over time are two facets of the same perceptual events. We describe a formal model to account for these and other data.

Adult

Multimodal artificial intelligence and machine learning in oncology: from data integration to precision cancer care.

Cancer remains a major global health burden, with approximately 20 million new cases and 9.7 million cancer-related deaths reported globally in 2022. While advances in radiological imaging, molecular profiling, and clinical data have enhanced the interpretation of disease progression, the availability of multiple such modalities still does not meet the needs of a large patient population. This narrative review focuses on the role of multimodal artificial intelligence and machine learning in bridging the gap in interpreting heterogeneous modalities to improve risk prediction, prognostic assessment, and treatment decision-making in precision oncology. Multimodal frameworks such as Pathomic Fusion illustrate how complementary histopathological and genomic information can be integrated for cancer diagnosis and prognostic modeling. Multimodal models have demonstrated potential in virtual biopsy, cancer screening, prognostic prediction, radiotherapy planning, intraoperative guidance, and clinical-trial design using digital twins and synthetic control arms. The major limitations of incorporating multimodal artificial intelligence and machine learning in oncology include data heterogeneity, demographic or institutional biases, and reproducibility challenges that hinder translation. Accordingly, appropriate data-governance strategies, fairness audits, and privacy-preserving approaches such as federated learning should be considered where appropriate. Future progress will depend on the development of standardized benchmarking datasets, robust external validation, seamless integration with electronic health records and picture archiving and communication systems, and the implementation of explainable, secure, and clinically validated multimodal artificial intelligence frameworks that support precision oncology in routine clinical practice.

deep learning

[Viewpoints on preparations for integrating data processing systems in routine microbiology diagnosis].

The enormously risen and further increasing numbers of examinations and tests in microbiological diagnostics within the last years need new methods for treatment. One possibility to meet the higher requirements for information of the clinic without loss in quality at constant staff is the integration of the microcomputer technique into the laboratory as direct "tool". Demands for a qualitatively high empirical antimicrobial chemotherapy, chemotherapy according to antibiotic susceptibility tests, indicated use of antimicrobial drugs and control measures of infectious processes in general are met only by means of a fast information processing. The microcomputer technique in the laboratory provides also the chance to automate still manually performed tests and comprises according to algorithm the strict observation of the diagnostic process and its control. The application of the microcomputer technique on the one hand means for the technical assistant the omission of much manually performed work, on the other hand enables work of higher quality and supports decisions in the diagnostic process. Mathematical and statistical calculations are no longer connected with great losses of activity. The actual need for information of the clinician is met in time in different ways.

Bacteriological Techniques

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software

SeqUIaSCOPE: multi-omics data integration platform for single-patient clinical oncology pathway exploration.

SUMMARY: SeqUIaSCOPE is an open-source platform designed for routine clinical oncology diagnostics through case-centric integration and visualization of genomic variants, fusion events, and expression profiles. The platform combines molecular-level validation via embedded genome browsing with systems-level interpretation through dynamic pathway visualization, enabling geneticists to assess how alterations converge across biological networks. Flexible reporting with customizable templates accommodates diverse institutional requirements, while secure cluster-based or local deployment ensures compliance with data protection policies, making advanced multi-omics diagnostics accessible to academic and clinical institutions. AVAILABILITY AND IMPLEMENTATION: SeqUIaSCOPE is freely available on GitHub at https://github.com/BioIT-CEITEC/sequiascope under the MIT license and archived at Zenodo (https://zenodo.org/records/21338445). Due to the sensitive nature of patient data, the repository provides simulated datasets that mimic the structure of real clinical data for testing and exploration. Documentation and a live demo accompany these datasets, allowing users to explore the application without any prior setup. The repository also includes a Helm chart for Kubernetes deployment and Docker containers for local deployment, ensuring compatibility across Linux, macOS, and Windows. No user registration is required, and all data remains on local or institutional infrastructure.

Humans

Prioritization of causal genes from genome-wide association studies by Bayesian data integration across loci.

MOTIVATION: Genome-wide association studies (GWAS) have identified genetic variants, usually single-nucleotide polymorphisms (SNPs), associated with human traits, including disease and disease risk. These variants (or causal variants in linkage disequilibrium with them) usually affect the regulation or function of a nearby gene. A GWAS locus can span many genes, however, and prioritizing which gene or genes in a locus are most likely to be causal remains a challenge. Better prioritization and prediction of causal genes could reveal disease mechanisms and suggest interventions. RESULTS: We describe a new Bayesian method, termed SigNet for significance networks, that combines information both within and across loci to identify the most likely causal gene at each locus. The SigNet method builds on existing methods that focus on individual loci with evidence from gene distance and expression quantitative trait loci (eQTL) by sharing information across loci using protein-protein and gene regulatory interaction network data. In an application to cardiac electrophysiology with 226 GWAS loci, only 46 (20%) have within-locus evidence from Mendelian genes, protein-coding changes, or colocalization with eQTL signals. At the remaining 180 loci lacking functional information, SigNet selects 56 genes other than the minimum distance gene, equal to 31% of the information-poor loci and 25% of the GWAS loci overall. Assessment by pathway enrichment demonstrates improved performance by SigNet. Review of individual loci shows literature evidence for genes selected by SigNet, including PMP22 as a novel causal gene candidate.

Genome-Wide Association Study