Search PubMedSearch

SEARCH · Search PubMed

Results for “Health data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Making waves: toward systems-level interpretation of hormonal and endogenous biomarkers in wastewater-based epidemiology.

Wastewater-based epidemiology (WBE) has proven invaluable for population health monitoring, most notably during the COVID-19 pandemic. Yet current WBE largely relies on exogenous markers such as drugs, pathogens, and their metabolites, limiting surveillance to what communities are exposed to. We argue for expanding WBE towards endogenous biomarkers, particularly hormones, which provide insights into physiological stress, metabolic function, and endocrine activity. Hormone-based WBE offers new opportunities to capture population-level biological responses to societal and environmental stressors, disasters, and chronic disease burdens at the community scale. This perspective outlines a systems-level framework for integrating hormonal signals in wastewater with clinical data, behavioral indicators, environmental factors, and digital markers to support more robust and context-aware public health surveillance. We highlight key technical considerations, interpretive challenges, and opportunities for translational pilot studies. By moving beyond exposure tracking toward more integrated interpretation of biological responses, hormone-informed WBE may contribute to more resilient, inclusive, and actionable public health infrastructure.

Humans

ClarID: A Human-Readable and Compact Identifier Specification for Biomedical Metadata Integration.

BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., 'CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W' that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (from that event); and (ii) a compact version named 'stub' (e.g., 'CT01001LTR0N401T1W') optimized for filenames, pipelines, and labeling.ClarID is implemented through an open-source command-line tool, ClarID-Tools, which processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID's utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID fills a critical gap between opaque accession numbers and rich metadata schemas by embedding key context directly into structured identifiers. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools.

Biosample identifiers

The positive known association design: a quality assurance method for occupational health surveillance data.

Quality control must be an integral component of an occupational health surveillance program. The positive known association design offers the occupational health physician a method to test, on a population basis (ie, high periodic medical surveillance examination participation rates by the employees), the quality of periodic medical surveillance data. Several well-established biological associations were evaluated and observed in this study, including a dramatic relation between white blood cell counts and smoking. We highly recommend that the positive known association design be incorporated in the quality assurance procedures of occupational health surveillance programs.

Adult

CAUSAL artificial intelligence and data-driven decision intelligence in personalized medicine: a review of healthcare informatics systems.

This review examines the integration of causal artificial intelligence (AI) and data-driven decision intelligence within healthcare informatics systems to advance personalized medicine and clinical decision-making. A narrative review methodology was employed, synthesizing interdisciplinary literature from major databases, including PubMed, Scopus, Web of Science, IEEE Xplore, and ScienceDirect. Studies focusing on causal inference, decision intelligence, and healthcare informatics applications in personalized medicine were included. Data were extracted on methodological approaches, healthcare settings, analytical techniques, and clinical applications, followed by thematic synthesis. Findings indicate that causal AI enhances clinical decision support by enabling estimation of treatment effects and simulation of intervention outcomes at the individual patient level. Integration of multimodal health data such as electronic health records, genomic data, and real-time monitoring improves prediction accuracy and supports tailored treatment strategies. Additionally, causal models improve interpretability, fostering clinician trust and facilitating transparent decision-making. Robust healthcare informatics infrastructures, including interoperable systems and data warehouses, were identified as critical enablers of causal analytics. Overall, causal AI represents a transformative advancement in healthcare analytics, supporting more informed, individualized, and evidence-based clinical decisions. Its integration within healthcare informatics systems has significant potential to improve patient outcomes and guide the future of intelligent, personalized healthcare delivery.

Precision Medicine

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease

The use of an information system for community health services planning and management.

The challenge of delivering health care on a more cost-effective and equitable basis has led to the formation of new kinds of organizations, which in turn require new kinds of management information systems. The system described in this article is used for planning resource allocation and monitoring the overall performance of Livingston Community Health Services, Inc., a rural community-owned health service designed to provide comprehensive care to a geographically defined target population. The Health Services Data System processes information on performance and productivity, effectiveness with respect to the target population, and billing and financial activities. Input consists of the results of a community census and two household surveys, patient registrations and patient services operating data, and financial data. Besides billing patients automatically, the system integrates financial, demographic, and health services utilization data to generate monthly summaries for the administrators, medical director, and community board. The article discusses several examples of these summaries, stratifying utilization by geographic location of residence, age, income, and race.

Accounting

Toward a unified approach: Considerations for bioinformatic and sequencing activities & data in wastewater surveillance of biologic public health threats.

Genomic technologies such as PCR and next-generation sequencing (NGS) have greatly advanced public health surveillance, especially during COVID-19, by enabling detailed tracking of pathogen spread, origins, and variants. While PCR is vital for targeted detection, falling NGS costs have made large-scale, high-throughput sequencing more feasible, supporting broader pathogen monitoring-including the detection of vaccine escape variants and new strains. Applying NGS to wastewater offers valuable population-level insights but faces challenges such as variable sample complexity, the need for skilled staff, suitable platforms, and robust IT infrastructure. Although there are currently a lot of efforts towards defining guidelines for sampling, analysis, and integrating wastewater data into public health policy, such as the recently published International Cookbook for Wastewater Practitioners, they often lack universal applicability, emphasizing the analytical approaches in favour of the NGS-based approaches. However, standardising protocols for sampling, sequencing, and analysis is crucial to ensure reliable, comparable data across surveillance systems worldwide. Pilot studies and continuous refinement are recommended to overcome implementation hurdles and fully realise the benefits of NGS in wastewater surveillance. This work attempts to outline these challenges and opportunities across the entire wastewater surveillance workflow, from data generation to reporting, and provide some concrete suggestions and considerations across the spectrum of activities. We further highlight that the infrastructure, funding and government-policy context in which surveillance operates acts as an enabling condition for these activities, and that technical standardisation alone is unlikely to deliver durable, comparable surveillance in its absence.

considerations

A One Health perspective: Genomic insights into temporal trends of antimicrobial resistance and zoonotic transmission risks in Escherichia coli from human and swine.

Antimicrobial resistance (AMR) poses a significant challenge within the One Health framework. By integrating genomic data from 824 E. coli isolates obtained from 22 swine farms in southwestern China with 8432 publicly available genomes from human and swine sources, this study provides comprehensive insights into the temporal trends and divergence of AMR in human and swine E. coli populations, the risk of AMR transmission from swine to human, and the evolutionary mechanisms underlying the human adaptation of ST2 strains. The results revealed an overall increase in AMR until approximately 2016, followed by a subsequent decline. However, resistance to tetracyclines, quinolones, and phenicols continues to exhibit an upward trend, highlighting the urgency of enhancing regulatory measures targeting these drugs. Horizontal gene transfer play pivotal roles in shaping distinct AMR profiles in human and swine strains. ST2 E. coli was identified as a major carrier of AMR in both human and swine, and also served as the primary reservoir of blaNDM-5 within the human-associated lineage. During evolution, ST2 E. coli underwent significant genetic changes, including the enrichment of blaNDM-5 and remodeling of virulence factors, facilitating its transition from a generalist lineage colonizing both human and swine to a human-adapted lineage.

Humans

From fragmentation to coordination: strengthening One Health research to support H5N1 preparedness in Cambodia.

OBJECTIVES: Highly pathogenic avian influenza A (H5N1) remains a major zoonotic threat, characterized by persistent transmission in Cambodia since its re-emergence in 2023. Despite strengthened surveillance and the establishment of the Inter-Ministerial Coordination Committee on One Health, limited integration of research across sectors constrains preparedness and response. This viewpoint examines how research supports the One Health system in Cambodia. METHODS: This viewpoint draws on insights obtained from the first national multistakeholder workshop on H5N1, held in March 2026. RESULTS: Fragmentation across epidemiological, clinical, behavioral, environmental, and genomic domains limits the generation of actionable evidence and delays its translation into policy. CONCLUSION: We propose the establishment of a multisectoral technical working group on H5N1 research embedded within the Inter-Ministerial Coordination Committee on One Health to align research priorities, strengthen data integration, and improve evidence-to-policy translation. This approach could enhance national preparedness while simultaneously positioning Cambodia as a model for coordinated One Health research in the Western Pacific region and beyond.

Avian influenza A (H5N1)

Multimodal artificial intelligence and machine learning in oncology: from data integration to precision cancer care.

Cancer remains a major global health burden, with approximately 20 million new cases and 9.7 million cancer-related deaths reported globally in 2022. While advances in radiological imaging, molecular profiling, and clinical data have enhanced the interpretation of disease progression, the availability of multiple such modalities still does not meet the needs of a large patient population. This narrative review focuses on the role of multimodal artificial intelligence and machine learning in bridging the gap in interpreting heterogeneous modalities to improve risk prediction, prognostic assessment, and treatment decision-making in precision oncology. Multimodal frameworks such as Pathomic Fusion illustrate how complementary histopathological and genomic information can be integrated for cancer diagnosis and prognostic modeling. Multimodal models have demonstrated potential in virtual biopsy, cancer screening, prognostic prediction, radiotherapy planning, intraoperative guidance, and clinical-trial design using digital twins and synthetic control arms. The major limitations of incorporating multimodal artificial intelligence and machine learning in oncology include data heterogeneity, demographic or institutional biases, and reproducibility challenges that hinder translation. Accordingly, appropriate data-governance strategies, fairness audits, and privacy-preserving approaches such as federated learning should be considered where appropriate. Future progress will depend on the development of standardized benchmarking datasets, robust external validation, seamless integration with electronic health records and picture archiving and communication systems, and the implementation of explainable, secure, and clinically validated multimodal artificial intelligence frameworks that support precision oncology in routine clinical practice.

deep learning

Culture and health-related schemas: a review and proposal for interdisciplinary integration.

We present a comprehensive review of anthropological, sociological, and psychological theory and data on the structure, content, and function of health-related schemas. Health psychology's need to integrate specific variables and principles from the other disciplines is highlighted. Suggestions for future research are offered, and the importance of cultural factors in health beliefs is emphasized.

Anthropology

Casualty encounters at a small rural hospital.

OBJECTIVE: To describe the reasons for encounter (RFE) at the casualty department of a small rural hospital and to highlight the value of the hospital to the community, and to health care workers, medical educators, and policy makers. SETTING: A small South Australian rural town with a population of about 4500 served by a 50-bed hospital that provides a 24 hour casualty service manned by the local three-person general practice on a fee-for-service basis. METHODS: Using an integrated computerised health information management system, data on all the RFE at the casualty department were accumulated over 9 months, coded with ICHPPC-2-Defined, analysed and transferred to a spreadsheet for presentation. RESULTS: There were sex variations in the various age groups with males presenting more commonly with accidents and injuries. The main reasons for encounter were injuries (35%), respiratory system problems (13%), ear problems (10%), infections (5%), ill-defined problems (5%), supplementary classification (5%). CONCLUSIONS: There is sufficient 'clinical material' for undergraduate and graduate training in the management of trauma and orthopaedic problems but insufficient for obstetric and abdominal surgical emergencies in small rural hospitals. Small rural hospitals must be supported and used effectively by educators and policy-makers to help rural doctors meet the needs of the 30% of the Australian population who do not live on the coastal fringe.

Adolescent

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

PURPOSE: The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results. We performed multiple modeling experiments integrating clinical and demographic data from electronic health records with genetic data to understand which decisions may affect performance. METHODS: Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from 2 large independent health systems, and polygenic risk scores (PRS) were generated across all patients of European ancestry with genetic data in the corresponding biobanks. Crohn's disease was studied based on its substantial genetic component, established electronic health records-based definition, and sufficient prevalence for training and testing. We investigated the impact of choices regarding the PRS integration method, training sample, model complexity, and performance metrics. RESULTS: Overall, our results showed that including PRS resulted in higher performance, but this gain was only robust in situations with limited clinical information. We found consistent performance increases from more compute-intensive models, such as random forest, but the impact of other decisions varied by site. CONCLUSION: This work highlights the importance of considering methodological decision points in interpreting the impact of PRS on prediction performance in clinical models.

Humans

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results across studies. Here, we performed multiple modeling experiments integrating clinical and demographic data from electronic health records (EHR) and genetic data to understand which decision points may affect performance. Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from two large independent health systems and polygenic risk scores (PRS) were generated across all patients with genetic data in the corresponding biobanks. Crohn's disease was used as the model phenotype based on its substantial genetic component, established EHR-based definition, and sufficient prevalence for model training and testing. We investigated the impact of PRS integration method, as well as choices regarding training sample, model complexity, and performance metrics. Overall, our results show that including PRS resulted in higher performance by some metrics but the gain in performance was only robust when combined with demographic data alone. Improvements were inconsistent or negligible after including additional clinical information. The impact of genetic information on performance also varied by PRS integration method, with a small improvement in some cases from combining PRS with the output of a clinical model (late-fusion) compared to its inclusion an additional feature (early-fusion). The effects of other modeling decisions varied between institutions though performance increased with more compute-intensive models such as random forest. This work highlights the importance of considering methodological decision points in interpreting the impact on prediction performance when including PRS information in clinical models.

Preprint

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article

Analysis of OSHA inspection data with exposure monitoring and medical surveillance violations.

Occupational Safety and Health Administration (OSHA) inspection data from the Integrated Management Information System (IMIS) enforcement data base are presented for lead, ethylene oxide, and formaldehyde for fiscal years 1985, 1987, and 1989, and are discussed with emphasis on exposure monitoring or medical surveillance section violations. These data suggest that the exposure monitoring section of these standards is more commonly used to cite workplaces below these standards than is the medical surveillance section. Medical surveillance violations more commonly resulted in fines, but there were no differences in the magnitude of the fines for exposure monitoring or for medical surveillance violations. Implications of these findings are discussed.

Environmental Exposure