Search PubMedSearch

SEARCH · Search PubMed

Results for “Data Science”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The Continuity Trap in Data Science Health Research.

Secondary use is now the ordinary condition of data science health research rather than an exception to it. Electronic health records collected for clinical care become prediction tools and inputs for generative AI; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable, indefinitely distributable cell lines. Governance has responded by emphasizing verifiable instruments such as provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, and they answer real questions about lineage, privacy, institutional responsibility, and accountability, but they are not sufficient to establish that a present use remains ethically justified. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We then define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into the remaining domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal rereview; it is a cross-domain inference error that can arise even in careful, good-faith review. We distinguish it from proxy closure, of which it is a continuity-specific subtype, and from Goodhart's and Campbell's laws, which describe how measures degrade once they become targets. We operationalize ethical continuity across 4 domains: provenance, semantics, authorization, and relational standing, developed in our Representational Veracity framework, and we show that these domains can diverge as data are linked, transformed, modeled, and redeployed. We identify the institutional mechanisms-provenance privilege, descriptor sedimentation, authorization fossilization, and community effacement-that cause auditable signals to be overread, and we examine how the US Health Insurance Portability and Accountability Act (HIPAA) of 1996, the General Data Protection Regulation, the European Health Data Space, US Food and Drug Administration guidance, the US National Institute of Standards and Technology (NIST) AI Risk Management Framework, and federated-learning governance can reduce risk while still inducing continuity traps. We apply the framework to consent and nonconsent settings, including public health, immunization, syndromic, and wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: rather than rereviewing every reuse, investigators and reviewers should identify the weakest continuity domain at the present data stage and impose a domain-matched safeguard, recorded in a short continuity statement. This reframing is intended for the committees, repositories, funders, and governance bodies that decide whether reuse may proceed, and it matters most in cross-border and low-resource settings. Provenance should begin ethical review; it should not end it.

Data Science

Practicing Data Science in Interactive Notebooks.

The Jupyter Notebook is a platform for interactive computing that displays code and results in the same browser, making it valuable for teaching, prototyping, data analysis, and collaboration. Its explicit and transparent structure greatly reproducibility while its backend server supports flexible deployment. In the past few years, Jupyter notebooks and similar tools have become increasingly popular. In this chapter, we will review key aspects of data analysis in a cloud environment and demonstrate common tasks for analyzing metabolomics data using template notebooks. This is an accompaniment to the basic bioinformatics tools and essential data science toolkit introduced in the first edition.

Software

The phenotype-genotype reference map: Improving biobank data science through replication.

Population-scale biobanks linked to electronic health record data provide vast opportunities to extend our knowledge of human genetics and discover new phenotype-genotype associations. Given their dense phenotype data, biobanks can also facilitate replication studies on a phenome-wide scale. Here, we introduce the phenotype-genotype reference map (PGRM), a set of 5,879 genetic associations from 523 GWAS publications that can be used for high-throughput replication experiments. PGRM phenotypes are standardized as phecodes, ensuring interoperability between biobanks. We applied the PGRM to five ancestry-specific cohorts from four independent biobanks and found evidence of robust replications across a wide array of phenotypes. We show how the PGRM can be used to detect data corruption and to empirically assess parameters for phenome-wide studies. Finally, we use the PGRM to explore factors associated with replicability of GWAS results.

Humans

Using cancer profiles to identify synthetic lethal therapeutic targets and predictive biomarkers in cancer gene dependency data.

MOTIVATION: Large scale loss-of-function screens utilising CRISPR or siRNA can provide profound insights into the importance of individual genes for the survival of a cancer cell and can drive the identification of therapeutic targets and biomarkers, and the development of targeted drugs. However, the analysis of these data and the substantial bodies of metadata that relate to them, is technically challenging and typically requires substantial expertise in data science and computer coding. RESULTS: To facilitate the analysis of cancer gene dependency data by cancer biologists and clinical scientists, we have developed DepMine-a computational toolkit providing a powerful system for framing complex queries relating cancer gene dependency to the underlying genetic changes that occur in cancer cells. DepMine identifies synthetic lethal relationships between putative target genes and complex 'cancer profiles' built from user-specified combinations of mutations, copy-number variation, and expression levels, and can refine these to optimal biomarker definitions for target dependency. AVAILABILITY: The Python implementation of DepMine and associated data files can be obtained at https://github.com/UOSbioinformaticslab/depmine and is free to academics and Not-For-Profit organisations. The DepMine release referenced in this paper is archived as DOI: 10.5281/zenodo.19570601.

Humans

EDAmame: interactive exploratory data analyses with explainable models.

SUMMARY: Complex tabular datasets comprising many diverse features can require specific expertise to interpret, posing a barrier to researchers with minimal data science experience. EDAmame is an interactive tool that simplifies initial analysis and visualization of these datasets, providing insights into data quality and feature relationships. By leveraging open-source machine learning frameworks in R, EDAmame allows researchers to perform effective exploratory data analysis without command-line or coding requirements. AVAILABILITY AND IMPLEMENTATION: A limited online version can be accessed at https://edamame.org.au/ or can be downloaded from https://doi.org/10.5281/zenodo.15356492. The app is developed in R Shiny and implements tidyverse and tidymodels packages.

Machine Learning

Giotto Suite: a multiscale and technology-agnostic spatial multiomics analysis ecosystem.

Emerging spatial multiomics technologies provide an increasingly large amount of information content at multiple scales. However, it remains challenging to efficiently represent and harmonize diverse spatial datasets. Here we present Giotto Suite, a suite of modular packages that provides scalable and extensible end-to-end solutions for multiscale and multiomic data analysis, integration and visualization. At its core, Giotto Suite is centered around an innovative data framework, allowing the representation and integration of spatial omics data in a technology-agnostic manner. Giotto Suite integrates molecular, morphology, spatial and annotated feature information to create a responsive and flexible workflow, as demonstrated by applications to several state-of-the-art spatial technologies. Furthermore, Giotto Suite builds upon interoperable interfaces and data structures that bridge the established fields of genomics and spatial data science in R, thereby enabling independent developers to create custom-engineered pipelines. As such, Giotto Suite creates an immersive and multiscale ecosystem for spatial multiomic data analysis.

Genomics

The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data.

The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.

Journal Article

Effects of phytosterols supplementation on hepatic lipid metabolism and metabolic outcomes in obese rodent models: a systematic review and meta-analysis.

This study aimed to synthesize and quantitatively assess the available evidence on the effects of phytosterol supplementation on hepatic lipid metabolism and obesity-related metabolic outcomes in obese rodent models, integrating biochemical, histological, and molecular evidence. A systematic search was conducted in electronic databases (PubMed, EMBASE, and Web of Science). Data on study design, population, intervention, outcomes, and risk of bias were extracted and analyzed. A quantitative meta-analysis was performed. Meta-analysis showed reductions in body weight, serum triglycerides, total cholesterol, LDL-C, VLDL-C, glucose, liver weight, hepatic cholesterol, hepatic triglycerides, and nonalcoholic fatty liver disease activity score. No significant changes were observed for adiposity index, HDL-C, insulin, or hepatic expression of PPARα, FAS, and SREBP1c. Conversely, CPT1A expression was significantly increased following PS supplementation. Subgroup analyses indicated that the beneficial effects on lipid and hepatic outcomes were generally consistent across rodent species (mice, rats, and hamsters), obesity induction models, and routes of administration, although the magnitude of responses varied between strains, with C57BL/6 mice showing more pronounced metabolic improvements. Additional analyses suggested that treatment duration and phytosterol composition may modulate specific outcomes, whereas dose-response meta-regression identified dose-dependent associations for serum and hepatic cholesterol, and PPARα expression in dietary supplementation studies. Overall, the available preclinical evidence suggests that phytosterol supplementation may improve several metabolic and hepatic outcomes in rodent models of obesity. However, the substantial heterogeneity across studies highlights the need for standardized experimental protocols and future clinical studies before these findings can be translated to human health.

Animals

AI In Leukemia Diagnostics: Complementing the Pathologist's Role.

Artificial intelligence (AI) is reshaping every stage of leukemia diagnostics, from digital morphology and multiparameter flow cytometry to next-generation sequencing, multi-omics analysis, and emerging computational frontiers such as quantum-inspired feature selection. This review outlines how contemporary AI tools can automate labor-intensive quantitation, flag diagnostically salient patterns, and standardize interpretation, while the pathologist or hematologist retains authority over validation, context-specific integration, and clinical decision-making. We present an illustrative "human-in-the-loop" workflow that embeds AI modules within current laboratory information systems, emphasizing points where expert oversight mitigates algorithmic bias and resolves discordant findings. We further map the validator-integrator role across morphology, flow cytometry, and genomic/multi-omic interpretation and provide practical training competencies and use cases for AI-assisted hematopathology. Beyond technical deployment, the article addresses the educational transformation required for sustainable adoption. Drawing on international competency frameworks, including the Digital Health Competencies in Medical Education Framework and recently proposed AI-specific Entrustable Professional Activities, we map core skills that future hematopathologists must master: data-science literacy, critical appraisal of AI outputs, and ethical governance. We highlight evaluated training models such as the Pathology Informatics Essentials for Residents curriculum, Stanford Artificial Intelligence in Machine and Imaging workshops, and College of American Pathologists bootcamps and propose integration strategies adaptable across resource settings. By pairing rigorous validation with targeted education, AI can elevate rather than eclipse the diagnostic role of the leukemia specialist, enabling more timely, reproducible, and personalized patient care.

Humans

Integrative Multiomics and Drug Sensitivity Profiling Reveal Potential Biomarkers and Therapeutic Strategies in Pediatric Solid Tumors.

UNLABELLED: Cure rates for childhood malignancies using established therapy protocols have increased to an average of 80% but have reached a plateau. Moreover, survival rates are particularly low for some pediatric tumors-such as high-risk group 3 medulloblastomas, osteosarcomas, Ewing sarcomas, high-risk neuroblastomas, and high-grade gliomas-and dismal for patients with relapsed malignancies. A functional drug response profiling platform for pediatric solid and brain tumors has been established within the INFORM program to identify patient-specific vulnerabilities and biomarkers and to unravel molecular mechanisms associated with drug response profiles for clinical translation. In this study, we performed a multiomics analysis using drug sensitivity profiles, as well as genomic and transcriptomic data, of 81 pediatric solid tumor samples. The integrative analysis suggested two multiomics signatures associated with drug sensitivity. One signature distinguished neuroblastoma samples with sensitivity to navitoclax, a BCL2 family inhibitor. A second signature was specific to a subset of Wilms tumors harboring the SIX1 (Q177R) hotspot mutation that displayed high expression of MGAM, PTPN14, STAT4, and KDM2B and high sensitivity to MEK inhibitors. A patient-specific causal interaction network analysis suggested possible molecular interactions between MEK inhibitors and the SIX1 mutation in Wilms tumor samples. In conclusion, the integration of drug sensitivity profiling and multiomics data revealed potential biomarkers that may be associated with drug sensitivity in pediatric solid tumors. Patient-specific causal interaction network analysis further elucidated the interaction between inhibitors and signature biomarkers, providing insights that may inform clinical translation. SIGNIFICANCE: The combination of multiomics analysis and drug sensitivity profiling identified two signatures related to drug sensitivity in pediatric solid tumors, contributing to the advancement of functional precision medicine and personalized treatment strategies. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Challenges and Opportunities in Analyzing Cancer-Associated Microbiomes.

The study of cancer-associated microbiomes has gained significant attention in recent years, spurred by advances in high-throughput sequencing and metagenomic analysis. Microbiome research holds promise for identifying noninvasive biomarkers and possibly new paradigms for cancer treatment. In this review, we explore the key computational challenges and opportunities in analyzing cancer-associated microbiomes (in tumor/normal tissues and other body sites, e.g., gut, oral, and skin), focusing on sequencing-driven strategies and associated considerations for taxonomic and functional characterization. The discussion covers the strengths and limitations of current analysis tools for identifying contamination, determining compositional bias, and resolving species and strains, as well as the statistical, metabolic, and network inferences that are essential to uncover host-microbiome interactions. Several key considerations are required to guide the choice of databases used for metagenomic analysis in such studies. Recent advances in spatial and single-cell technologies have provided insights into cancer-associated microbiomes, and Artificial Intelligence-driven protein function prediction might enable rapid advances in this field. Finally, we provide a perspective on how the field can evolve to manage the ever-growing size of datasets and generate robust and testable hypotheses. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Modeling Early-Onset Cancer Kinetics Reveals Changes in Underlying Risk and the Impact of Population Screening.

UNLABELLED: Recent studies have reported increases in early-onset cancer cases (diagnosed less than 50 years of age) and raised questions about whether the increase is related to earlier diagnosis from nonspecific medical tests as reflected by decreasing tumor-size-at-diagnosis (apparent effects) or actual increases in underlying cancer risk (true effects), or both. The classic Multistage Clonal Expansion (MSCE) model assumes cancer detection at the first malignant cell's emergence, although later modifications have included lag-times or stochasticity in detection to represent the delay in tumor detection. In this study, we introduced an approach to explicitly incorporate tumor-size-at-diagnosis in the MSCE framework accounting for improvements in cancer detection over time to distinguish between apparent and true increases in early-onset cancer incidence. The model was structurally identifiable and provided better parameter estimation than the classic model. The model was applied to colorectal, breast, and thyroid cancers to examine changes in cancer risk while accounting for detection improvements over time in three representative birth cohorts (1950-1954, 1965-1969, and 1980-1984). The analyses suggested accelerated carcinogenic events and shorter mean sojourn times (the average time from the first malignant cell emergence to cancer detection) in more recent cohorts. Furthermore, using this model to examine the screening impact on the incidence of breast and colorectal cancers, for which both have established screening protocols, provided results that align with well-documented differences in screening effects between these cancers. These findings underscore the importance of incorporating tumor-size-at-diagnosis in cancer modeling and support true increases in early-onset cancer risk in recent years for breast, colorectal, and thyroid cancers. SIGNIFICANCE: A model of early-onset cancer trends that distinguishes true risk from detection effects accurately captures cancer kinetics, trends in cancer progression, and the impact of screening, which could inform cancer prevention strategies. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Pithos - a scalable and secure data container for FAIR-compliant research data management in life sciences.

Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.

Biological Science Disciplines

Identification of CD55 as a downstream factor of EP4 receptor signaling in colorectal cancer cells.

Prostaglandin E2 (PGE2) signaling through the E-type prostanoid 4 (EP4) receptor has been implicated in the pathophysiology of colorectal cancer (CRC). We herein identified decay-accelerating factor, also known as CD55, as a novel CRC-associated downstream factor of the EP4 receptor. The integration of transcriptomic profiling of PGE2-stimulated HCA-7 human colon cancer cells with analyses of cancer genomic databases predicted CD55 as a potential EP4 receptor-regulated target. Inhibitor-based experiments showed the induction of CD55 after a PGE2 stimulation required the EP4 receptor and Gi protein in HCA-7 cells, whereas protein kinase A signaling was dispensable. In combination with a toxicogenomic database analysis, p38 mitogen-activated protein kinase (MAPK) was identified as the predominant effector connecting the EP4 receptor to CD55 upregulation. A single-cell RNA-seq re-analysis of human CRC tissues revealed CD55 upregulation and p38 MAPK-related gene set enrichment in epithelial cells expressing the EP4 receptor, suggesting that this induction mechanism may operate in a subset of epithelial cells in clinical specimens. Collectively, these results delineate a PGE2/EP4 receptor/Gi protein/p38 MAPK signaling axis that induces CD55 expression in HCA-7 cells and epithelial tumor cells, provide new mechanistic clues for understanding the regulation of complement regulatory molecule CD55 expression by prostaglandin signaling.

Humans

Scalable, open-access and multidisciplinary data integration pipeline for climate-sensitive diseases.

Climate-sensitive infectious diseases pose an important challenge for human, animal and environmental health and it has been estimated that over half of known human pathogenic diseases can be aggravated by climate change. While climatic and weather conditions are important drivers of transmission of vector-borne diseases, socio-economic, behavioural, and land-use factors as well as the interactions among them impact transmission dynamics. Analysis of drivers of climate-sensitive diseases require rapid integration of interdisciplinary data to be jointly analysed with epidemiological (including genomic and clinical) data. Current tools for the integration of multiple data sources are often limited to one data type or rely on proprietary data and software. To address this gap, we develop a scalable and open-access pipeline for the integration of multiple spatio-temporal datasets that requires only the declaration of the country and temporal range and resolution of the study. The tool is locally deployable and can easily be integrated into existing climate-disease-modelling applications. We demonstrate the utility of the tool for dengue modelling in Vietnam where epidemiological data are legally required to remain local. We include a pipeline for bias correction of climate data to enhance their quality for downstream modelling tasks. The Dengue Advanced Readiness Tools-Pipeline empowers users by simplifying complex download, correction, and aggregation steps, fostering data-driven discovery of relationships between infectious diseases and their drivers in space and time, and enhancing reproducibility in research. Additional modules and datasets can be added to the existing ones to make the pipeline extendable to use cases other than the ones presented here.

automated workflows

Use of Wearable Sensors in Angelman Syndrome: A Systematic Review.

BACKGROUND: Wearable sensors are a promising method for collecting clinical trial outcome data for people with Angelman syndrome (AS). However, there has yet to be a systematic probe into the ways in which wearable sensors have been successfully used in AS. The current study aims to provide a quantitative summary of wearable sensors used in AS, including contexts of use and psychometric properties, and to present key narrative highlights. METHOD: Literature searches were performed in three electronic databases: APA PsycInfo, PubMed and Web of Science Core Collection. Data items were categorized into four categories: sample characteristics, study methodological details, wearable sensor characteristics and psychometric properties assessed. Sample characteristics included sample size, age, biological sex, race/ethnicity and cognitive/developmental functioning. Study methodological details were subdivided into study design and setting. Wearable sensor characteristics included sensor type, placement site, means of attachment, assessed construct and sensor-related data loss. Psychometric properties assessed included reliability and validity of sensor-derived data. RESULTS: We identified 16 articles through our systematic review. Wearable sensors were used to study sleep (n = 10, 62.5%), language (n = 2, 12.5%), gait (n = 2, 12.5%), caregiver proximity (n = 1, 6.3%), EEG power (n = 1, 6.3%), and arousal (n = 1, 6.3%) in AS through actigraphs, vocalization recorders, inertial sensors, radio-frequency identification watches, wireless EEG caps, and functional near-infrared spectroscopy caps, respectively. Findings from these studies broadly indicate that wearable sensors are feasible, reliable and valid for assessing a range of behaviours relevant to AS. CONCLUSIONS: Wearable sensors are a promising solution to enhance assessments in AS. However, with the small extant literature characterized by small sample sizes and restricted focus on a few relevant features in AS, there remains ample opportunities to explore the use of wearable sensors in people with AS. Additional studies will better inform clinical decision-making and ultimately improve the lives of people with AS and their families.

Humans

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article