Search PubMedSearch

SEARCH · Search PubMed

Results for “Public data reuse”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

8 recordsLinked to original sources

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics

Computational metabolomics at scale: from open data to insight.

Metabolomics data are currently generated at scale thanks to the evolution of technologies that have led to marked improvements in the number of metabolites detected, spanning all chemical classes. These data are increasingly submitted to public repositories for data reuse, integration, and interpretation. Despite the availability of public resources and associated computational tools, the field still lacks a widely adopted, consistent data and analytics infrastructure capable of transforming this wealth of information into scientific insight. Indeed, the metabolomics field is just now scratching the surface of being able to harness the power of new computational technologies. In this review, we summarize discussions from the "Dagstuhl-Seminar 24181 Computational Metabolomics: Towards Molecules, Models, and their Meaning" with a focus on public data availability, open data standards, data and knowledge integration, and education. Our goal is to raise awareness and adoption of the latest open science resources while highlighting key areas needing further development.

Metabolomics

The European Health Data Space and the Secondary Use of Sensitive Health Data.

INTRODUCTION: The European Health Data Space (EHDS) is one of the European Union's most ambitious data-governance projects. It aims to create a common framework through which electronic health data can be accessed and reused across Member States for care, research, innovation, policy, and public-interest purposes. Its practical viability depends not only on digital infrastructure, but also on legal, ethical, and organisational harmonisation, particularly for genetic and genomic data. METHODS: This paper examines the EHDS with emphasis on the secondary use of health data. It reviews the EHDS institutional architecture, discusses Finland's Findata as a national model for structured access, and analyses challenges for data holders and data donors, including interoperability, governance burdens, privacy protection, residual re-identification risk, and genomic-data sensitivity. RESULTS: A cross-border cancer-genomics case study shows that the EHDS can streamline data discovery and the routing of access requests, but does not by itself eliminate legal fragmentation, heterogeneous ethics review, and consent-related barriers. DISCUSSION: Effective implementation will require harmonisation beyond infrastructure, including clearer consent standards, more consistent ethics procedures, interoperable metadata, and proportionate safeguards for genomic data.

Electronic Health Records

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49 h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning

The Continuity Trap in Data Science Health Research.

Secondary use is now the ordinary condition of data science health research rather than an exception to it. Electronic health records collected for clinical care become prediction tools and inputs for generative AI; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable, indefinitely distributable cell lines. Governance has responded by emphasizing verifiable instruments such as provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, and they answer real questions about lineage, privacy, institutional responsibility, and accountability, but they are not sufficient to establish that a present use remains ethically justified. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We then define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into the remaining domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal rereview; it is a cross-domain inference error that can arise even in careful, good-faith review. We distinguish it from proxy closure, of which it is a continuity-specific subtype, and from Goodhart's and Campbell's laws, which describe how measures degrade once they become targets. We operationalize ethical continuity across 4 domains: provenance, semantics, authorization, and relational standing, developed in our Representational Veracity framework, and we show that these domains can diverge as data are linked, transformed, modeled, and redeployed. We identify the institutional mechanisms-provenance privilege, descriptor sedimentation, authorization fossilization, and community effacement-that cause auditable signals to be overread, and we examine how the US Health Insurance Portability and Accountability Act (HIPAA) of 1996, the General Data Protection Regulation, the European Health Data Space, US Food and Drug Administration guidance, the US National Institute of Standards and Technology (NIST) AI Risk Management Framework, and federated-learning governance can reduce risk while still inducing continuity traps. We apply the framework to consent and nonconsent settings, including public health, immunization, syndromic, and wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: rather than rereviewing every reuse, investigators and reviewers should identify the weakest continuity domain at the present data stage and impose a domain-matched safeguard, recorded in a short continuity statement. This reframing is intended for the committees, repositories, funders, and governance bodies that decide whether reuse may proceed, and it matters most in cross-border and low-resource settings. Provenance should begin ethical review; it should not end it.

Data Science

Application of qualifying variants for genomic analysis.

MOTIVATION: Qualifying variants (QVs) are genomic alterations selected by defined criteria within analysis pipelines. Although crucial for both research and clinical diagnostics, QVs are often seen as simple filters rather than dynamic elements that influence the entire workflow. In practice these rules are embedded within pipelines, which hinders transparency, audit, and reuse across tools. A unified, portable specification for QV criteria is needed. RESULTS: Our aim is to embed the concept of a "QV" into the genomic analysis vernacular, moving beyond its treatment as a single filtering step. By decoupling QV criteria from pipeline variables and code, the framework enables clearer discussion, application, and reuse. It provides a flexible reference model for integrating QVs into analysis pipelines, improving reproducibility, interpretability, and interdisciplinary communication. Validation across diverse applications confirmed that QV based workflows match conventional methods while offering greater clarity and scalability. AVAILABILITY AND IMPLEMENTATION: The source code and data are accessible at the Zenodo repository https://doi.org/10.5281/zenodo.17414191. Manuscript files are available at https://github.com/DylanLawless/qvApp2025lawless. The QV framework is available under the MIT licence, and the dataset will be maintained for at least two years following publication.

Genomics

Whole-genome surveillance supports hazard profiling of Escherichia coli lineages in recycled water treatment systems.

UNLABELLED: The use of treated wastewater is increasingly important for sustainable water management under a changing climate, yet conventional monitoring based on Escherichia coli enumeration provides limited insight into strain diversity and associated public health hazards. Here, we applied longitudinal whole-genome sequencing (WGS) to 180 E. coli isolates collected across the treatment continuum of a recycled water facility, from influent to final effluent. Genomic analysis revealed extensive strain-level heterogeneity, comprising 88 sequence types across eight phylogroups, with greater diversity in influent than in treated effluent. Phylogenetic comparisons with contextual Australian genomes indicated clustering with strains associated with companion animals, wild birds, humans, and livestock, suggesting multiple potential source reservoirs rather than a single dominant origin, although source contributions were not definitive. Despite a >90% reduction in total E. coli loads, isolates recovered from upstream and downstream stages exhibited broadly comparable virulence factor and antimicrobial resistance gene (ARG) profiles, suggesting that, within the cultured isolate collection, reductions in abundance exceeded shifts in genomic composition. To assess operational relevance, we prototyped a genomics-informed hazard framework integrating virulence determinants, ARGs, plasmid-associated mobility, and reuse-specific exposure context. Using this framework, 92.8% of isolates were classified as low hazard, and 7.2% as moderate hazard, with no isolates meeting criteria for high or critical hazard classifications. These findings demonstrate that genomic profiling of indicator organisms can reveal population structure and hazard heterogeneity not captured by conventional enumeration alone, and can provide a practical basis for incorporating genomic information into hazard-informed monitoring of recycled water systems. IMPORTANCE: Routine recycled water monitoring relies largely on culture-based E. coli counts, which indicate regulatory compliance but provide limited insight into strain diversity, persistence, and genomic characteristics relevant to public health. Using longitudinal whole-genome sequencing, we show that genetically distinct E. coli lineages, including isolates carrying combinations of virulence and antimicrobial resistance determinants, can persist through advanced treatment despite substantial reductions in overall E. coli loads. While most isolates were classified as low genomic hazard and no high- or critical-hazard isolates were detected, these findings demonstrate that conventional enumeration alone cannot distinguish between genetically diverse lineages with differing hazard potential in highly treated systems. By integrating genomic data into a hazard classification framework, this study demonstrates an applied approach to contextualize E. coli detections and distinguish low-risk background populations from isolates with elevated genomic hazard profiles. This work supports the use of genomic profiling of indicator organisms to improve surveillance, inform treatment performance assessment, and enable more risk-based management of recycled water systems.

Escherichia coli

A dataset of estimated heterozygous individual and carrier couple frequencies for pan-ancestry carrier screening.

The data described in this publication supported the development and evaluation of pan-ancestry reproductive carrier screening panels for autosomal recessive (AR) and X-linked (XL) conditions. Raw data included combined sets of DNA variants in 1,350 AR/XL genes obtained from the ClinVar and gnomAD databases. The dataset enabled calculations of positive yield for individuals and couples across both ancestry-specific and pan-ancestry, optimised "Goldilocks"-ranked gene panels, addressing population-specific variations in the frequencies of heterozygous individuals and carrier couples. The positive yield analysis offered a performance metric for carrier screening panels, facilitating the modeling of screening performance for panels of varying sizes and composition and providing resources for optimizing panel content to ensure equity across underrepresented genetic ancestries The dataset can support ongoing research into the equitable application of carrier screening and offers significant reuse potential for refining population genetic screening practices, validating computational models, and developing frameworks to update carrier screening panels in alignment with evolving genomic data, including in underrepresented and minority populations.

Carrier screening