Search PubMedSearch

PubMed · 42567273

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

Abstract

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Explore related subjects

Keep this discovery

BibTeXRIS

Alessandra Vittorini Orgeas, Christoph W Sensen. 2026-08-07. A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.. https://doi.org/10.1016/j.jbiotec.2026.06.018

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (Ψ) represents one of the most abundant and conserved RNA modifications. Ψ provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of Ψ sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel Ψ site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA Ψ-site prediction. The Ψ modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA Ψ-site prediction. Meta-PseU offers a new framework for robust Ψ-site identification by using long sequences.

Pseudouridine

Genetic determinants of gestational diabetes mellitus in thai pregnant women: role of GCKR, CDKAL1, TCF7L2, NEDD1, and CMIP variants.

BACKGROUND: Gestational diabetes mellitus (GDM) has a high global prevalence and arises from complex interactions between genetic predisposition and environmental factors. GDM is associated with metabolic disturbances and chronic low-grade inflammation, both of which contribute to its pathogenesis. This study aimed to investigate the association between GDM and 135 single-nucleotide polymorphisms (SNPs) across 20 genes related to metabolic traits. METHODS: In this case-control study, 152 pregnant women with GDM and 684 pregnant women with normal glucose tolerance (NGT) who underwent antenatal examination at Siriraj Hospital, Bangkok, were enrolled. Clinical data and blood samples were collected from all participants. Genomic DNA was isolated and subjected to whole-genome sequencing using the DNBSEQ-T7RS high-throughput sequencing platform. Genotype analyses were performed using R software, and haplotype analyses were conducted using the online SNPStats software. RESULTS: After adjusting for maternal age and pre-pregnancy body mass index, polymorphisms in TCF7L2 (rs34872471, rs7901695, rs4506565, rs7903146, rs12243326, and rs12255372), NEDD1 (rs10431408, rs11830756, rs249579, rs249585, and rs4762339), CMIP (rs2306115 and rs201681534), CDKAL1 (rs4710942), GCKR (rs2293572 and rs2293571), and GCK (rs5883890) were significantly associated with the risk of GDM. Haplotype analysis demonstrated that the TCF7L2 rs12243326-rs12255372 CA haplotype was associated with a decreased risk of GDM (OR = 0.44, 95% CI: 0.23-0.81), while the NEDD1 rs249579-rs249585-rs4762339 GGT haplotype was associated with an increased risk of GDM (OR = 1.40, 95% CI: 1.08-1.82). CONCLUSIONS: These findings suggest that genetic variations in TCF7L2, NEDD1, CMIP, CDKAL1, GCK, and GCKR contribute to GDM susceptibility in the Thai population.

Humans

Making waves: toward systems-level interpretation of hormonal and endogenous biomarkers in wastewater-based epidemiology.

Wastewater-based epidemiology (WBE) has proven invaluable for population health monitoring, most notably during the COVID-19 pandemic. Yet current WBE largely relies on exogenous markers such as drugs, pathogens, and their metabolites, limiting surveillance to what communities are exposed to. We argue for expanding WBE towards endogenous biomarkers, particularly hormones, which provide insights into physiological stress, metabolic function, and endocrine activity. Hormone-based WBE offers new opportunities to capture population-level biological responses to societal and environmental stressors, disasters, and chronic disease burdens at the community scale. This perspective outlines a systems-level framework for integrating hormonal signals in wastewater with clinical data, behavioral indicators, environmental factors, and digital markers to support more robust and context-aware public health surveillance. We highlight key technical considerations, interpretive challenges, and opportunities for translational pilot studies. By moving beyond exposure tracking toward more integrated interpretation of biological responses, hormone-informed WBE may contribute to more resilient, inclusive, and actionable public health infrastructure.

Humans