Search PubMedSearch

SEARCH · Search PubMed

Results for “Bibliometric mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

101 records · Page 3Linked to original sources

Smoking Cue Reactivity in Relation to Uncertain-Threat and Reward-Anticipation Networks: A Coordinate-Based fMRI Meta-Analysis.

BACKGROUND: Smoking is a concerning medical and social problem, yet how the brain links stress to continued smoking is still not well understood. This coordinate-based meta-analysis identified convergent activations for smoking cues, uncertain threat, and reward anticipation, and examined co-activation patterns to clarify whether smoking-cue activity in smokers relates to threat and reward activity in non-addicted controls. METHODS: We conducted a coordinate-based activation likelihood estimation (ALE) meta-analysis of 102 fMRI studies (N = 3,068), including 30 studies on smoking cue reactivity (n = 945), 48 on reward anticipation (n = 1,413), and 24 on uncertain threat processing (n = 1,621). We performed single, conjunction, and contrast analyses, followed by meta-analytic connectivity modeling (MACM) of key regions. RESULTS: Single analysis revealed smoking engaged bilateral ACC (-1.1, 46.2, -1.1), uncertain threat engaged bilateral insula (left = -33.6, 22, 4.3, right = 41.3, 22, 1), reward anticipation activated thalamus (0.9, 1.3, -3.1) and medial frontal gyrus (2.7, 6.6, 52.1). Conjunction and contrast analyses showed shared or unique activation for each task in its respective regions. MACM showed ACC co-activation with thalamus and medial frontal gyrus, while insula co-activated with ACC and inferior frontal gyrus. CONCLUSIONS: Each process converged in a separate region, with no overlap between the smoking-cue map and either the threat or reward map. The ACC nonetheless co-activated with reward-related regions and shared network membership with the threat-related insula. On this basis we hypothesize a shift in motivation from stress-driven reward toward cue-driven craving, to be tested within subjects, and identify candidate neuromodulation targets for preventing stress-precipitated relapse.

fMRI

Identifying stakeholder behaviors for competency-based pharmacy education: A stage 1 behavior change wheel analysis.

INTRODUCTION/OBJECTIVES: Competency-Based Pharmacy Education (CBPE) is a strategic priority for preparing graduates to meet evolving healthcare needs. However, efforts to implement CBPE can stall due to behavioral challenges among faculty, administrators, preceptors, and learners. This study aimed to apply Stage 1 of the Behavior Change Wheel (BCW) to identify stakeholder-specific behaviors and associated determinants needed to implement the five core components of CBPE. METHODS: A multi-method approach grounded in the BCW, the Capability, Opportunity, Motivation - Behavior (COM-B) model, and the Theoretical Domains Framework (TDF) was used. Data were gathered through (1) targeted literature review; (2) structured focus groups with competency-based education experts and pharmacy education stakeholders; and (3) an iterative consensus process. Behaviors were mapped to the five CBPE components: (1) defined competencies, (2) developmental progression, (3) tailored instruction, (4) authentic experiential learning, and (5) programmatic assessment, and then mapped to COM-B and TDF constructs. RESULTS: Over fifty stakeholder-specific behaviors were identified and specified across the CBPE framework. This revealed shared barriers such as limited instructional design knowledge (psychological capability), insufficient assessment of infrastructure (physical opportunity), and misaligned professional identity (reflective motivation). Key TDF domains included knowledge, environmental context, beliefs about capabilities, and professional roles. The behavioral problem statements, specifications, and determinants were identified to support future intervention planning. CONCLUSION: This Stage 1 analysis provides a behaviorally grounded foundation for CBPE implementation by identifying stakeholder behaviors and conditions that enable change. These findings will inform the development of readiness-to-change assessments and targeted interventions (BCW Stages 2 and 3), supporting scalable and sustainable CBPE transformation in pharmacy education.

Education, Pharmacy

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (Ψ) represents one of the most abundant and conserved RNA modifications. Ψ provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of Ψ sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel Ψ site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA Ψ-site prediction. The Ψ modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA Ψ-site prediction. Meta-PseU offers a new framework for robust Ψ-site identification by using long sequences.

Pseudouridine

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny

Emerging Principles in Spatial Functional Genomics.

Spatial transcriptomic and proteomic atlases have enabled mapping of gene programs within intact tissues, but these measurements remain largely descriptive and do not define the mechanisms controlling tissue biology. Pooled CRISPR screening provides scalable causal interrogation of gene function but remains largely confined to dissociated systems that lack spatial context. In vivo spatial functional genomics (SFG) bridges these approaches by integrating genetic perturbations with in situ transcriptomic and proteomic readouts to measure gene function within intact tissue ecosystems. By preserving spatial organization, SFG enables interpretation of perturbations through effects on cell-cell interactions, diffusible signals, multicellular niches, and tissue architecture. Here, we outline key design axes of SFG: perturbation strategy, barcoding strategy, and phenotypic readout. We discuss computational challenges, including spatial autocorrelation, neighborhood dependence, and context-aware null modeling, and highlight how SFG reveals non-cell-autonomous, architecture-dependent mechanisms of gene function, advancing toward predictive models of tissue organization and gene function.

Genomics

Single nucleus multiomics reveals an early inflammatory response to high-fat diet in mouse islets.

In periods of sustained hyper-nutrition, pancreatic β-cells undergo functional compensation through transcriptional upregulation of gene programs driving insulin secretion. This adaptation is essential for maintaining systemic glucose homeostasis and metabolic health. Using single nuclei multiomics, we have mapped the early transcriptional adaptive mechanisms in murine islets of Langerhans exposed to high-fat diet (HFD) for 1 and 3 wk. We show that β-cells exhibit the largest transcriptional response to HFD, characterized by early activation of pro-inflammatory eRegulons and down-regulation of β-cell identity genes, particularly in a distinct subset of β-cells. These observations extend to humans, where the prevalence of an β-cells with a high inflammatory signature is increased in diabetes. Collectively, these observations point to cellular crosstalk through pro-inflammatory signaling as a central and early driver of β-cell dysfunction that limits the compensatory capacity of β-cells, which is closely linked to the development of diabetes.

Animals

Systematic evaluation of one-dimensional-to-two-dimensional near-infrared spectroscopy transformations with deep learning for quantifying coconut sap adulteration.

Near-infrared (NIR) spectroscopy have limitations when combined with deep learning (DL) algorithms because they rely on low-dimensional datasets. Therefore, we investigated the potential of transforming one-dimensional (1D) NIR spectra into two-dimensional (2D) spectrograms using synchronous and asynchronous techniques and the continuous wavelet transform (CWT) and their effectiveness by integrating with DL for detecting adulteration in coconut sap. NIR spectra (12,500-4000 cm-1) were collected from binary mixtures (0%-100%;w/w). The performance of all DL (convolutional neural networks-CNN, AlexNet and ResNet) models was compared with that of partial least squares (PLS). The models were ranked in the mentioned order based on their performances: 2D-CWT > 2D-asynchronous > 2D-synchronous > 1D/2D-PLS. The important features of the best model can be explained and visualized using gradient-weighted-class-activation-mapping. The findings highlight that the 1D-to-2D NIR data transformation combined with DL is a highly robust approach because it addresses the feature representation gap in NIR data and effectively captures the spatial-spectral correlations.

Spectroscopy, Near-Infrared

A chromosome-level, haplotype-resolved genome assembly for the barn owl, Tyto alba.

Recent advances in long-read sequencing have enabled near telomere-to-telomere (T2T) assemblies across diverse taxa. However, avian genomes remain challenging due to numerous microchromosomes, small, typically < 20Mb, DNA molecules that are gene-, GC-, and repeat-rich. As a consequence, microchromosomes are often missing from genome assemblies. Here, we present a chromosome-level, haplotype-resolved genome assembly for the Western barn owl (Tyto alba). Using a trio-binning strategy with Illumina parental reads combined with PacBio HiFi and Oxford Nanopore Technologies data, we generated two phased contig sets. These were scaffolded into 40 linkage groups using a linkage map. Comparative analyses identified unplaced HiFi scaffolds corresponding to microchromosomes, which we integrated into six additional microchromosomes using long reads information. The two assemblies present 46 chromosomes, matching the karyotype of the species. They exhibit strong synteny between parental haplotypes, except for a &#x223c;38 Mb complex region on chromosome 7 containing nested inversions. This high-quality reference provides a haplotype-resolved and chromosome-level genome for Strigiformes, enabling fine-scale studies of structural variation and avian genome evolution.

Tyto alba

Deciphering CD8+ T cell exhaustion in human cancers through single-cell and spatial transcriptomics.

Exhausted CD8+ T cells (Tex) within the tumor microenvironment (TME) represents a critical barrier limiting anti-tumor immune responses. Tex cells are characterized by upregulated inhibitory immune checkpoint receptors, reduced cytotoxicity, and functional heterogeneity. Their genomic features and regulatory networks remain poorly defined, and only a minority of patients respond to immune checkpoint blockade (ICB) therapy. Single-cell RNA sequencing (scRNA-seq), through high-resolution transcriptomic profiling, has revealed diverse Tex subpopulations, identified subpopulation-specific marker genes and regulatory pathways. Spatial transcriptomics has further mapped the spatial distribution of Tex and their interaction networks with immune cells, tumor cells, and stromal cells, elucidating the impact of spatial heterogeneity on Tex functionality. Current studies indicate that the exhausted state of Tex is dynamic and modifiable, with functional differences among subpopulations closely associated with tumor progression and therapeutic response. However, the genomic characteristics, epigenetic regulation, and spatial interaction mechanisms of Tex require further exploration. This review summarizes recent advances in high-resolution omics technologies for precisely dissecting Tex heterogeneity, functional features, and interactions with other cells. It emphasizes the central value of optimizing Tex-targeted tumor immunotherapy strategies, providing theoretical foundations and directional guidance for developing more effective anti-tumor immunotherapies.

Humans

Cross-species variant-to-function analyses implicate MEIS1 in conferring sleep abnormalities and impaired cerebellar development.

Genome-wide association studies (GWAS) have identified numerous loci for insomnia, yet functional validation of effector genes remains limited because most risk variants lie in noncoding regions, and the true causal gene is not known. Here, we use prior human cell-based variant-to-gene mapping to nominate six insomnia effector genes and test them in zebrafish, a tractable diurnal vertebrate model well suited for sleep phenotyping. Our CRISPR-based behavioral screening identifies the MEIS1 ortholog, meis1b, as a regulator of sleep maintenance, with crispants displaying impaired nighttime-specific sleep maintenance and increased sleep latency. Comparative chromatin analyses reveal conserved regulatory architecture spanning the human insomnia-associated locus and selectively implicate meis1b, whereas the duplicated ohnolog meis1a was dispensable. Developmental profiling further shows that meis1b is expressed in cerebellar granule progenitors, paralleling human MEIS1 expression, and that its disruption impairs cerebellar development. Together, these findings establish zebrafish as an efficient vertebrate platform for functional interrogation of GWAS candidates and support an evolutionarily conserved cerebellar role for MEIS1 in sleep maintenance.

Animals

Ramu stunt virus genome reveals previously unreported segments and nucleocapsid domain duplication in Mechlorovirus.

Ramu stunt virus (RmSV), a member of the genus Mechlorovirus within the family Phenuiviridae, was previously described as a six-segmented RNA virus infecting sugarcane. In this study, we re-examined type material and additional isolates using high-throughput sequencing and RT-PCR validation, revealing that RmSV possesses a nine-segmented genome, making it the largest reported in the Phenuiviridae. This expanded architecture includes duplicated RNA segments (RNA 2a and RNA 2b) encoding nucleocapsid-like proteins and two novel segments (RNA 7 and RNA 8). Comparative analysis showed that RNA 2a and 2b share about 84% amino acid identity, while RNA 5 encodes a third nucleocapsid homolog, indicating unprecedented domain redundancy. Structural modeling confirmed that all three nucleocapsid proteins maintain a conserved fold despite low sequence identity, with electrostatic mapping suggesting differential RNA-binding potential. Additionally, RNA 6 encodes a hypothetical protein structurally similar to the rice stripe virus disease-specific S-protein, implicating a role in symptom development. Transcript abundance analysis revealed RNA 6 as the most highly expressed segment across isolates. These findings revise the genomic composition of RmSV, highlight mechanisms of genome plasticity and adaptive evolution in plant-infecting bunyaviruses, and underscore practical implications for diagnostic assay design, resistance breeding, and biosecurity surveillance.

Genome, Viral

Quantitative assessment of the fingerprint evidential value using machine learning.

Fingerprints as physical evidence have long supported criminal investigation and adjudication. In practice, however, fingerprint identification relies mainly on examiners' experience. Furthermore, expert opinions tend to be categorical, even though the opinions with the same conclusion could differ substantially in evidential strength. To quantitatively assess fingerprint evidential value, this study proposes a machine learning-based framework as an interpretable decision-support tool. A lightweight residual one-dimensional convolutional neural network was constructed, incorporating channel recalibration and a similarity-driven attention mechanism to learn adaptive contribution weights for different matched minutiae (minutiae for short). Controlled experiments revealed that the predicted evidential value increased with the number of minutiae and was significantly influenced by the quality of minutiae. With 10 minutiae, the mean predicted scores were 4.49, 7.00, and 9.09 for blurred, moderately blurred, and clear minutiae, respectively. Multiple regression analysis indicated that replacing a pair of blurred minutiae with a pair of clear minutiae increased the score by 0.492, whereas replacing it with a pair of moderately blurred minutiae increased the score by only 0.216. By mapping predicted scores to graded levels of evidential strength, the framework contributes to a paradigm shift from categorical expert opinions to graded ones, helping courts evaluate fingerprint evidence more scientifically.

Humans

Plant cis-regulatory grammar: Decoding the multidimensional code of transcriptional regulation for programmable crop engineering.

Cis-regulatory elements (CREs) orchestrate the spatiotemporal precision of gene expression that underlies plant development, adaptation, and domestication. Decoding the cis-regulatory grammar of plant genomes remains a central challenge in modern biology, with profound implications for programmable crop engineering. Here, recent conceptual and technological advances are synthesized to reshape our understanding of plant CREs. This review first argues that CRE function is not only an intrinsic property of DNA sequence alone but also emerges from a multidimensional context, including chromatin accessibility, histone modifications, three-dimensional genome topology, and cell type-specific regulatory landscapes. Furthermore, the convergence of single-cell epigenomics, high-throughput functional assays, and CRISPR-based dissection has begun to unravel this contextual grammar, revealing the computational principles governing transcriptional regulation. Critically, we propose that artificial intelligence (AI) platforms are catalyzing an ongoing transition from descriptive discovery to predictive engineering, wherein these platforms outperform natural evolution in designing synthetic CREs. Finally, a roadmap is outlined toward a plant regulatory grammar foundation model, which will enable truly predictive engineering of gene expression when fine-tuned for specific tasks. Collectively, the integration of single-cell resolution maps, precise genome editing, AI-driven design, and regulatory-compliant delivery systems promises to transform our ability to reprogram plant gene regulation for next-generation agriculture, bridging the gap between foundational regulatory biology and tangible crop improvement.

artificial intelligence

AI-driven snapshot hyperspectral imaging for on-line sorting systems in food industry: From real-time sensing to intelligent decision-making.

High-throughput food sorting requires rapid, non-destructive detection of external defects, foreign materials, and internal quality attributes in heterogeneous food matrices. Conventional scanning hyperspectral imaging may suffer from motion-induced spatial-spectral mismatches, whereas snapshot hyperspectral imaging (S-HSI) captures spectral images within a single integration time. However, its advantage is limited by trade-offs in resolution, signal-to-noise ratio (SNR), reconstruction uncertainty, and calibration stability, which are further amplified by variable tissue structure, surface reflection, moisture, and fat distribution in foods. This review critically examines artificial intelligence (AI)-driven S-HSI for on-line food sorting within a sensing-representation-decision-execution framework. Compact architectures are compared according to their physical constraints, food-sorting suitability, and ability to support mapping between spectral responses and physicochemical quality attributes. AI strategies are reviewed for spectral reconstruction, image restoration, spatial-spectral representation, band selection, uncertainty-aware decision-making, and edge implementation. AI can partially compensate for snapshot-specific limitations, but current evidence remains largely limited to laboratory or prototype studies. Future work should link system performance to food safety and quality outcomes by reporting throughput, decision latency, calibration drift, missed-detection risk, false-rejection cost, and closed-loop sorting success.

Hyperspectral Imaging

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans

Development of the zebrafish foveal analogue: a quantitative atlas of high-acuity zone growth and retinal regionalisation.

The vertebrate retina contains specialised regions for high-acuity vision, exemplified by the human fovea and its zebrafish analogue, the high-acuity zone (HAZ). Despite the widespread use of zebrafish to model retinal disease, a stage-resolved quantitative reference describing normal eye, photoreceptor layer (PRL) and lens growth has been lacking. Here, we apply contrast-enhanced micro-computed tomography (micro-CT) to construct the first three-dimensional micro-CT normative atlas of wild-type zebrafish eye development across five larval stages [3, 5, 7, 10 and 18&#x2005;days post-fertilisation (dpf)], mapping circumferential PRL thickness, eye and lens morphology, and compartment growth rates. Regional PRL thickening within the temporo-ventral region of the expected HAZ emerged by 5&#x2005;dpf and was sustained by a localised redistribution of growth, persisting and extending towards the optic nerve through 18&#x2005;dpf. The PRL, lens and eye grew through four phases, alternating between disproportionate PRL expansion and coordinated growth, while the eye remodelled from a nasal-dominant to a temporo-ventral-dominant form. This regional specialisation was protracted relative to gross ocular growth and could proceed independently of it, paralleling the extended postnatal maturation of the human fovea. This atlas provides a quantitative baseline for distinguishing disease-induced changes from normal variation, supporting zebrafish models of foveal hypoplasia and related disorders.

Animals

Integrated analysis uncovers exogenous induction and molecular regulation of erinacine A accumulation in Hericium erinaceus.

Erinacine A, a cyathane-type diterpenoid mainly from Hericium erinaceus mycelia, exhibits prominent neurotrophic and neuroprotective activities, making it a promising candidate for managing neurodegenerative diseases. However, its low abundance and unclear genetic regulatory mechanisms hinder its application as a nutraceutical. This study aimed to decipher its regulatory mechanisms and enhance production. Four exogenous inducers were screened, with salicylic acid (SA) and ergosterol (ERG) significantly increasing erinacine A content by 62.21% and 146.70% at 20 days, respectively. Transcriptome and WGCNA of inducer-treated sample identified darkorange and magenta modules associated with erinacine A biosynthesis, with the eri gene cluster enriched in the darkorange module and eriG and eriF as hub genes. Forward genetic analysis via QTL mapping of the HeD127 dikaryon population revealed significant phenotypic variation in erinacine A content (0.341-13.085&#x202f;mg/g) and identified two loci (erA-1 and erA-2) explaining 18.63% of phenotypic variation. Integrating these forward and reverse genetic analyses revealed that salicylic acid and ergosterol synergistically regulate core carbon metabolic pathways to augment acetyl-CoA supply for the mevalonate pathway, suppressed competitive metabolism, enhanced diterpene skeleton construction and structural modification. These results deepen our understanding of the genetic and molecular basis governing accumulation of erinacine A, and facilitate its application in neuroprotective pharmaceuticals.

Diterpenes

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers