Search PubMedSearch

PubMed · 42178371

MarkerMatch: a proximity-based probe-matching algorithm for joint analysis of copy-number variants from different genotyping arrays.

Abstract

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to detect CNVs, which can be used in genetic association tests. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (those present on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has led to excessive reduction in overall sensitivity since arrays can have an undesirably low probe overlap. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4906 individuals genotyped across three different arrays, we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g. use of consensus probes only). We further demonstrate that MarkerMatch matches the CNV detection from current practice in terms of F1 score and PPV for larger CNVs. We also optimize MarkerMatch parameters, DMAX and Method, and find an optimal DMAX setting at 10 kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis. AVAILABILITY: The R package for MarkerMatch is available at: https://github.com/FranjoIM/MarkerMatch. The code used for analysis and implementation is available at: https://doi.org/10.5281/zenodo.18460979. The live notebook is available at https://fivankovic.notion.site/2026-markermatch.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Franjo Ivankovic, Dongmei Yu, James Shen, Lingyu Zhan, Maria Niarchou, Ariadne Kaylor, Laura Domènech, Tyne W Miller-Fleming, Luz M Porras, Paola Giusti-Rodríguez, Roel A Ophoff, Jeremiah M Scharf, Carol A Mathews. 2026-06-01. MarkerMatch: a proximity-based probe-matching algorithm for joint analysis of copy-number variants from different genotyping arrays.. https://doi.org/10.1093/bioinformatics%2Fbtag341

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

PScnv: personalized self-normalizing CNV detection with a hierarchical multi-phase framework.

MOTIVATION: Accurate detection of copy number variations (CNVs) from targeted panel sequencing remains challenging due to limited genomic coverage and pronounced sample-specific biases. Existing normalization strategies, including baseline-cohort, matched-control, and single-sample approaches, often struggle to balance noise suppression with adaptability, leading to inconsistent performance across heterogeneous samples. RESULTS: We present PScnv, a personalized self-normalizing framework for robust CNV detection from panel sequencing data. PScnv integrates a pre-built panel-of-normals (PoN) with sample-intrinsic stable chromosomes through ridge-regression normalization to generate individualized log2 ratio profiles with reduced systematic variation. CNVs are then identified using a hierarchical multi-phase segmentation pipeline incorporating z-score pre-partitioning, kernel-based correction, and circular binary segmentation. In 139 clinical tumor samples with orthogonal FISH validation at MET, ERBB2, and MTAP, PScnv showed improved accuracy and robustness over existing methods that do not require patient-matched normal samples, provided that a pre-built PoN cohort is available. AVAILABILITY: Source code is available for academic use at https://github.com/lvws/PScnv.

DNA Copy Number Variations

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

Absolute copy number aware CNV calling of sub-megabase segments in ultra-low coverage single-cell DNA sequencing data.

Recent advances in ultra-low coverage whole-genome sequencing (WGS) of single cells have enabled detailed analysis of copy number variation at a throughput approaching that of single-cell RNA sequencing. However, downstream computational methods have not seen comparable advances and are largely adaptations of deep sequencing methodology with reduced precision. Here, we present ASCENT, a computational method built to take full advantage of modern direct tagmentation-based WGS at ultra-low depth. Using joint segmentation with high-resolution bins, we accurately detect small segments, achieving accurate copy number profiles even at 100 000 reads per cell. ASCENT implements true absolute copy state inference for single cells, based on statistical modeling of coverage rather than comparison to a reference, while taking variable segment copy state into account. Further, ASCENT implements per-segment copy-neutral loss of heterozygosity (LOH) calling without the need for non-tumor or bulk WGS reference. When applied to a pediatric B-ALL sample, ASCENT finds copy-neutral LOH in a small segment and a minor subclone defined by breakpoints missed in bulk WGS. Thus, by applying appropriate computational methods, single-cell WGS provides clear advantages over bulk, even at a relatively low cell number and sequencing depth.

DNA Copy Number Variations